What the stand-in technique is
The stand-in technique is the camera route for creating a product reference that doesn't exist yet. When you have the physical product but no photograph of the state you need (worn on a face, half-folded, tilted just so, three items arranged), you photograph anyone (yourself, a friend, your partner in their pajamas) wearing, holding, or arranging the product, then use that photo purely as a fit and pose reference: proportions, scale, position. The identity comes from elsewhere. Use it when the product is physically in your hands and no existing photo shows the pose you need. It is one of three ways to create a missing reference, alongside pose-match (the AI route, when photos are all you have) and the sketch (the drawn route, when you have neither). The camera route has the highest truth of the three, because it is a real photograph of the real product in the real pose.
The reason it works is a limitation of the model, stated plainly: AI is bad at making an object take a position it doesn't see. It was never shown the folded state, the worn state, the tilted state, so it invents the hidden geometry, and it invents it wrong. The stand-in removes the guessing. You give the compositor a real photograph of the exact arrangement, and it only has to swap the identity in, not imagine the shape from nothing.
Pick by what you're holding
The stand-in is one of three siblings in the "create the missing reference" family, and the choice between them is decided by what is physically in front of you, not by preference. Holding the physical product? Stand-in, the camera route: photograph anyone wearing or arranging it, and let identity come from elsewhere. Only have photos, no product in hand and no reshoot possible? Pose-match, the AI route: have the model build the missing angle. Have neither, and words can't carry the position you need? Sketch, the drawn route: a rough doodle of the placement. One question, three answers.
The trade-off is truth versus availability. The stand-in has the higher truth of the three, but it needs someone to actually hold the physical product, which in the common freelancer case is nobody. Pose-match needs no physical product at all, only a clean photo, which is why it's the route you fall back to when a stand-in shot is one nobody can take. This page is the deep dive on the camera route; the route map is where the full diagnosis lives, and it also points back to the stand-in as the bluntest recovery move of all: no usable photo at all? Take one.
How to shoot it: the product is real, everything else is disposable
The whole trick is that the AI doesn't care about the stand-in, only the product. Your partner in their pajamas works. The wrong person, the wrong face, the wrong outfit, the wrong room, none of it matters, because you are only extracting the fit and the pose from the shot, not the identity. Put the real product on a real body in the real position, photograph it (an iPhone is enough), then in the prompt keep the product from your shot and replace the identity from another image. The instruction reads like this, verbatim: "Using the face and glasses from [IMAGE1] as a base for the composition."
So the stand-in shot is a scaffold, not a final image. It carries three things the model struggles to invent: where the product sits relative to the body, how big it is against real anatomy, and what pose it holds. Those are exactly the failure modes that wreck a from-scratch generation. Once the model has a true photograph of that arrangement, the rest of the job (rendering a specific model, dropping in a scene) is a swap it can do, not a guess it has to make.
The product-only version: fold it yourself
The stand-in isn't only for products worn on a body. It works product-only too, and the reason is the same one-line rule: AI is bad at making an object take a position it doesn't see. If you need the glasses half-folded, or tilted at an angle no photo captured, don't ask the model to imagine that pose. Fold or tilt the glasses yourself, shoot that as the arrangement reference, and feed it in. You've replaced an invention the model gets wrong with a photograph of the real thing in the real position.
This is the cleanest case for the technique, because there's no identity to swap at all: the object in your reference and the object in the output are the same object. If the pose in your shot is already the pose you want and you only need a new background or surface, you've crossed into lock-and-outpaint territory: freeze the exact pixels and build the world around them for zero product drift. Use the stand-in when the pose itself is what you can't get any other way.
The advanced version: 4-image role assignment
When you need a specific model wearing specific glasses accurately, the stand-in becomes a four-image job with one rule per image: separation of concerns. Each image does exactly one job, so nothing blends. Image one is the real person wearing the glasses (your stand-in shot), and it carries fit alone: proportions, face-to-frame ratio, where the frames sit relative to the eyebrows. Image two is the model's headshots, and it carries identity: who to render. Image three is the environment, carrying setting. Image four is the product on white, carrying design: the exact shape, materials, and details.
It works because fit is not design, and an explicit identity anchor ("render [image2]'s face") stops the model from blending the stand-in's face with the model's. Close-up framing forces the model to prioritize the glasses-face relationship over everything else. Always attach the fit rules on top: frames below the eyebrows with a visible gap of skin, roughly 70 to 75 percent of the face width, sitting where a real pair would sit. The stand-in gives the true placement; the fit rules make sure the render keeps it.
The precondition nobody mentions: someone has to hold it
There is one hard precondition on the whole technique, and it's the thing most write-ups skip: someone must actually hold the physical product. Often that someone is the client, not you. In the common freelancer case, you have the client's photos and a spec sheet, but the physical frames are on another continent, and there is nobody to put them on a face and shoot them. When that's true, do not advise a stand-in shot nobody can take. It is the wrong route for that situation, however good it is on paper.
The fallbacks when nobody can hold the product are clear. Use lock-and-outpaint if the pose you already have is usable. Reuse a proven prompt from a similar past project. Or generate-and-confirm the missing reference with pose-match, then have whoever knows the product best approve it. The stand-in is the highest-truth route only when its precondition is met; the moment it isn't, availability beats truth, and the AI route is the honest recommendation.
An eyewear account: the flagship application
A premium eyewear account is the clearest real-world use of the technique. The brand makes acetate fronts with titanium temples, and for their product drops Bertrand photographed the physical frames himself on an iPhone, deliberately imitating competitor compositions, then had Dezygn clean, enhance, and uniform them. The stand-in work was the composition itself: as he put it, "the composition I had stolen from a competitor" did the work. The competitor's shot supplied the arrangement; the real frames supplied the product; the AI supplied the finish.
That established a repeatable recipe, now saved in the Awa campaign instructions for the account: remove lens stickers, uniform the background to the brand's #ECE8E7, diffused cool studio light, preserve exact frame details, four to five variations at 2K in 4:3, floating with a subtle shadow. The lifestyle side follows the same angle-matching discipline the technique demands: a frontal pose gets a front-view reference, a three-quarter pose gets a three-quarter reference. The stand-in shot is what makes that matching possible, because you can pose the real frames to the exact angle the layout needs.
The roofing hardware job: when you can't shoot, borrow reality
The stand-in has a sibling move for when the product is a technical assembly, and a roofing hardware job is where it got confirmed. The client makes a rooftop pipe-support product (a rubber block, strut, and clamps), and the hero shot needed the whole load path to read correctly: pipe to clamp to strut channel to threaded rods to block to roof. The AI failed three ways in a row: editing clamps into the shot, regenerating from four ingredients with the load-path logic spelled out, and simplified packshots all broke. The diagnosis was blunt: "The AI has little knowledge of this product arrangement in its training dataset."
The fix was the stand-in principle applied through a found photograph instead of a self-shot one. Rather than shoot a twin, the move was to find a photo of a similar real arrangement, use it as the base, and swap in the product with a simple replace prompt, so the logical structure was inherited from reality rather than invented. Then lock-and-outpaint the low-angle sky hero. It took six minutes once the reference was found. The lesson, in Bertrand's words, is that he "spent way too much time before switching to the reference image." For an unfamiliar technical assembly, a real reference image replaces the spatial reasoning the model doesn't have, whether you shoot that reference or find it.
Key Takeaways.
- The stand-in technique is the camera route for creating a missing reference: photograph anyone holding or wearing the real product, then use that shot purely as a fit and pose reference while identity comes from elsewhere.
- It exists because AI is bad at making an object take a position it doesn't see; a real photograph of the arrangement removes the guessing.
- Pick by what you're holding: physical product means stand-in, photos only means pose-match, neither means a sketch. The stand-in has the highest truth of the three.
- The AI doesn't care about the stand-in, only the product; even your partner in their pajamas works, because you extract fit and pose, not identity.
- Hard precondition: someone must actually hold the product, often the client, not you; when nobody can, fall back to lock-and-outpaint, a reused prompt, or pose-match.
- For unfamiliar technical assemblies, a real reference image (shot or found) replaces the spatial reasoning the model doesn't have, as the roofing hardware job confirmed.
Ready to Put This Into Practice?
Dezygn gives you the AI creative tools, training, and community to turn these insights into real results for your clients.
Start FreeRelated Resources.
How to Get Product Accuracy with Nano Banana
Product accuracy in AI product photography is a discipline: name the defect, diagnose it, route it to the right technique, judge the output. Measured, not luck.
Read guidePose-Match: Directing AI Models Without Losing the Product
What pose-match is and when to use it: the route for creating a product reference that doesn't exist yet, plus the eval that says use one jump, not a chain.
Read guideThe Product Accuracy Route Map: Diagnosis to Technique
How do you fix an inaccurate AI product image? Name the defect on one of 8 fidelity axes, then route to the exact technique that clears it. The full map.
Read guide