The AI Product Photography Glossary: 40 Terms, Plainly Defined
By the ORA Lab team · Updated 18 September 2026 · 10 min read
Key takeaways
- Shared vocabulary is a working tool: teams that name things precisely brief faster, QA sharper, and see through vendor marketing quicker.
- The forty terms group into five conversations: how generation works, how fidelity is judged, how images are crafted, how video behaves, and how the work is bought and run.
- A handful of these terms carry entire decisions: anchoring versus reference, keeper rate, regeneration test, and cost per usable image decide most tool purchases on their own.
- Most vendor claims translate into three or four of these words, and the translation is the evaluation: ask which term a feature actually is, and the demo gets much shorter.
- Every definition links to the article that treats it fully, so this page doubles as a map of the whole library.
Every maturing field reaches the moment where its vocabulary needs writing down, not because the words are difficult but because half the confusion in meetings turns out to be two people using one term for different things, or different terms for one thing. AI product imagery is past that moment: buyers compare 'AI photoshoots' against 'product staging' against 'virtual photography' and the tools behind the labels may be identical or opposites. This glossary is the working vocabulary as we use it across this library: forty terms, plain definitions, grouped by the conversation they belong to, each linked to the article that treats it properly.
How generation works
- Generative AI imagery: images produced by a model that learned visual patterns from training data, rather than captured by a camera or built from 3D geometry.
- Diffusion model: the dominant architecture, which generates by removing noise step by step until an image emerges, each step guided by your inputs. Explained fully in our marketer's explainer.
- Training data: the image-and-caption pairs a model learned from. The model keeps statistical patterns, never the images themselves.
- Latent space: the model's compressed internal representation of images, where prompts locate neighbourhoods of similar-looking outputs rather than exact addresses.
- Prompt: the text input describing what to generate. In product work, best used for the scene, never the product, per the prompt-writing method.
- Conditioning: any input that steers generation, text, reference images, style locks, or a locked first frame. Every tool feature is a form of it.
- Reference image: an image supplied to pull output toward its look. Conditions strongly but guarantees nothing: results are 'inspired by', not 'identical to'.
- Anchoring: the commerce-grade step up from reference, where the product's geometry and materials are held as a constraint the generation must preserve. The core idea behind product fidelity.
- Hallucination: a generated detail with no basis in the input, the model filling an unpinned gap with a statistical guess.
- Regression to the average: the tendency of unconstrained generation to drift toward the most typical version of a thing, the mechanism behind lost stones and generic drape.
How fidelity is judged
- Product fidelity: how exactly a generated image preserves the real product's identity: geometry, materials, colour, and countable detail.
- Count test: zooming into a generated frame and counting everything countable (stones, links, buttons, stitches) against the source photo. The first gate in the 12-test protocol.
- Regeneration test: running the identical input twice; anchored systems return the same product in varied scenes, reference-based systems return two subtly different products.
- Label integrity: exact reproduction of printed text, logos, and typography on packaging and dials, the beauty and watch categories' version of the count test.
- Impossible light: reflections or shadows that contradict the scene's light source, the reliable tell on reflective products.
- Phantom reflection: a mirror surface showing the source photo's environment (studio white, lightbox edges) instead of the generated scene.
- Scale anchor: an in-frame reference (doorway, rug, wrist, collarbone) that forces and verifies true product size, load-bearing in furniture scenes.
- Drape truth: fabric falling as that fabric genuinely falls, in the size shown, the apparel fidelity axis defined in the flat-lay guide.
- True-colour grade: finishing with minimal stylistic colour bias, mandatory where colour is the purchase decision (shades, leathers, marketplace mains).
- Keeper rate: the fraction of generations good enough to ship. The hidden multiplier behind every real cost calculation.
How images are crafted
- Packshot: the clean product-on-plain-background photograph, the workhorse of listings and the usual source photo for generation.
- Ghost mannequin: apparel photographed on a mannequin that is edited out, leaving an invisible-body shape. Compared against worn imagery here.
- On-model / worn frame: the product shown on a person, real or generated, where scale and desire live.
- Virtual try-on: the umbrella term for four different technologies (live AR, selfie try-on, seller-side generation, marketplace features), mapped in the try-on landscape.
- Scene brief: the six-layer specification of a generated setting: product context, setting, surface, lighting, camera, grade.
- Hint: a one-line statement of intent given to a system that already holds the brand's locked look, the art-direction alternative to per-image prompting.
- Visual system / locked look: a brand's standing imagery decisions (surfaces, light character, grade) defined once and applied to every generation.
- Translation sheet: the one-page distillation of a moodboard into five enforceable decisions, per the campaign method.
- Batch consistency: a set of images reading as one photographer's work, the property that separates catalogs from image piles.
- Upscaling: enlarging an image with AI after generation. Useful, but can invent micro-detail, so fidelity-critical zoom areas need rechecking.
How video behaves
- Image-to-video: generating motion from a photograph that becomes the clip's first frame, the commerce-safe mode per the comparison.
- Text-to-video: generating motion from a sentence alone; unmatched for concepting, disqualified for product footage because the product is invented.
- Temporal consistency: whether the product stays exactly itself across every frame, video's version of the count test.
- Motion brief: the three-line video instruction: one camera move, what is alive in the scene, and intensity, per the photo-to-video workflow.
- Camera move: the named viewpoint motion (push-in, orbit, parallax drift, tilt-up) that defines a clip's character.
- Seamless loop: a clip whose last frame hands back to its first, the format product pages and reels reward.
How the work is bought and run
- Credit: a metering unit consumed per action, not per image; upscales, revisions, and video seconds draw different amounts, decoded in the pricing guide.
- Cost per usable image: total spend divided by shipped keepers, the only number that compares tools honestly.
- Content velocity: the sustained rate at which a brand ships fresh imagery, the operating metric generation actually changes.
- Commercial usage rights: what a tool's terms let you do with outputs (listings, ads, packaging), and whether those rights survive cancellation, covered in the licensing guide.
- Training opt-out: the setting that keeps your uploads from being used to improve the vendor's models, a hard requirement for unreleased collections.
- Zero-inventory visuals: launch-grade imagery generated from one physical sample to sell before manufacturing, with the one-sample rule as its integrity line.
Vocabulary is cheap leverage: an afternoon with these forty terms makes every future briefing shorter and every demo more transparent. If you want to watch the load-bearing ones proven live rather than defined, and bring your hardest SKU: anchoring, the count test, and keeper rate all reveal themselves inside the first twenty minutes.
Frequently asked questions
- What is anchoring in AI product photography?
- Anchoring means the system treats your product photo as a constraint the generation must preserve: geometry, materials, and countable details are held fixed while only the scene is generated. It differs from a reference image, which merely pulls output toward a similar look. The regeneration test (same input twice) reveals which a tool does.
- What is keeper rate and why does it matter?
- Keeper rate is the fraction of generations good enough to ship. It matters because it multiplies every price: a tool at half the per-generation cost but a third of the keeper rate is more expensive per usable image. It varies more between tools than sticker prices do, which is why trials on your own products beat pricing-page comparisons.
- What is temporal consistency in AI video?
- Whether objects, especially your product, stay exactly themselves across every frame of a generated clip. It is video's version of the count test: a clip that warps a ring at second four is misrepresenting the product sixty times a second. QA means scrubbing slowly and watching the product, not the scene.
- What is the difference between a prompt and a hint?
- A prompt specifies a scene's implementation in detail, surface, light, camera, and grade, and gets retyped per image. A hint states intent in a sentence ('terrace at dusk, warmer') and relies on a system that already holds the brand's locked visual identity. Prompting is a per-image skill; hints are art direction over a standing system.
- What does hallucination mean in AI product images?
- A detail in the output with no basis in your input: the model filled a gap with a statistical guess drawn from its training patterns. Harmless in scene elements, disqualifying on the product itself, where a hallucinated stone, label character, or seam misrepresents what the customer will receive.