The Product-Fidelity Checklist: 12 Tests Before You Trust an AI Photo Tool
By the ORA Lab team · Updated 25 August 2026 · 9 min read
Key takeaways
- Vendor galleries prove what a tool did once for someone else's products; the only evaluation that matters is a structured test on your own SKUs, and it takes one afternoon.
- The 12 tests split into four groups: identity (does the product survive?), physics (does light behave?), consistency (does a catalog cohere?), and workflow (does it survive contact with your team?).
- Test with your three hardest SKUs, not your easiest, a multi-stone piece, a reflective surface, a textured material. Tools that pass the hard cases pass everything.
- Score pass/fail, not impressions: count stones at 200% zoom, lay outputs beside the source, check the same prompt twice. Feelings about beauty are how bad tools get bought.
- Any tool failing an identity test is disqualified regardless of how it scores elsewhere, beautiful-but-wrong is the most expensive failure mode in commerce imagery.
Every AI photo tool demos beautifully. That's not cynicism, it's selection: the gallery you're shown is the best output the vendor ever produced, on products chosen because they generate well. Your evaluation problem is different, you need to know what the tool does to your products, on an ordinary Tuesday, in the hands of whoever runs your listings. The good news is that this is testable, cheaply, in an afternoon, before any contract is signed.
What follows is the test protocol we recommend buyers run on every tool they shortlist, including ours. It's twelve tests in four groups, each with a concrete pass condition, designed around the failure modes that actually surface after purchase: products that quietly change, light that doesn't add up, catalogs that drift apart, and workflows that die in week three. Bring three SKUs and a free trial; leave with a decision you can defend.
Before you start: pick the right three SKUs
The single biggest evaluation mistake is testing with easy products. A matte ceramic vase generates beautifully in every tool on the market; it tells you nothing. Choose your three hardest: something countable (a multi-stone ring, a tufted headboard with visible buttons), something reflective (polished metal, glass, high-gloss lacquer), and something textured (woven cane, brushed finish, engraved detail). Photograph each cleanly against a plain background, the capture checklist covers this in ten minutes, because a fair test requires a fair input. Then run all twelve tests on each tool with identical inputs.
Group 1, Identity: does the product survive?
- 1.The count test. Generate three scenes from your countable SKU. Zoom to 200% and count everything countable, stones, prongs, buttons, links, seams, against the source photo. Pass: every count matches in all three outputs. One missing pavé stone is a fail, not a rounding error; it's a misrepresented product on a listing.
- 2.The silhouette test. Overlay or eyeball the product's outline against the source at matched angles. Proportions, curvature, and hardware placement must hold. Pass: you cannot find a shape difference a customer could photograph in a complaint.
- 3.The material test. Check that gold stayed the same gold, the stone kept its cut and saturation, the fabric kept its weave. Pass: a colour-accurate screen shows no drift a merchandiser would flag.
- 4.The regeneration test. Run the identical prompt on the identical input twice. The scene may vary; the product may not. Pass: both outputs depict the same object, this is the fastest way to expose tools that regenerate rather than preserve.
Group 2, Physics: does the light add up?
- 1.The shadow test. Find the scene's implied light source, then audit every shadow: does the product's shadow fall the right way, with softness matching the light's character, connecting the object to the surface it sits on? Pass: no floating products, no shadows pointing two directions.
- 2.The reflection test (critical for jewellery, glass, and gloss). Reflective surfaces must mirror the scene they sit in, a ring on a marble slab should carry marble tones in its band, not a phantom studio. Pass: reflections reference the actual environment.
- 3.The scale test. Place the product beside implied references (hands, furniture, rooms). Pass: dimensions read true, the 12mm pendant doesn't render statement-sized, the armchair doesn't dwarf its side table.
Group 3, Consistency: does a catalog cohere?
- 1.The batch test. Generate the same scene style for all three SKUs. Pass: the outputs read as one photographer's work, same light character, same grade, same mood, because a catalog is a body of work, not a stack of one-offs.
- 2.The revision test. Take one output and request a specific change: 'same scene, warmer light' or 'move the camera lower.' Pass: the tool changes what you asked and only what you asked. Tools that regenerate everything on every edit turn art direction into a slot machine.
- 3.The format test. Check what the tool exports against what your channels need, marketplace-compliant white-background mains, social crops, campaign ratios. Pass: your channel requirements come out without a manual editing pass appended to every image.
Group 4, Workflow: does it survive your team?
- 1.The operator test. Have the person who will actually use the tool, not the founder, not the agency, produce one listing-ready image unassisted. Pass: they succeed within an hour, and would volunteer to do it again. Tools that require a prompt specialist have a hidden headcount cost.
- 2.The economics test. Price a realistic month: your SKU count × images per SKU × revision rate (assume 2 to 3 attempts per keeper while learning). Pass: the real per-usable-image cost, not the advertised per-generation cost, fits your content budget.
Scoring and deciding
| Result | Verdict | Action |
|---|---|---|
| Any Group 1 fail | Disqualified | No price or polish compensates for a changed product |
| Group 1 pass, physics fails | Risky | Usable for drafts; expect manual QA on everything customer-facing |
| Groups 1 to 2 pass, consistency fails | Single-shot tool | Fine for one-off images; will fight you at catalog scale |
| Groups 1 to 3 pass, workflow fails | Right tool, wrong fit | Revisit if pricing or operator experience changes |
| 12/12 | Adopt | Run a 20-SKU pilot and measure listing performance |
Two habits make the scorecard honest. First, write results down as you go, memory smooths over failures, especially when the outputs are pretty. Second, run the identical protocol on every shortlisted tool in the same sitting with the same three SKUs; sequential evaluations weeks apart measure your mood as much as the tools. The whole protocol takes about ninety minutes per tool, which is roughly the cost of one badly chosen subscription's first week.
We publish this checklist because we're confident about where ORA lands on it, identity anchoring is the part of the problem we built the company around. Run all twelve tests on us alongside anyone else on your shortlist: , bring your three hardest SKUs, and we'll run the protocol with you live.
Frequently asked questions
- How do I test if an AI product photo tool is accurate?
- Run identity tests on your own hardest SKUs: generate three scenes, zoom to 200%, and count every countable detail (stones, prongs, links, buttons) against the source photo; check silhouette and materials; then run the identical prompt twice and confirm the product doesn't change between runs. Any failure disqualifies the tool for listing imagery.
- Which products should I use to evaluate an AI photography tool?
- Your three hardest, not your easiest: one countable product (multi-stone jewellery, tufted furniture), one reflective one (polished metal, glass), one textured one (weave, engraving, brushed finish). Easy products generate well in every tool and reveal nothing about where a tool breaks.
- What is the fastest single test of an AI photo tool?
- The regeneration test: run the identical prompt on the identical source photo twice. Preservation-based tools return the same product in a varied scene; regeneration-based tools return two subtly different products. It takes five minutes and exposes the architectural difference that matters most for commerce.
- How long does a proper AI tool evaluation take?
- About ninety minutes per tool once your three test SKUs are photographed: four identity tests, three physics tests, three consistency tests, and two workflow tests, each with a written pass/fail. Evaluate all shortlisted tools in one sitting with identical inputs so the comparison is fair.
- What does it mean if a tool makes beautiful images but fails the count test?
- It's an invention-based generator: it produces a plausible new product rather than preserving yours. That's fine for moodboards and concepting, but disqualifying for listings and campaigns, where the image is a factual claim about the item a customer will receive, and the mismatch surfaces as returns and marketplace flags.