You already know the tell. Skin with no pores, a shadow falling from a light source that doesn't exist, a logo on a cap that's almost the real one but not quite. It's not that the model doesn't look like a person. It's that nothing in the frame agrees with anything else, and your eye clocks the disagreement before your brain can name it.
It isn't a resolution problem
Every serious image model can render a convincing face at this point. The failure isn't in the model's capability, it's in what nobody bothered to specify. "Generate a woman wearing this jacket" is not a brief. It's a coin flip on lens choice, light source, skin texture and eight other decisions that a real casting and a real photographer would never leave unmade.
Real photographs have flaws. Sensor grain in low light. A slight motion blur if the shutter speed didn't quite keep up. Skin that catches light unevenly because skin is not a flat surface. Renders don't have those flaws unless someone specifies them, and specifying them is a photography skill, not a prompting trick.
The four things that actually fix it
Camera and lens spec, stated explicitly. "85mm, f/2, shot on a body that renders skin the way a real portrait lens does" produces a different image than a generic instruction, and the difference is visible immediately.
Real skin detail, not smoothed skin. Pores, minor texture, uneven tone where light grazes the face. The instinct to make everything flawless is exactly what makes AI imagery read as AI.
One light logic, held across the whole set. A single key light direction and a single fill ratio, applied consistently frame to frame, is what separates a coherent shoot from six unrelated images that happen to feature the same outfit.
A locked model. Casting a model once, as a matched headshot and full-body pair, then reusing that exact person across every future shoot, is what makes a brand's AI imagery start to look like a brand rather than a slot machine.
Garments are the harder problem
Faces get most of the attention, but garments are where AI fashion imagery usually falls apart first. Back views, ribbing, seam placement, how a fabric actually drapes over a shoulder: these are the details that separate a convincing product shot from one that quietly undermines trust in the product itself. That only holds if the actual garment photos are the reference for every generation, not a text description of the garment.
Consistency is a direction problem, not a prompting problem. That's the entire discipline, and it's the difference between a set of images and a brand.
None of this requires believing AI model photography is flawless. It requires the same thing a real casting shoot has always required: someone making deliberate decisions before a single frame gets produced, and holding every subsequent frame to those decisions.