This is the single hardest problem in synthetic production and the one that decides whether a piece of work is a demo or a campaign. Audiences tolerate a great deal of oddity in a generated frame; they do not tolerate a presenter whose jaw changes between two adverts in the same set.
The test is mechanical: stack the frames and flick through them at speed. Drift in the jawline, the eye spacing or the hairline shows up immediately, in a way that studying one frame at a time hides. A fail goes back to the identity sheet. It never goes to inpainting, because patching a face teaches you nothing and produces a second face that also drifts.
How do you keep an AI character consistent across scenes?
Train the identity once from a sheet of twenty-plus stills at varied angles under even light, then generate every scene from that trained identity rather than from a fresh description. Reference images alone drift by the fourth shot.
Why does my character’s face change between shots?
Because each shot is being generated from the prompt rather than from a locked identity. Text cannot specify a face precisely enough to reproduce it, and small wording changes move it further.