Watch enough generative video and you learn to feel the moment it starts to go. Nothing obvious happens. The shot simply stops being convincing somewhere around the two-thirds mark, and if you scrub back you find a hand that gained a finger during a gesture, or hair that settled into a shape it could not have reached, or a background object that quietly changed material.
That is temporal coherence failing, and it is the defining production constraint of this medium in 2026.
Why it happens
A video model is not simulating a world and photographing it. It is producing a sequence of frames that are each plausible and mutually consistent, under a constraint that gets harder to satisfy the further it travels from its conditioning.
Early frames are anchored: to the reference, to the first frame, to the prompt. Later frames are anchored mostly to earlier generated frames, so small errors compound. Nothing catastrophic occurs at any single step, which is precisely why the failure is hard to spot. It is a slow drift, not a break.
The failure order
The degradation is not random. Things go in roughly this order, which is useful because it tells you what to look for and when.
| ORDER | WHAT DRIFTS | HOW IT READS TO A VIEWER |
|---|---|---|
| 1 | Fine texture: fabric weave, skin pore, hair strand | A slight softening nobody consciously notices |
| 2 | Small rigid objects: jewellery, buttons, cutlery | Something in the frame is subtly wrong |
| 3 | Extremities: fingers, feet, ears | The obvious tell, and the one people name |
| 4 | Physical logic: gravity, contact, occlusion | The shot stops feeling real |
| 5 | Identity: face structure, proportion, age | It is no longer the same person |
| 6 | Scene topology: geometry, object persistence | The space is incoherent |
The practically important line is three. Extremity failure is the first one an untrained viewer reliably catches, which means your usable clip length is the point just before it, not the point where the model stops producing frames.
The technique: cut before the drift
Cinema solved the problem of shots that cannot run forever a century ago, and the solution was editing. Generative production inherits the answer wholesale.
- Generate longer than you need. Ask for eight seconds when you want three. The cost difference is trivial and it gives you choice.
- Take the stable opening. The first portion is anchored hardest and therefore best.
- Cut on motion. A cut during a movement hides the discontinuity between two independently generated clips, which is the same reason editors have always cut on action.
- Vary shot size across the cut. Two similar-sized shots joined together announce the join. Wide to close does not.
- Never cross-dissolve two generative clips. A dissolve holds both images on screen simultaneously, which is exactly the condition under which an audience compares them and notices they do not match.
What extends usable length, and what does not
Things that genuinely help:
- Less motion in frame. A slow move on a mostly static subject holds far longer than a subject in complex action.
- Fewer articulated objects. Hands are the enemy of length. A composition where hands are out of frame or still buys seconds.
- A simpler background. Every additional object is another thing that has to persist.
- Locked-off camera. Camera movement compounds with subject movement and halves the budget.
- First-and-last-frame conditioning where the model supports it, which re-anchors the end of the clip rather than letting it float.
Things that do not help, despite being widely recommended: adding "consistent, coherent, stable" to the prompt, raising resolution, and generating the same clip repeatedly hoping for a longer stable window. The third one is expensive and the ratio does not improve with attempts.
Testing for it before the client does
Two checks, neither of which takes long.
First, scrub the clip backwards at speed. Drift that is invisible forwards is obvious in reverse, because the eye is no longer being carried by the motion and starts comparing frames instead.
Second, pull the first and last frame and put them side by side at full size. Anything that has changed between them and should not have is your drift, quantified. If the face is different, the clip is dead regardless of how good the middle looks.
Where this is going
Coherent windows have lengthened steadily, and models with native scene-extension and reference conditioning now chain generations well past the single-clip limit. The drift has not been eliminated; it has been pushed further out and made more graceful.
The production discipline does not change when the window lengthens. Shot-length discipline, cutting on motion, varying shot size and checking the first frame against the last are just editing, and editing was always the part that made sequences work.
The short version, with the related terms and the questions people actually ask about it.
READ THE GLOSSARY DEFINITION →