Models do not manipulate pixels; they manipulate a much smaller encoded representation and decode it at the end. That is why generation is tractable at all, and why a small change to a prompt can produce a large change to a frame. You have moved to a different neighbourhood.
The practical use is interpolation. Moving smoothly between two points in latent space is what produces a morph, a style blend or a coherent transition, and it is the mechanism underneath first–last frame video.
Why does changing one word change the whole image?
Because the word moved the conditioning to a different region of latent space. Nearby regions look similar; distant ones do not, and prompts do not map to distance in an intuitive way.