The format is unforgiving because there is nothing else in frame to distract from an error. Every drift in the silhouette, every misrendered character on a label, every physically impossible reflection is the subject of the shot.
Which is exactly why it is the format where a plate-locked pipeline earns its cost most visibly. A product cinematic generated from a description will be beautiful and wrong. One generated from a verified plate will be correct, and correctness is what the client is buying.
Keep readable type out of frame wherever the composition allows, and composite real type in post where it does not.
Why are product cinematics harder than they look?
There is nothing else in the frame. Every error is on the subject, at full size, and the audience is looking directly at the thing that is wrong.
How do you handle packaging text in a product cinematic?
Compose it out of frame where possible and composite real type in post where not. Asking a model to render a label the audience can read is the lowest-yield thing in generative production.
