The practical gain is that a brief, a reference frame and a voice note can go into the same conversation and come out as a shot list, a plate and a script. The handoffs that used to lose information between tools happen inside one context instead.
The practical risk is the same thing. A model that can see your reference will also confidently describe what is not in it, so multimodal input does not remove the need for a claim gate. It moves it earlier.
What can multimodal models do that text models cannot?
Read a frame and act on what is in it: critique a composition, extract a palette, match a reference, or take a rough sketch and turn it into a specification.