Describing a body position in words is unreliable and always has been. A skeleton removes the ambiguity: the model gets the geometry directly and spends its capacity on everything else.
For synthetic presenters this is the difference between a performance you directed and one you accepted. Extract the pose from a reference performance, apply it to your trained identity, and the gesture is yours rather than the model’s idea of the word "gesturing".
Hands remain the hard case. A pose skeleton fixes where the hand is and not how many fingers it has, so hands in frame still need the same gating they always did.
Does pose conditioning fix hands?
It fixes where the hand is, not what it is made of. Finger count and contact points still need checking at full resolution.
Where does a pose reference come from?
Any footage of the movement you want, including a phone video of somebody in the office doing the gesture. The extraction keeps the geometry and discards the person.
