ControlNet answers the question text cannot: exactly where. You supply structure (the pose of a figure, the depth of a room, the outline of a product) and the model fills it with the look you asked for while respecting the geometry you gave it.
In production the point is repeatability. The same depth map rendered in four registers gives four styles of the same shot, which is how a campaign holds its layout while changing its skin. It is also the cheapest fix for a composition the model keeps refusing to build.
What is ControlNet used for?
Locking composition, pose or layout while letting style vary. Common inputs are pose skeletons, depth maps, Canny edge traces and segmentation masks.
Does ControlNet work for video?
Structural conditioning exists for video too, and it is what keeps motion tied to a real reference. The failure mode moves from wrong geometry to jitter between frames, which is a temporal coherence problem.