Back to glossary

ControlNet

A technique for conditioning a diffusion model's output on structure — like pose, depth, or edges — instead of a prompt alone.

ControlNet conditions a diffusion model's output on an extra structural input — a pose skeleton, a depth map, an edge outline — so generation follows that structure precisely rather than relying on the prompt alone to describe composition.

It's the difference between asking for "a person waving" and handing the model an exact skeleton of a waving pose to follow. That extra control is especially valuable for video, where consistent framing across frames matters more than in a single image.

See also: LoRA, Temporal Consistency.

Put these ideas into motion.

Download

Frequently asked questions

ControlNet is a technique that conditions a diffusion model's generation on an extra structural input — a pose skeleton, a depth map, an edge outline — so the output follows that structure precisely instead of relying on the prompt alone.

Prompts alone give loose control over composition and framing. Feeding a model a structural reference lets creators pin down exactly where a subject is and how it's posed, frame after frame.

Ready to tell your story?

Describe an idea and watch the agent animate it — on your Mac, in minutes.