Overview
Text prompts describe desired content but do not precisely place every body part, object, edge, or depth relationship. ControlNet adds a structural reference that influences the generation. It can guide pose, depth, outlines, line art, scribbles, or segmented regions while the positive prompt controls the new visual content.