Data Generation and Annotation
(a) Poster Creation. We use GPT-5.5, Claude Opus 4.6, and Gemini 3.5 Flash to build content and style repositories. Sampled themes, items, and styles are organized by GPT-5.5 into structured poster plans, including text, background, objects, and their interactions, and rendered into diverse posters with Qwen-Image-2512, Z-Image, and ERNIE-Image.
(b) Annotation Generation. For each poster, SAM-3 segments visual instances, FLUX.2-klein-9B relights or reposes them when needed, and Chandra-OCR-2 extracts text and locations. All these annotations, including bounding boxes, reference images, and textual identifiers, are directly rendered onto the Spatial Canvas, while the structured plans provide the corresponding Text Specifications.
(c) Compo-200K. Each sample pairs a generated poster with its Spatial Canvas and Text Specifications, yielding approximately 200K training samples with fine-grained structured annotations.