arXiv Preprint 2610.12230


From Prompting to Composing:
A Spatial Canvas Interface for Poster Generation

Introducing Compo, a poster generation model adapted from a pretrained image editing model without task-specific architecture design to interpret Spatial Canvas Interface and their associated Text Specifications.

Yitong Wang1,2* Fangyun Wei2* Jinjing Zhao3,2 Sirui Zhang4,2
Hongyang Zhang5 Dong Chen2 Bo Dai6† Yan Lu2

*Equal Contribution †Corresponding Author

Compo teaser image

Spatial Canvas and Text Specifications as Standard TI2I Interface

Our interface supports four complementary forms of binding: Semantic Binding associates a textual concept with a target region; Identity Binding associates a reference image with a region to preserve its visual identity; Text Binding specifies both an exact text string and its desired location; and Pixel Binding directly places visual content that should be preserved.
All these elements, including bounding boxes, reference images, and textual identifiers, are directly rendered onto the Spatial Canvas to form a single unified image input.

Abstract

Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words. We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance. Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications. Compo supports both direct inference, where users explicitly construct the canvas, and agentic mode, where a high-level request is automatically translated into a planned Spatial Canvas. To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch. We further introduce a benchmark that evaluates adherence to individual binding types and their joint composition. Experiments show that Compo achieves stronger compositional controllability than both general-purpose image generation models and dedicated poster generation systems while maintaining high visual quality. By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.

Data | Training | Inference | Visualization

Data Generation and Annotation

Extracted architecture crop showing the prompt-based and direct video relighting route

(a) Poster Creation. We use GPT-5.5, Claude Opus 4.6, and Gemini 3.5 Flash to build content and style repositories. Sampled themes, items, and styles are organized by GPT-5.5 into structured poster plans, including text, background, objects, and their interactions, and rendered into diverse posters with Qwen-Image-2512, Z-Image, and ERNIE-Image.

(b) Annotation Generation. For each poster, SAM-3 segments visual instances, FLUX.2-klein-9B relights or reposes them when needed, and Chandra-OCR-2 extracts text and locations. All these annotations, including bounding boxes, reference images, and textual identifiers, are directly rendered onto the Spatial Canvas, while the structured plans provide the corresponding Text Specifications.

(c) Compo-200K. Each sample pairs a generated poster with its Spatial Canvas and Text Specifications, yielding approximately 200K training samples with fine-grained structured annotations.

Training

Extracted architecture crop showing the prompt-based and direct video relighting route

(a) Supervised Fine-Tuning (SFT). We fine-tune LongCat-Image-Edit and Qwen-Image-Edit-2511 on Compo-200K, where each sample pairs a target poster with its Spatial Canvas and Text Specifications, enabling Compo to understand structured composition inputs.

(b) DiffusionNFT as Reinforcement Learning (RL). We further refine the SFT models with DiffusionNFT using a binding-aware multi-reward objective, combining text, instance similarity, and layout rewards to improve text rendering, instance fidelity, and spatial adherence.

Inference

Extracted architecture crop showing the environment-map and render-based relighting route

(a) Direct Inference. Users manually construct a Spatial Canvas and specify the corresponding Text Specifications, providing explicit control over element positions, binding types, visual references, and generation requirements.

(b) Agentic Mode. Users can also provide a high-level natural-language request. The agent automatically extracts the required elements, plans their layout, selects appropriate binding types, and retrieves visual references when needed. It then constructs a Spatial Canvas and Text Specifications, which are passed to Compo to generate the final poster.

Visualization

Extracted architecture crop showing the environment-map and render-based relighting route

Visualizations at various resolutions across diverse trending themes and visual elements. Our Compo models can generate visually coherent and plausible poster designs based on diverse trending themes and visual elements that are unseen during training, demonstrating robust generalization and practical applicability.

Results

Comparison with representative baselines

Extracted architecture crop showing the prompt-based and direct video relighting route

Visual comparison with representative baselines (Part 1). We compare Compo with dedicated poster generation methods, including PosterMaker, PosterOmni, and CreatiDesign, as well as HiDream-O1-Image, the best-performing general text-and-image-to-image baseline in our benchmark. Transparent pink boxes highlight noticeable errors.

Comparison with representative baselines

Extracted architecture crop showing the environment-map and render-based relighting route

Visual comparison with representative baselines (Part 2). We compare Compo with dedicated poster generation methods, including PosterMaker, PosterOmni, and CreatiDesign, as well as HiDream-O1-Image, the best-performing general text-and-image-to-image baseline in our benchmark. Transparent pink boxes highlight noticeable errors.

More Results

Input and Output Gallery

The following results showcase a series of Spatial Canvas inputs and the corresponding posters generated by our Compo models.

For clarity and visual consistency, the Spatial Canvas examples shown in the pipeline diagram and qualitative comparisons above are redrawn with a white background to match the overall white visual style of the paper.

In practice, however, the Spatial Canvas used during dataset curation and model training follows the format shown below. These examples therefore reflect the actual inputs directly received by Compo and their corresponding outputs. More details are included in the appendix of our paper.

To facilitate the construction of the Spatial Canvas conveniently, we reuse some segmented visual assets. Due to occlusion, certain assets may be partially incomplete, but this generally does not affect the overall functionality of the model.

Extracted architecture crop showing the environment-map and render-based relighting route

Resources

Citations

BibTeX
@misc{wang2026compo,
      title={From Prompting to Composing: A Spatial Canvas Interface for Poster Generation}, 
      author={Yitong Wang and Fangyun Wei and Jinjing Zhao and Sirui Zhang and Hongyang Zhang and Dong Chen and Bo Dai and Yan Lu},
      year={2026},
      eprint={2610.12230},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.12230}, 
}