ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Control
Researchers have introduced ChArtist, a novel domain-specific diffusion model designed to automate the creation of pictorial charts, which blend data visualization with artistic elements for effective storytelling. Addressing the challenge of balancing rigid chart structures with flexible visual aesthetics, ChArtist utilizes a Diffusion Transformer (DiT) architecture. The model features two distinct control mechanisms: spatial control, achieved through a new skeleton-based representation that encodes data structure without rigid outlines, and subject-driven control, which incorporates visual characteristics from reference images. To manage these inputs, the system employs adaptive position encoding and Spatially Gated Attention. The team also compiled a large-scale dataset of 30,000 triplets containing skeletons, reference images, and final charts to support model fine-tuning. Additionally, a unified data accuracy metric was proposed to evaluate the faithfulness of generated charts. This work highlights the potential of generative AI to move beyond general-purpose conditions toward task-specific representations for data-driven visual storytelling, offering a significant advancement in computer vision and automated graphic design.
Wire timeline
ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Control
Researchers have introduced ChArtist, a novel domain-specific diffusion model designed to automate the creation of pictorial charts, which blend data visualization with artistic elements for effective storytelling. Addressing the challenge of balancing rigid chart structures with flexible visual aesthetics, ChArtist utilizes a Diffusion Transformer (DiT) architecture. The model features two distinct control mechanisms: spatial control, achieved through a new skeleton-based representation that encodes data structure without rigid outlines, and subject-driven control, which incorporates visual characteristics from reference images. To manage these inputs, the system employs adaptive position encoding and Spatially Gated Attention. The team also compiled a large-scale dataset of 30,000 triplets containing skeletons, reference images, and final charts to support model fine-tuning. Additionally, a unified data accuracy metric was proposed to evaluate the faithfulness of generated charts. This work highlights the potential of generative AI to move beyond general-purpose conditions toward task-specific representations for data-driven visual storytelling, offering a significant advancement in computer vision and automated graphic design.
cs.AI updates on arXiv.org