Alibaba Qwen open-sources Qwen-Image-2.1, a 7B unified image generation and editing model
Alibaba's Qwen team open-sourced Qwen-Image-2.1, a lightweight 7-billion-parameter model that unifies text-to-image generation and image editing in a single architecture. It natively supports transparent RGBA images, up to 10 reference images for localized editing, and tasks like panoramas and infographics. The model is available with open weights on GitHub, ModelScope, and Hugging Face, with a live demo on Hugging Face Spaces and integration into SGLang-Diffusion for serving.
IllustrationEditorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Reporting timeline
SGLang-Diffusion deploys Qwen-Image-2.1 for text-to-image and editing tasks
In a post on X, Alibaba Qwen announced that SGLang-Diffusion now serves Qwen-Image-2.1, enabling text-to-image generation, multi-image editing, and transparent RGBA output. The post thanks the SGL project for day-0 support and invites users to try the new capabilities. This marks an integration of the Qwen-Image-2.1 model into the SGLang-Diffusion serving framework, expanding its functionality for AI-powered image creation and manipulation.
Read sourceQwen-Image-2.1 × @HuggingApps: live demo on Spaces! 🖼 One single checkpoint for generation and editing. Try it in your browser, no setup needed. 👇
Alibaba's Qwen team, in collaboration with HuggingApps, has released a live demo of Qwen-Image-2.1 on Hugging Face Spaces. The model is a single checkpoint capable of both image generation and editing. Users can try the demo directly in their browser without any setup required. This release provides public access to the latest iteration of Qwen's image-focused AI model, enabling hands-on testing of its generation and editing capabilities.
Read sourceQwen Open-Sources Qwen-Image-2.1, a Lightweight High-Performance Image Model
On September 20, according to an announcement from Qwen Large Model, the company open-sourced Qwen-Image-2.1, an image model designed to balance generation quality, inference efficiency, and usage costs. The model integrates text-to-image generation and image editing within a single architecture, with the visual generation component containing only 7 billion parameters. It natively supports generating and editing transparent images. Key features include: a lightweight architecture that balances quality and computational cost; native transparency support for generating standard or transparent images and editing transparent layers; comprehensive editing capabilities with support for up to 10 reference images to enhance local editing and portrait/product fidelity; and improved text layout, character lighting and shadows, and detail rendering for more aesthetically pleasing results.
Show 2 older updatesHide older updates
Alibaba Qwen releases Qwen-Image-2.1 with open weights and native transparency
Alibaba's Qwen team announced the release of Qwen-Image-2.1, described as the most balanced and cost-effective image generation model in the Qwen-Image series. The model is now available with open weights. It is a unified model for both image generation and editing, delivering top-tier quality in a lightweight 7B architecture that outperforms most closed-source models. Key highlights include compact and exceptionally fast inference for multi-image inputs, native support for generating and editing RGBA layers for seamless compositing and text editing, versatile high-fidelity editing with support for up to 10 reference images, and broad coverage of tasks like panoramas, infographics, and virtual try-ons with realistic textures and typography. Links to the blog, GitHub, Model Scope, and Hugging Face repositories were provided.
Read sourceQwen Open-Sources Qwen-Image-2.1: 7B Model for Unified Generation and Editing
The Qwen team has open-sourced Qwen-Image-2.1, a unified model that integrates text-to-image generation and image editing into a single system. The visual generation component features only 7 billion parameters and natively supports the generation and editing of transparent images. The model accommodates up to 10 reference images and enables localized editing via circular, scribble, or independent masks. It improves inference efficiency through a mixed-granularity attention architecture and KV cache reuse, while enhancing text rendering, portrait lighting, and fidelity for people and products. Additionally, it covers tasks such as panoramic images, infographics, and storyboards. Official documentation details the model's capabilities, allowing readers to assess its usability in scenarios like design.
Read source