ZAYA1-VL-8B Technical Report Released on arXiv
Researchers have published a technical report for ZAYA1-VL-8B, a compact mixture-of-experts vision-language model built upon the in-house ZAYA1-8B language model. Despite its small footprint, the model demonstrates performance competitive with leading base models like Molmo2-4B and InternVL3.5-4B, while surpassing others such as Qwen2.5-VL-3B in image understanding, reasoning, and counting benchmarks. The architecture features two key innovations: vision-specific LoRA adapters integrated into the LLM to boost modality-specific capacity without adding experts, and bidirectional attention over image tokens to enhance visual comprehension. The report details the complete training pipeline, including data composition, sequence packing, and attention masking schemes. The model comprises 9.2 billion total parameters, with only 1.4 billion active parameters including the vision encoder. Developed by a team including Hassan Shapourian and colleagues, ZAYA1-VL-8B is now publicly available on Hugging Face, offering an efficient solution for advanced multimodal AI tasks.
Wire timeline
ZAYA1-VL-8B Technical Report Released on arXiv
Researchers have published a technical report for ZAYA1-VL-8B, a compact mixture-of-experts vision-language model built upon the in-house ZAYA1-8B language model. Despite its small footprint, the model demonstrates performance competitive with leading base models like Molmo2-4B and InternVL3.5-4B, while surpassing others such as Qwen2.5-VL-3B in image understanding, reasoning, and counting benchmarks. The architecture features two key innovations: vision-specific LoRA adapters integrated into the LLM to boost modality-specific capacity without adding experts, and bidirectional attention over image tokens to enhance visual comprehension. The report details the complete training pipeline, including data composition, sequence packing, and attention masking schemes. The model comprises 9.2 billion total parameters, with only 1.4 billion active parameters including the vision encoder. Developed by a team including Hassan Shapourian and colleagues, ZAYA1-VL-8B is now publicly available on Hugging Face, offering an efficient solution for advanced multimodal AI tasks.
cs.AI updates on arXiv.org