DeepSeek Open-Sources DeepSeek-OCR 2 with New Visual Encoding Architecture
Chinese AI startup DeepSeek has officially released and open-sourced its latest optical character recognition model, DeepSeek-OCR 2. This new iteration is built on the proprietary DeepEncoder V2 architecture, which marks a significant shift from traditional rigid scanning-based visual encoding to a semantic reasoning approach. This innovation allows AI systems to dynamically rearrange image components based on context and meaning, aiming for more human-like machine vision capabilities. According to the company, the model significantly enhances data compression efficiency, requiring only 256 to 1,120 visual tokens to process complex document pages, thereby reducing computational costs for downstream large language models. In benchmark tests conducted on OmniDocBench v1.5, DeepSeek-OCR 2 achieved an overall score of 91.09%, representing a 3.73% improvement over its predecessor, with particularly strong performance in reading order recognition. This release underscores the intensifying efforts among Chinese AI developers to enhance foundational models and open-source capabilities amidst growing global competition in the fields of large language models and multimodal AI systems.
Wire timeline
DeepSeek Open-Sources DeepSeek-OCR 2 with New Visual Encoding Architecture
Chinese AI startup DeepSeek has officially released and open-sourced its latest optical character recognition model, DeepSeek-OCR 2. This new iteration is built on the proprietary DeepEncoder V2 architecture, which marks a significant shift from traditional rigid scanning-based visual encoding to a semantic reasoning approach. This innovation allows AI systems to dynamically rearrange image components based on context and meaning, aiming for more human-like machine vision capabilities. According to the company, the model significantly enhances data compression efficiency, requiring only 256 to 1,120 visual tokens to process complex document pages, thereby reducing computational costs for downstream large language models. In benchmark tests conducted on OmniDocBench v1.5, DeepSeek-OCR 2 achieved an overall score of 91.09%, representing a 3.73% improvement over its predecessor, with particularly strong performance in reading order recognition. This release underscores the intensifying efforts among Chinese AI developers to enhance foundational models and open-source capabilities amidst growing global competition in the fields of large language models and multimodal AI systems.
TechNode