The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark
A new research paper introduces KnotBench, a rigorous benchmark designed to evaluate the diagrammatic reasoning capabilities of Vision-Language Models (VLMs). The study utilizes a corpus of over 858,000 images derived from 1,951 prime-knot prototypes to test models on tasks such as equivalence judgment, move prediction, and identification. Results indicate that leading models, including Claude Opus 4.7 and GPT-5, struggle significantly with structural reasoning despite accurate visual perception. Most tasks yielded scores near or below random baselines, and no model successfully produced strictly correct symbol transcriptions. While enabling 'thinking-mode' reasoning improved accuracy modestly, particularly for GPT-5, it failed to bridge the gap between perception and operational simulation. The findings suggest that current VLMs can identify visual features but lack the apparatus to simulate logical moves on those features, highlighting a critical limitation in AI's ability to handle complex spatial and topological reasoning.
Wire timeline
The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark
A new research paper introduces KnotBench, a rigorous benchmark designed to evaluate the diagrammatic reasoning capabilities of Vision-Language Models (VLMs). The study utilizes a corpus of over 858,000 images derived from 1,951 prime-knot prototypes to test models on tasks such as equivalence judgment, move prediction, and identification. Results indicate that leading models, including Claude Opus 4.7 and GPT-5, struggle significantly with structural reasoning despite accurate visual perception. Most tasks yielded scores near or below random baselines, and no model successfully produced strictly correct symbol transcriptions. While enabling 'thinking-mode' reasoning improved accuracy modestly, particularly for GPT-5, it failed to bridge the gap between perception and operational simulation. The findings suggest that current VLMs can identify visual features but lack the apparatus to simulate logical moves on those features, highlighting a critical limitation in AI's ability to handle complex spatial and topological reasoning.
cs.AI updates on arXiv.org