Vector Institute Releases Open-Source State of Evaluation Study for Global AI Models
The Vector Institute has published its inaugural State of Evaluation study, assessing eleven leading global artificial intelligence models using sixteen diverse benchmarks. This initiative marks a significant step toward transparency in the AI sector, as Vector is open-sourcing both the evaluation results and the underlying code. The study examines a mix of open-source and closed-source frontier models, including recent releases like DeepSeek-R1, GPT-4o, and Llama-3.1. Developed by Vector’s AI Engineering team, the assessment utilizes benchmarks pioneered by researchers Wenhu Chen and Victor Zhong, covering both single-turn knowledge tasks and complex agentic scenarios that simulate real-world decision-making. By providing an interactive leaderboard and reproducible methods, the institute aims to help developers, policymakers, and users better understand model accuracy, reliability, and safety. The project also involves collaboration with Google DeepMind to validate dangerous capabilities benchmarks. This effort underscores Vector’s role in establishing standardized, accountable evaluation frameworks for large language models, facilitating responsible deployment amidst rapid technological advancements.
Wire timeline
Vector Institute Releases Open-Source State of Evaluation Study for Global AI Models
The Vector Institute has published its inaugural State of Evaluation study, assessing eleven leading global artificial intelligence models using sixteen diverse benchmarks. This initiative marks a significant step toward transparency in the AI sector, as Vector is open-sourcing both the evaluation results and the underlying code. The study examines a mix of open-source and closed-source frontier models, including recent releases like DeepSeek-R1, GPT-4o, and Llama-3.1. Developed by Vector’s AI Engineering team, the assessment utilizes benchmarks pioneered by researchers Wenhu Chen and Victor Zhong, covering both single-turn knowledge tasks and complex agentic scenarios that simulate real-world decision-making. By providing an interactive leaderboard and reproducible methods, the institute aims to help developers, policymakers, and users better understand model accuracy, reliability, and safety. The project also involves collaboration with Google DeepMind to validate dangerous capabilities benchmarks. This effort underscores Vector’s role in establishing standardized, accountable evaluation frameworks for large language models, facilitating responsible deployment amidst rapid technological advancements.
Vector Institute for Artificial Intelligence