Vector Institute Releases Independent Evaluation of Leading AI Models
Canada’s Vector Institute has unveiled the results of its independent evaluation of eleven leading large language models (LLMs), providing an objective assessment of their performance. The study, titled State of Evaluation, tested both open-source models like DeepSeek-R1 and closed-source commercial models such as OpenAI’s GPT-4o and Google’s Gemini 1.5 against sixteen comprehensive benchmarks. These benchmarks cover critical areas including general knowledge, coding capabilities, and cyber-safety. In a significant move for transparency, Vector Institute has open-sourced the benchmarks, underlying code, and results via an interactive leaderboard. This initiative aims to foster accountability, collaboration, and trust in artificial intelligence by allowing researchers, developers, and policymakers to independently verify results and understand model strengths and limitations. The project builds on Vector’s previous contributions to AI safety standards, including the development of MMLU-Pro and the Inspect Evals platform. By offering these resources, the institute seeks to help organizations deploy AI systems safely and responsibly, addressing the rapid evolution of frontier AI technologies while promoting rigorous, standardized evaluation methods across the global AI community.
Wire timeline
Vector Institute Releases Independent Evaluation of Leading AI Models
Canada’s Vector Institute has unveiled the results of its independent evaluation of eleven leading large language models (LLMs), providing an objective assessment of their performance. The study, titled State of Evaluation, tested both open-source models like DeepSeek-R1 and closed-source commercial models such as OpenAI’s GPT-4o and Google’s Gemini 1.5 against sixteen comprehensive benchmarks. These benchmarks cover critical areas including general knowledge, coding capabilities, and cyber-safety. In a significant move for transparency, Vector Institute has open-sourced the benchmarks, underlying code, and results via an interactive leaderboard. This initiative aims to foster accountability, collaboration, and trust in artificial intelligence by allowing researchers, developers, and policymakers to independently verify results and understand model strengths and limitations. The project builds on Vector’s previous contributions to AI safety standards, including the development of MMLU-Pro and the Inspect Evals platform. By offering these resources, the institute seeks to help organizations deploy AI systems safely and responsibly, addressing the rapid evolution of frontier AI technologies while promoting rigorous, standardized evaluation methods across the global AI community.
Vector Institute for Artificial Intelligence