Vector Institute Introduces HumaniBench for Human-Centered Evaluation of Multimodal AI Models
The Vector Institute for Artificial Intelligence has introduced HumaniBench, the first comprehensive benchmark designed to evaluate large multimodal models (LMMs) through a human-centered lens. Moving beyond traditional metrics like accuracy and speed, HumaniBench assesses AI alignment with seven key human values: fairness, ethics, understanding, reasoning, language inclusivity, empathy, and robustness. The benchmark utilizes a curated dataset of 32,000 image-question pairs derived from real-world news articles, annotated via a GPT-4o workflow and expert verification. It features seven diverse tasks, including scene understanding, empathetic captioning, and visual grounding. In initial tests involving 15 leading LMMs, such as GPT-4o, Gemini 2.0, Llama 3.2, and Qwen, results indicated that while proprietary models excelled in reasoning and empathy, open-source models performed better in visual grounding and robustness. Crucially, nearly all models exhibited discrepancies in handling demographic attributes like age, race, and language. This initiative highlights significant gaps in current AI systems regarding ethical judgment and inclusivity, offering a structured framework for developers to improve responsible and equitable AI development.
Wire timeline
Vector Institute Introduces HumaniBench for Human-Centered Evaluation of Multimodal AI Models
The Vector Institute for Artificial Intelligence has introduced HumaniBench, the first comprehensive benchmark designed to evaluate large multimodal models (LMMs) through a human-centered lens. Moving beyond traditional metrics like accuracy and speed, HumaniBench assesses AI alignment with seven key human values: fairness, ethics, understanding, reasoning, language inclusivity, empathy, and robustness. The benchmark utilizes a curated dataset of 32,000 image-question pairs derived from real-world news articles, annotated via a GPT-4o workflow and expert verification. It features seven diverse tasks, including scene understanding, empathetic captioning, and visual grounding. In initial tests involving 15 leading LMMs, such as GPT-4o, Gemini 2.0, Llama 3.2, and Qwen, results indicated that while proprietary models excelled in reasoning and empathy, open-source models performed better in visual grounding and robustness. Crucially, nearly all models exhibited discrepancies in handling demographic attributes like age, race, and language. This initiative highlights significant gaps in current AI systems regarding ethical judgment and inclusivity, offering a structured framework for developers to improve responsible and equitable AI development.
Vector Institute for Artificial Intelligence