Vector Institute Benchmarks xAI’s Open-Source Grok-1 Model
The Vector Institute for Artificial Intelligence has conducted a comprehensive benchmark of xAI’s newly open-sourced Grok-1 model. Released on March 17, 2024, Grok-1 is a massive 314-billion parameter Mixture of Experts (MoE) model, making it the largest publicly available language model to date. Despite its size, which is significantly larger than Llama-2-70B and Mixtral 8x7B, the Vector Institute’s analysis reveals underwhelming performance results. Testing on subsets of the MMLU dataset showed that Grok-1 falls short of current state-of-the-art open-source models. Furthermore, safety evaluations using the RealToxicityPrompts dataset indicated that Grok-1, which lacks fine-tuning or alignment training like RLHF, exhibits significantly higher toxicity levels compared to closed-source alternatives. The model requires substantial hardware resources, specifically eight A100 GPUs, and currently suffers from slow inference speeds due to unoptimized implementation. Released under the Apache 2.0 license, Grok-1 allows commercial use but presents challenges for practical deployment without further community-driven optimization and safety alignment.
Wire timeline
Vector Institute Benchmarks xAI’s Open-Source Grok-1 Model
The Vector Institute for Artificial Intelligence has conducted a comprehensive benchmark of xAI’s newly open-sourced Grok-1 model. Released on March 17, 2024, Grok-1 is a massive 314-billion parameter Mixture of Experts (MoE) model, making it the largest publicly available language model to date. Despite its size, which is significantly larger than Llama-2-70B and Mixtral 8x7B, the Vector Institute’s analysis reveals underwhelming performance results. Testing on subsets of the MMLU dataset showed that Grok-1 falls short of current state-of-the-art open-source models. Furthermore, safety evaluations using the RealToxicityPrompts dataset indicated that Grok-1, which lacks fine-tuning or alignment training like RLHF, exhibits significantly higher toxicity levels compared to closed-source alternatives. The model requires substantial hardware resources, specifically eight A100 GPUs, and currently suffers from slow inference speeds due to unoptimized implementation. Released under the Apache 2.0 license, Grok-1 allows commercial use but presents challenges for practical deployment without further community-driven optimization and safety alignment.
Vector Institute for Artificial Intelligence