Wire flash
TechApodex launches Apodex 1.1 model with score of 44 on AI Index
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
Apodex has launched Apodex 1.1, its first model evaluated on the Artificial Analysis Intelligence Index, scoring 44. The model demonstrates strong performance in agentic tasks compared to peers in its intelligence tier, including Kimi K2.6 (45), MiniMax-M3 (45), and Inkling (42). On the GDPval-AA v2 agentic benchmark, Apodex 1.1 achieves an Elo of 1348, ahead of DeepSeek V4 Pro (1333), Qwen3.7 Max (1308), and Kimi K2.6 (1202). On TerminalBench v2.1, it scores 70%, ahead of Kimi K2.6 (66%) and behind Qwen3.7 Max (75%). The model uses approximately 17,000 output tokens per task, making it less token-efficient than some peers but cheaper per task at about $0.05, compared to Qwen3.7 Max ($0.07) and Kimi K2.6 ($0.06). However, it shows modest knowledge accuracy with a -21.9 score on AA-Omniscience, 32% accuracy, and a 78.4% hallucination rate. The model has a 256K token context window, pricing of $0.30/$3.00 per 1M input/output tokens, and is available via Apodex's first-party API.
Source report
Apodex has released Apodex 1.1, a proprietary model that scored 44 on the Artificial Analysis Intelligence Index, demonstrating strong performance in agentic tasks compared to other models in its intelligence tier.
Model Positioning
Apodex 1.1 is the lab's first model evaluated on Artificial Analysis. The company previously launched Apodex 1.0 and Apodex 1.0 mini. With an Intelligence Index score of 44, Apodex 1.1 sits in a similar tier alongside:
- Kimi K2.6 (45)
- MiniMax-M3 (45)
- Inkling (42)
Within this group, Apodex 1.1 stands out on agentic and knowledge work evaluations, but shows trade-offs in knowledge reliability and frontier academic reasoning.
Key Results
Agentic Workflows
- On GDPval-AA v2 (real-world agentic work benchmark), Apodex 1.1 achieves an Elo of 1348, ahead of:
- DeepSeek V4 Pro (1333)
- Qwen3.7 Max (1308)
- Kimi K2.6 (1202)
- On TerminalBench v2.1 (agentic coding and terminal use benchmark), it scores 70%, ahead of Kimi K2.6 (66%) and just behind Qwen3.7 Max (75%)
Token Efficiency
- Apodex 1.1 uses approximately 17,000 output tokens per task on average across the Artificial Analysis Intelligence Index
- More token efficient than DeepSeek V4 Pro (16,842 tokens)
- Less token efficient than peers with similar Intelligence Index scores:
- Qwen3.7 Max (9,391 tokens)
- MiniMax-M3 (8,133 tokens)
Cost Efficiency
- Costs approximately $0.05 per Intelligence Index Task, making it relatively attractive compared to peer models in its intelligence tier
- This is primarily driven by cheap pricing, slightly offset by higher tokens per task
- At ~$0.05, Apodex 1.1 is cheaper than:
- Qwen3.7 Max (~$0.07)
- Kimi K2.6 (~$0.06)
- While not on the Cost vs. Intelligence Pareto frontier, it lands in the most attractive quadrant
Knowledge Accuracy and Reliability
- Scores -21.9 on AA-Omniscience
- Achieves 32% accuracy on individual questions
- 78.4% hallucination rate
- Attempts to answer 87% of questions rather than declining to respond
Additional Model Details
| Feature | Specification | |---|---| | Context window | 256K tokens | | Pricing | $0.30 / $3.00 per 1M input/output tokens | | Cache-hit price | $0.03 | | Availability | Apodex first party API |
Source
ArtificialAnlysNeutral / independent
Part of this Story
Apodex launches Apodex 1.1 AI model scoring 44 on AA Index, excelling in agentic tasks over higher-tier peers