Apodex launches Apodex 1.1 AI model scoring 44 on AA Index, excelling in agentic tasks over higher-tier peers
Apodex released Apodex 1.1, its first model evaluated on the Artificial Analysis Intelligence Index, scoring 44. It demonstrates strong agentic performance, achieving an Elo of 1348 on GDPval-AA v2 and 70% on TerminalBench v2.1, outperforming models like DeepSeek V4 Pro and Kimi K2.6 in agentic tasks despite a 78.4% hallucination rate and modest knowledge accuracy. The model is available via Apodex's first-party API.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page itself is projected from evidence records.
- Current automated evidence projection
Cross-source coverage
Wire timeline
Apodex 1.1 trains AI agents for complete job takeover with verifiable work outputs
Apodex AI has announced Apodex 1.1, a new AI model trained to complete entire jobs rather than just return answers. The system is designed to handle raw files, datasets, spreadsheets, papers, images, or code, inspect them, choose methods, run code, maintain task state, recover from failures, and produce verifiable outputs. Apodex 1.1 employs two scaling paths: Environment Scaling expands executable file, search, and code situations the model learns from, while Agentic Coordination Scaling trains it to split long tasks across agents, bring partial results into shared state, and revise plans when evidence changes. The AgentOS system keeps files, tool state, artifacts, dependencies, and coordination state persistent across work, allowing subagents to return useful results before all branches finish so the main agent can redirect unfinished work without discarding valid progress. Officially published results place Apodex 1.1 with Agent Team in the leading performance band across professional work, finance, science, reasoning, and search.
Apodex 1.1 scores 44 on AA Index and 70% on Terminal-Bench v2.1
A post on X by DataChaz promotes the AI model Apodex 1.1, claiming it is an 'agentic powerhouse in disguise' despite a low score of 44 on the AA Index. The post highlights that the model achieved a 70% score on Terminal-Bench v2.1, suggesting it excels at real-world execution tasks rather than trivia. The post contrasts the two benchmark results to argue that the AA Index score is misleading regarding the model's practical capabilities. No further details about the model's developer, release date, or methodology are provided in the post.
Apodex released Apodex 1.1, its proprietary model that reached 44 on the Artificial Analysis Intelligence Index and performs strongly on agentic tasks versus models in its tier. on GDPval-AA v2, which measures real-world agentic work, it
Apodex has released Apodex 1.1, its proprietary AI model, which achieved a score of 44 on the Artificial Analysis Intelligence Index. The model demonstrates strong performance on agentic tasks compared to other models in its tier. On the GDPval-AA v2 benchmark, which measures real-world agentic work, Apodex 1.1 reached an Elo rating of 1,348, outperforming several models with higher general intelligence scores. The model's performance appears concentrated around professional and agentic tasks rather than being evenly distributed across evaluation suites. This distinction is highlighted as important for model selection, particularly for agentic execution loops where the key question is how often a model completes a given goal with provided tools.
Show 2 older updatesHide older updates
Apodex 1.1 achieves Elo of 1348 on GDPval-AA v2 benchmark, showing agentic skills
Apodex 1.1, a new version of an AI model, has achieved an Elo rating of 1348 on the GDPval-AA v2 benchmark. This result indicates strong agentic capabilities, meaning the model can effectively perform tasks and make decisions in complex environments. The benchmark score serves as a quantitative measure of the model's performance relative to other systems. The announcement was made via a post on X, highlighting the model's progress in the field of artificial intelligence. The specific details of the model's architecture or training methodology were not disclosed in the post. This development is significant for researchers and developers tracking advancements in AI agent performance.
Apodex has launched Apodex 1.1, a proprietary model scoring 44 on the Artificial Analysis Intelligence Index with strong performance in agentic tasks compared to models in its intelligence tier Apodex 1.1 is Apodex's first model on
Apodex has launched Apodex 1.1, its first model evaluated on the Artificial Analysis Intelligence Index, scoring 44. The model demonstrates strong performance in agentic tasks compared to peers in its intelligence tier, including Kimi K2.6 (45), MiniMax-M3 (45), and Inkling (42). On the GDPval-AA v2 agentic benchmark, Apodex 1.1 achieves an Elo of 1348, ahead of DeepSeek V4 Pro (1333), Qwen3.7 Max (1308), and Kimi K2.6 (1202). On TerminalBench v2.1, it scores 70%, ahead of Kimi K2.6 (66%) and behind Qwen3.7 Max (75%). The model uses approximately 17,000 output tokens per task, making it less token-efficient than some peers but cheaper per task at about $0.05, compared to Qwen3.7 Max ($0.07) and Kimi K2.6 ($0.06). However, it shows modest knowledge accuracy with a -21.9 score on AA-Omniscience, 32% accuracy, and a 78.4% hallucination rate. The model has a 256K token context window, pricing of $0.30/$3.00 per 1M input/output tokens, and is available via Apodex's first-party API.