Wire flash
TechNIST launches AITE platform for safe AI model evaluation
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
The National Institute of Standards and Technology (NIST) has launched the AI Technology Evaluation (AITE) platform, a voluntary testing program that provides researchers with an isolated testbed environment to safely evaluate AI models. AITE offers blind, non-training data for objective assessments of model capabilities, initially focusing on image analysis tasks in quantum science, genomics, and public safety using large vision language models. The platform aims to establish a universal rubric for evaluating AI performance and determining the state of the art. Data providers must submit original, non-public datasets with meaningful tasks, while model developers submit their models for testing. The first evaluations begin in August 2026. This initiative is part of the Trump administration's broader strategy to advance AI safety through voluntary model submissions, following a renegotiated deal with Google DeepMind, Microsoft, and xAI under the Commerce Department's Center for AI Standards and Innovation.
Source report
The National Institute of Standards and Technology (NIST) launched a new program on Monday, granting researchers access to an isolated testbed environment to safely evaluate artificial intelligence models against various commands.
About AITE
The AI Technology Evaluation, or AITE, is a voluntary testing vehicle focused on AI model safety analysis. It provides blind data for models to process when completing tasks, enabling objective insights and evaluations of model capabilities. Notably, the evaluation data is not intended to serve as training data for the models.
Initial Focus Areas
Initially, AITE will focus on conducting image analysis tasks using large vision language models across three domains:
- Quantum science
- Genomics
- Public safety
Additional tasks will be made available in the future.
Core Objective
AITE’s fundamental goal is to offer a universal rubric to effectively evaluate AI models' capabilities and determine the state of the art for model performance.
“The infrastructure provided by NIST will provide common data, metrics, and scoring to help developers understand the performance of their models,” the press release said.
Participation Requirements
Both data providers and model providers working with AITE will need to submit materials related to the testing:
- Data providers are asked to submit original datasets that are inaccessible publicly, along with a “meaningful” task suited for the data.
- Model developers will submit their AI models to be tested using the datasets.
The first set of evaluations will begin in August 2026.
Broader Context
The formation of AITE is the latest step in the Trump administration’s strategy to work with major AI developers in advancing model safety through voluntary model submissions.
The Commerce Department announced a renegotiated deal in May between the agency and three companies—Google DeepMind, Microsoft, and xAI—to evaluate their models through the Center for AI Standards and Innovation.
Source
Defense One - All ContentWestern
Part of this Story
NIST launches AITE platform for safe AI model evaluation