Society
NIST launches a sequestered testbed so AI benchmarks stop grading their own homework
8:00 AM ET · July 31, 2026
NIST’s Technology Test and Evaluation Division announced AITE, a new program built around a sequestered testbed environment where AI models can be evaluated against data and tasks they have never encountered during training, directly addressing the growing concern that publicly available benchmarks become progressively less meaningful once their answers have likely leaked into training data somewhere on the internet. Under the program, data providers submit original datasets and meaningful, real world tasks and receive performance measurements on their own data in return, while model providers submit their AI systems for blind testing and get comparative results against other participating models, all without either side seeing the other’s underlying data or model weights directly. The first three tasks focus on image analysis using large vision language models across three specific domains: quantum science, genomics, and public safety, chosen as areas where genuinely novel test data is available and where model performance has real world stakes. NIST said additional tasks will be added over time as more data providers and model developers join the program, though it did not specify a start date for the initial round of evaluations. The effort sits alongside NIST’s other recent AI work, including a July workshop on securing AI data center infrastructure, as part of a broader push by the agency to build measurement and evaluation infrastructure that keeps pace with how quickly AI capabilities, and the benchmarks meant to track them, are moving.