Research

Life After Benchmark Saturation: A Case Study of CORE-Bench

June 23, 2026

We show that when a benchmark's accuracy saturates, six other dimensions of agent performance remain informative — construct validity, out-of-distribution generalizability, efficiency, reliability, model versus scaffold contributions, and uplift f...

Log Analysis is Necessary for Credible Evaluation of AI Agents

May 08, 2026

We argue that log analysis — the systematic tracking and analysis of the inputs, execution, and outputs of an AI agent — is necessary to overcome validity threats in agent evaluation, presenting a taxonomy of threats and a set of guiding principles.

Towards a Science of AI Agent Reliability

February 21, 2026

We propose twelve metrics decomposing AI agent reliability along four dimensions — consistency, robustness, predictability, and safety — and evaluate 14 models, finding that recent capability gains have yielded only small improvements in reliability.