Interactive Evals Are Killing the Benchmark
ARC-AGI-3 launched in March 2026. Humans score 100%. Frontier AI scores under 1%. The 99-point gap is the smaller story — the bigger one is why static reasoning benchmarks have stopped predicting whether a production agent will actually work.