Research
Benchmarks for dimensions of intelligence that traditional evaluations miss.
August 14, 2026
A planned evaluation of how models hold up when work stretches across hours, revisions, and changing context—not just a single prompt.