Evals
August 14, 2026
plannedLong-horizon performance
A planned evaluation of how models hold up when work stretches across hours, revisions, and changing context—not just a single prompt.
Most capability evaluations ask a model one question and score the answer.
That is a reasonable way to measure some things. It is a poor way to measure the work people actually want models to do: research that unfolds over days, writing that needs revision, plans that collide with reality, and tasks where the important signal is whether the model stays coherent, honest, and useful as context accumulates.
This evaluation is a first sketch of that problem.
We want to measure whether a model can:
- Hold a goal across a long sequence of actions
- Notice when its earlier work is wrong and correct it
- Ask for the information it is missing
- Degrade gracefully instead of hallucinating confidence
The first version will be small, public, and intentionally incomplete. The point is not to declare a winner. The point is to make a dimension of intelligence visible that current leaderboards flatten away.
Details, tasks, and scoring will be published here as the evaluation takes shape.