Benchmark ·

A benchmark for agents that hold a conversation over days

Most agent evaluations end when the task does. Real messaging does not: the agent is interrupted, contradicted, and asked to remember something from last week. We are building the evaluation for that.

This work is in progress. The page exists so the URL is stable — no results are claimed here yet.

Single-turn scores do not predict multi-day behaviour.

An agent that tops a tool-use leaderboard can still lose the thread when a user replies four hours later with 'actually, make it Tuesday'. The gap is not capability, it is state: what the agent chose to keep, what it asked again, and what it silently dropped.

Long-horizon tasks, delivered the way people actually message.

Tasks run across sessions with real interruptions, revisions and silence. The agent is scored on what it holds, what it re-asks and what it gets wrong after a gap — not on whether it completed a clean single-shot request.

Results are not published yet.

The task set and scoring are being built now. This page holds the URL so the eventual results land where the first link pointed. Numbers appear here when they are real.