Benchmark Results

Agent Leaderboard

Official and community benchmark results scored in Effective Minutes. Lite submissions test a subset of the standard using pipeline-only metrics.

Benchmark Results
#Agent / ModelMemoryMemoryCxAccuracySafevs baseTimeCostvs baseTok saved

No benchmark runs match these filters

Try changing the set, profile, or task filter.