Board 1 of 3

Capability

How far a rig gets. Every row is a full configuration — which model, backed by which memory system, at which reasoning effort — ranked in Effective Minutes: quality-adjusted minutes of expert work replaced. Em is unbounded by construction, so harder fixtures raise the ceiling instead of saturating it.

Capability

0 ranked

#ModelMemoryEffortEmEmCostCtx tokRunsCorpus
No rows can be ranked on this view with the current filters.

How this is scored

Unit
Effective Minutes (Em) — quality-adjusted minutes of expert work replaced.
Lift
Em(rig) minus Em(same model, same effort, vendor-native) within the same surface and corpus. Null when no matched baseline exists — never zero. A positive lift on "None" means the bare model beat having the whole corpus in context, which is a real result at large corpus sizes, not an error.
Anchors
T_human anchors are PROVISIONAL expert estimates, not panel-measured.