Independent evidence for LLMs + memory

See where models fail.
See when memory helps.

ScoreCrux turns model runs into inspectable results. Compare reasoning, coding, context and memory behaviour at real scale — with the receipts behind every score.

ScoreCrux model and memory benchmark analysis

Published scale results

More context is not automatically better.

Read the full scale study
2M+ token corpus

Full-corpus stuffing reduced core recall on two strong models.

In the published Delta runs, adding the whole corpus made the answer worse than using the model without it.

Claude Sonnet44%28%bare  →  stuffed
GPT-5.428%8%bare  →  stuffed

Results are scoped to the named benchmark, models, corpus and arms. ScoreCrux publishes the methodology and evidence so claims can be checked rather than generalised blindly.

Start with your question

Get to a useful answer quickly.

Benchmark surfaces

Eight ways to interrogate performance.

View all results

A result you can interpret

Effective Minutes anchor performance to useful work.

“23 Em” means 23 quality-adjusted minutes of expert work replaced by the agent. If the safety gate fails, the result is zero.

0 EmUnsafeSafety gate failed
<1 EmLowTrivial or poor quality
1–10 EmRoutineReasonable quality
10–60 EmSignificantExpert work replaced
>60 EmComplexOne hour or more
01

Time-anchored

Minutes of useful expert work, not an opaque score.

02

Safety-gated

A destructive or unsafe session scores zero.

03

Decomposable

Inspect time, information, continuity, safety and cost.

04

Reproducible

Fixtures, versions, receipts and methods remain visible.