● tamper-evident, re-derivable benchmark

AtlasVector vs a frontier model

AtlasVector's full pipeline — multi-desk debate, the number-ledger re-derivation, and the grounding gate — against a raw frontier-LLM baseline on the same finance tasks. Regime & reason bucketed, sha256-chained, and independently re-derivable. We publish the hash; you can check it.

atlas-full-pipeline
86.7%
accuracy · re-derived + grounded
vs+6.7 pts
frontier-llm (synthetic baseline)
80.0%
raw single-model
15 tasksAtlasVector wins 13–2win rate 73.3%
The difference that matters — grounding

Accuracy is close (86.7% vs 80.0%). What separates AtlasVector is provenance: it grounds its claims in cited sources and re-derives every number, so you can check the work. A raw single-model baseline does neither — it emits prose, not sourced claims, so it carries no citation layer to ground or cover.

Claims grounded in cited sourcesfaithfulness
AtlasVector86.7%
single-model baseline0.0%
Sources it should cite, citedcoverage
AtlasVector86.7%
single-model baseline0.0%

Baseline shown is a deterministic, illustrative stand-in. The grounding gap is structural to any single-model baseline — it has no citation layer to score — not an empirical horse-race result.

By market regime
RegimenAtlasVectorBaselineΔ ptsW–L
CREDIT2100.0%50.0%+50.02–0
EQUITY7100.0%85.7%+14.37–0
RISK2100.0%100.0%+0.02–0
WEALTH1100.0%100.0%+0.01–0
MACRO250.0%50.0%+0.01–1
VOL10.0%100.0%-100.00–1
By reasoning type
ReasonnAtlasVectorBaselineΔ ptsW–L
capital-return1100.0%100.0%+0.01–0
cashflow1100.0%100.0%+0.01–0
leverage1100.0%0.0%+100.01–0
liquidity1100.0%100.0%+0.01–0
profitability2100.0%100.0%+0.02–0
quality1100.0%0.0%+100.01–0
risk2100.0%100.0%+0.02–0
strategy1100.0%100.0%+0.01–0
valuation2100.0%100.0%+0.02–0
rates250.0%50.0%+0.01–1
options10.0%100.0%-100.00–1

Tamper-evident · independently re-derivable

chain tip (sha256)abe060c7fa1515bfc4f18719b0cf5fcf8f5b67ba9370d6e25a26176f6fef1ba0
versionpubbench-v1 · golden finbench-v1
Atlas full pipeline (debate + number-ledger + grounding gate) vs a raw frontier-LLM baseline on the same finbench tasks. Baseline is a deterministic synthetic stand-in (no live frontier key) — every cell is synthetic and the whole leaderboard is sha256-chained + independently re-derivable at /verify.