● long-horizon agent tasks · signed

Agent leaderboard

A single model can write. Can it build a reconciled model, run multi-hop research, and catch its own errors? AtlasVector's agent pipeline (orchestration + reconciliation + self-falsification) vs a raw frontier-agent baseline on the same long-horizon tasks — sha256-chained and independently re-derivable.

Task success rate
100.0%
vs baseline 8.3%
Avg quality score
100.0%
vs baseline 52.1%
Self-falsification catch-rate
100.0%
vs baseline 25.0%
Faithfulness
100.0%
vs baseline 52.1%
atlas-agent-pipeline vs frontier-llm-agent (synthetic baseline)
AtlasVector won 11 of 12 tasks (91.7% win rate), baseline 0.
By task type
TasknAtlasVectorBaselineΔ ptsW–L
Build a reconciled model4100.0%0.0%+100.04–0
multi-hop4100.0%0.0%+100.04–0
Self-falsification catch4100.0%25.0%+75.03–0
By market regime
RegimenAtlasVectorBaselineΔ ptsW–L
CREDIT2100.0%0.0%+100.02–0
MACRO2100.0%0.0%+100.02–0
RISK2100.0%0.0%+100.02–0
VOL1100.0%0.0%+100.01–0
WEALTH1100.0%0.0%+100.01–0
EQUITY4100.0%25.0%+75.03–0

Tamper-evident · independently re-derivable

chain tip (sha256)4bde314488356334ca0363b7de8a3d7213c3adeb26083a997b218000b3069b8a
versionagentbench-v1
Atlas agent pipeline (orchestration + reconciliation + self-falsification) vs a synthetic frontier-agent baseline on the SAME long-horizon agent tasks: build-a-reconciled-model, multi-hop research, and self-falsification catch-rate. Baseline is a deterministic synthetic stand-in (no live frontier key) — every cell is synthetic and the whole leaderboard is sha256-chained + independently re-derivable at /verify.