● long-horizon agent tasks · signed
Agent leaderboard
A single model can write. Can it build a reconciled model, run multi-hop research, and catch its own errors? AtlasVector's agent pipeline (orchestration + reconciliation + self-falsification) vs a raw frontier-agent baseline on the same long-horizon tasks — sha256-chained and independently re-derivable.
Task success rate
100.0%
vs baseline 8.3%
Avg quality score
100.0%
vs baseline 52.1%
Self-falsification catch-rate
100.0%
vs baseline 25.0%
Faithfulness
100.0%
vs baseline 52.1%
atlas-agent-pipeline vs frontier-llm-agent (synthetic baseline)
AtlasVector won 11 of 12 tasks (91.7% win rate), baseline 0.
By task type
| Task | n | AtlasVector | Baseline | Δ pts | W–L |
|---|---|---|---|---|---|
| Build a reconciled model | 4 | 100.0% | 0.0% | +100.0 | 4–0 |
| multi-hop | 4 | 100.0% | 0.0% | +100.0 | 4–0 |
| Self-falsification catch | 4 | 100.0% | 25.0% | +75.0 | 3–0 |
By market regime
| Regime | n | AtlasVector | Baseline | Δ pts | W–L |
|---|---|---|---|---|---|
| CREDIT | 2 | 100.0% | 0.0% | +100.0 | 2–0 |
| MACRO | 2 | 100.0% | 0.0% | +100.0 | 2–0 |
| RISK | 2 | 100.0% | 0.0% | +100.0 | 2–0 |
| VOL | 1 | 100.0% | 0.0% | +100.0 | 1–0 |
| WEALTH | 1 | 100.0% | 0.0% | +100.0 | 1–0 |
| EQUITY | 4 | 100.0% | 25.0% | +75.0 | 3–0 |
Tamper-evident · independently re-derivable
chain tip (sha256)4bde314488356334ca0363b7de8a3d7213c3adeb26083a997b218000b3069b8a
versionagentbench-v1
Atlas agent pipeline (orchestration + reconciliation + self-falsification) vs a synthetic frontier-agent baseline on the SAME long-horizon agent tasks: build-a-reconciled-model, multi-hop research, and self-falsification catch-rate. Baseline is a deterministic synthetic stand-in (no live frontier key) — every cell is synthetic and the whole leaderboard is sha256-chained + independently re-derivable at /verify.