AtlasVector vs a frontier model
AtlasVector's full pipeline — multi-desk debate, the number-ledger re-derivation, and the grounding gate — against a raw frontier-LLM baseline on the same finance tasks. Regime & reason bucketed, sha256-chained, and independently re-derivable. We publish the hash; you can check it.
Accuracy is close (86.7% vs 80.0%). What separates AtlasVector is provenance: it grounds its claims in cited sources and re-derives every number, so you can check the work. A raw single-model baseline does neither — it emits prose, not sourced claims, so it carries no citation layer to ground or cover.
Baseline shown is a deterministic, illustrative stand-in. The grounding gap is structural to any single-model baseline — it has no citation layer to score — not an empirical horse-race result.
| Regime | n | AtlasVector | Baseline | Δ pts | W–L |
|---|---|---|---|---|---|
| CREDIT | 2 | 100.0% | 50.0% | +50.0 | 2–0 |
| EQUITY | 7 | 100.0% | 85.7% | +14.3 | 7–0 |
| RISK | 2 | 100.0% | 100.0% | +0.0 | 2–0 |
| WEALTH | 1 | 100.0% | 100.0% | +0.0 | 1–0 |
| MACRO | 2 | 50.0% | 50.0% | +0.0 | 1–1 |
| VOL | 1 | 0.0% | 100.0% | -100.0 | 0–1 |
| Reason | n | AtlasVector | Baseline | Δ pts | W–L |
|---|---|---|---|---|---|
| capital-return | 1 | 100.0% | 100.0% | +0.0 | 1–0 |
| cashflow | 1 | 100.0% | 100.0% | +0.0 | 1–0 |
| leverage | 1 | 100.0% | 0.0% | +100.0 | 1–0 |
| liquidity | 1 | 100.0% | 100.0% | +0.0 | 1–0 |
| profitability | 2 | 100.0% | 100.0% | +0.0 | 2–0 |
| quality | 1 | 100.0% | 0.0% | +100.0 | 1–0 |
| risk | 2 | 100.0% | 100.0% | +0.0 | 2–0 |
| strategy | 1 | 100.0% | 100.0% | +0.0 | 1–0 |
| valuation | 2 | 100.0% | 100.0% | +0.0 | 2–0 |
| rates | 2 | 50.0% | 50.0% | +0.0 | 1–1 |
| options | 1 | 0.0% | 100.0% | -100.0 | 0–1 |