Don't trust us. Verify us.
Every public claim AtlasVector makes seals to one tamper-evident audit chain — the head-to-head benchmark, the long-horizon agent leaderboard, and the self-falsified house verdicts all re-derive from a single sha256 root. Re-walk the chain yourself.
The Glass Box
AtlasVector vs a frontier model
Agent-task leaderboard
The reliability diagram — predicted vs realized hit probability. We bin every served prediction by the probability the model assigned it, then plot that bin's mean predicted probability against how often the call actually landed. A well-calibrated model's predicted probabilities match realized frequencies, so its dots sit on the diagonal. Overconfidence shows as dots below the line; underconfidence, above it.
The VaR backtest — Basel's traffic-light test applied to our own engine. Every day the risk desk makes a 99% one-day Value-at-Risk forecast; the next day the book realizes a P&L. An exception is a realized loss worse than the forecast. Over a rolling 250-day window you expect ~2.5 exceptions: ≤4 green, 5–9 yellow, ≥10 red. Kupiec (coverage) and Christoffersen (independence) tests back the light. This is glass-box applied to our own risk numbers — we grade ourselves in public.
One root, not three claims
Benchmark, leaderboard, and verdicts share one chain. You can't fabricate a track record retroactively without breaking every downstream hash.
It attacks itself
The agent runs a RED-TEAM desk and a self-falsification gate on its own conclusions, and we publish the catch-rate.
Independently verifiable by any party
Every proof surface ships a /verify endpoint. The numbers aren't "trust the vendor" — they're recompute-and-check.
Data provenance
Market data is real-time for stocks, ETFs, indices and crypto; illustrative figures are labeled as such. The reasoning structure and the cryptographic seal are always real.