OQOmniQuantBeta🛢️ COMMODITIES
Live
Cohort
Window
All markets · last 90 days · all horizons (default)
🌐 Every market · track record

The whole track record.
Every slice, including misses.

Every prediction — across all markets — scored on the real close-to-close move, broken down by horizon, confidence, sector, region, model and regime. Calibration, risk-adjusted returns and platform discipline included. No cherry-picking.

95% confidence interval: 51.9%–52.5% · All markets

23,290 outcomes (18.1%) moved less than the symbol's noise band and are excluded as indeterminate — not counted for or against.

52.2%directional
105,343
validated predictions
50.9%
high-conviction (≥70%)
separate cohort · ranking, not odds

Top-conviction calls · held 25d

highest-conviction decile · 556 of 5,566 served calls · horizon auto-selected as the highest-accuracy slice
+3.02%
mean return per call, held to horizon
60.6%
directional accuracy of that slice

Live (served) predictions only — shadow-challenger calls, which users never saw, are excluded. No stop-loss and no price target: just the direction, held to the horizon. This is a research signal, not investment advice.

The trust funnel

last 90 days · no confidence floor

Every slot the system tried, through to the cohort the headline accuracy is computed on — including the ones we declined to call. Hover a stage for its exact definition.

Attempted
2,423,595
Abstained
1,657,648
Predicted
765,947
Pending Validation
374,547
Validated
389,878
60.0% accurate
Net Predictions (KPI cohort)
122,043
50.9% accurate

Accuracy over time

Rolling 7-day directional accuracy vs the 50% coin-flip line — All markets.

2026-06-1950% baseline (dashed)2026-09-17

Recent validation activity

most recent settle 2026-09-17

Predictions scored per day, by SETTLE date — the day the outcome actually resolved, not the day the call was made. This is the only series here that reaches today, which makes it the honest answer to “is the scoring engine running right now?”. Every chart below is by generation date and tapers at recent dates because those predictions have not matured yet.

Settled (30d)
68,925
predictions scored
Most recent day
1,359
2026-09-17
Days covered
31
with settlement activity
Hit rate that day
32.1%
correct / settled
-3462,2214,788
2026-08-18settledcorrect2026-09-17

Cumulative accuracy

Running accuracy across every scored prediction through each date — All markets. It converges as the sample grows: early swings shrink as evidence accumulates, so a flat right-hand tail is the sample getting large, not performance freezing. Ends short of today by design, because recent predictions are still maturing.

48.6%58.7%68.8%
2026-06-19cumulative accuracy · dashed = 50% coin flip2026-09-17

Daily accuracy vs 7-day rolling

Daily bars are noisy on thin days, so the rolling line is the one to read. Both are by generation date on the canonical cohort — the last stretch of each line is built from only the short-horizon calls that have matured, which is why it wanders more than the rest.

-6%23%52%81%
2026-06-19dailyrolling 7d50% coin flip2026-09-17

Scored predictions per day

Sample size behind each point of the accuracy charts above, by generation date. Spikes are scheduled batch prediction runs. The last few days read low because those predictions are still maturing — not because generation stopped.

02,2254,4492026-06-19: 4512026-06-20: 102026-06-22: 5552026-06-23: 2972026-06-24: 6832026-06-25: 6122026-06-26: 1,6972026-06-27: 8202026-06-29: 3772026-06-30: 5522026-07-01: 6092026-07-02: 6472026-07-03: 1462026-07-04: 12026-07-05: 2452026-07-06: 9422026-07-07: 5392026-07-08: 9982026-07-09: 9632026-07-10: 4632026-07-11: 752026-07-12: 212026-07-13: 7922026-07-14: 6172026-07-15: 2112026-07-16: 6692026-07-17: 6672026-07-18: 1152026-07-19: 1,8312026-07-20: 3,7132026-07-21: 1,6122026-07-22: 2,1052026-07-23: 1,8152026-07-24: 2,3912026-07-25: 792026-07-26: 162026-07-27: 2,3912026-07-28: 2,6562026-07-29: 2,2322026-07-30: 2,4422026-07-31: 4,4492026-08-01: 1712026-08-02: 482026-08-03: 2,3232026-08-04: 2,2542026-08-05: 2,3792026-08-06: 2,3702026-08-07: 3,2932026-08-08: 3652026-08-09: 182026-08-10: 2,0932026-08-11: 2,7902026-08-12: 2,7132026-08-13: 2,3422026-08-14: 3,4072026-08-15: 4322026-08-16: 132026-08-17: 2,3812026-08-18: 2,7562026-08-19: 3,2512026-08-20: 2,8672026-08-21: 2,5322026-08-22: 2282026-08-23: 372026-08-24: 2,4992026-08-25: 3,0472026-08-26: 2,9812026-08-27: 1,9342026-08-28: 2,4852026-08-29: 2732026-08-30: 112026-08-31: 2,3502026-09-01: 4,2002026-09-02: 2,7542026-09-03: 3,3272026-09-04: 4,4342026-09-05: 5392026-09-06: 432026-09-07: 1,9772026-09-08: 3,3252026-09-09: 2,9912026-09-10: 3,4142026-09-11: 4,3712026-09-12: 7932026-09-13: 232026-09-14: 3,1742026-09-15: 3,4192026-09-16: 3,1272026-09-17: 1,204
2026-06-19scored predictions (validated only)2026-09-17

Daily breakdown by maturation date

last 14 days due

Each row is the complete cohort of predictions that came due that day, so the counts do not taper the way the generation-date charts above do. Click a row for the per-horizon split.

Due dateCohortScoredCorrectPendingAccuracyRolling 7dStatus

By horizon

1d
49.3%2,815 calls
5d
48.3%22,030 calls
7d
47.8%25,694 calls
10d
53.2%20,409 calls
15d
54.7%19,301 calls
20d
59.6%3,127 calls
25d
59.6%4,956 calls
30d
64.0%4,687 calls
45d
61.7%2,320 calls

Bullish vs bearish

Bullish
52.2%105,343 up-calls

Walk-forward accuracy

Accuracy on the most recent N days only — is recent live performance holding up?

Last 30 days
39.0%
n=26,857
Last 60 days
48.6%
n=83,028
Last 90 days
51.0%
n=121,880

Do the probabilities mean what they say?

Calibration data withheld. An automated integrity check flagged this batch of calibration bins as internally inconsistent, so the curve is not shown rather than shown wrong. It returns automatically once the check clears.

Does confidence mean anything?

Mean gap 15.8ppWorst band 39.4ppn=105,343

Actual directional hit-rate by the model’s own stated confidence bucket — higher confidence should mean more calls come out right. This is a different measurement from the reliability diagram above, which checks whether the calibrated probabilities match realized frequencies. Ranking well here and calibrating badly there is a real, common combination, so read them separately.

Verdict: Overconfident — across the confidence bands the stated probability runs ahead of the hit-rate we actually deliver, by 14.7 percentage points on average once each band is weighted by how many predictions are in it. Treat high-confidence calls as a ranking, not as literal odds.Measured on the 105,343 scored predictions in the table below — 4 bands over-stated · 2 on target. Worst band: Extreme (90%+), stated 90.0% and delivered 50.6% over 18,930 calls (−39.4pp). Stated confidence is the band's mean stated confidence.
02550751000255075100Very Low (<50%): said 48%, delivered 49.4% (n=19,447)Low (50-60%): said 55%, delivered 55.6% (n=29,614)Medium (60-70%): said 60%, delivered 53.9% (n=8,193)High (70-80%): said 74%, delivered 53.2% (n=18,276)Very High (80-90%): said 86%, delivered 47.5% (n=10,883)Extreme (90%+): said 90%, delivered 50.6% (n=18,930)stated confidence %actual hit rate %
Dashed line is perfect calibration. Points below it are overconfident; circle size is the number of predictions in that bucket.
Confidence bandPredictionsWe saidWe deliveredDelivered − stated
Very Low (<50%)19,44747.8%
49.4%
+1.6pp
Low (50-60%)29,61454.7%
55.6%
+0.9pp
Medium (60-70%)8,19360.4%
53.9%
−6.5pp
High (70-80%)18,27674.3%
53.2%
−21.1pp
Very High (80-90%)10,88386.0%
47.5%
−38.5pp
Extreme (90%+)18,93090.0%
50.6%
−39.4pp

Calibration by model

lower Brier / ECE is better

The headline numbers above are the blend of every model that produced a scored call. Split out, they disagree — which is the useful part. Delivered − stated is realized accuracy minus stated confidence, the same convention as the band table above: negative means that model talks a better game than it plays.

ModelScoredAvg confidenceAccuracyDelivered − statedBrierECE
v1842,30257.6%46.0%−11.6pp0.2670.116
modular6,11458.3%40.8%−17.5pp0.2760.175
shadow1,55158.8%45.3%−13.6pp0.2690.136
besttoo few to judge2658.0%42.3%−15.6pp0.2670.156
news_causal_v2too few to judge756.2%0.0%−56.2pp0.3160.562

Risk-adjusted returns

all markets · all horizons · last 90 days · n=143,727

Measured on the disciplined book — the signals we actually commit to after the confidence gate, not every raw signal. This is the money-weighted result.

Computed on all markets · all horizons · last 90 days, 143,727 closed trades.

Book repaired. 268,849 trades in the ledger — 3,818 re-walked against the venue their prediction actually named, and 590 excluded as unresolvable (386 on a price series frozen at a single close, 204 whose venue could not be confirmed).Trades whose candles were walked against a venue the prediction did not name have been re-walked against the venue it did name, holding entry, stop, target, horizon and direction fixed so that only the venue changed. Rows whose correct venue could not be confirmed, or whose series is frozen at a single close (a data gap rather than a tape), are excluded rather than published as flat trades.
Sharpe
0.49
risk-adj (annualized)
Sortino
1.08
downside-only
Profit factor
1.34
gross win / loss
Win rate
48.6%
winning trades
Calmar
2.76
book return / max DD · 88d series
Expectancy
+0.71%
per trade (trade-weighted)
Max drawdown
-8.03%
peak-to-trough
Kelly fraction
0.06
implied fraction · ≤0 = no position

Tail risk & efficiency

How bad the bad trades get. VaR is the return the worst 1-in-20 (5%) and 1-in-100 (1%) trades beat; CVaR is the average of that tail — the number that matters when the tail actually arrives.

VaR 5%
-8.50%
worst 1-in-20 trade
CVaR 5%
-10.12%
avg of worst 5%
VaR 1%
-12.00%
worst 1-in-100 trade
CVaR 1%
-12.35%
avg of worst 1%
Information ratio
0.52
excess return / tracking error
Risk / reward
1.42
avg win / avg loss

Returns by horizon (holding-period)

How a user actually experiences it: hold from the call to its horizon. Sharpe here is annualized the correct way — per-trade × √(252 / horizon-days), not a calendar-year compound.

HorizonnAvg hold returnSharpe (holding-period)Win rateMax DD
1d2,868-0.271%-0.9749.0%-19.43%
5d22,336+0.206%0.2548.3%-10.39%
7d26,534+0.452%0.4548.2%-11.22%
10d26,015+0.547%0.3650.5%-16.05%
15d27,394+1.080%0.5453.4%-12.10%
20d5,532+2.150%0.7858.4%-21.49%
25d7,169+2.143%0.6657.0%-12.57%
30d11,362+3.659%0.9158.7%-30.47%
45d8,782+3.816%0.7559.7%-17.54%
60d3,131+5.380%0.7865.3%-22.35%

Trade funnel

Generated
434,465
Traded
268,849
Rejected
165,616
Abstained
1,657,535
Why we gate. Follow every raw signal — both directions, no confidence floor, including the low-conviction tails and long-dated horizons — and the naive backtest is negative. The book above is positive. Two things separate it from that firehose: it is long-only — short calls are still generated and scored, but not sold — and it applies the confidence floor. The long-only cohort is the larger of the two effects, and we do not attribute it to the gate.

Strategy returns (gated book)

last 90 days

Cumulative return of the disciplined trade book — the same daily equity series the Sharpe, Calmar and Max-drawdown figures above are built from, not the raw-signal firehose.

Basis: all markets · all horizons · last 90 days. Compounded over 88 trading days (2026-06-20 → 2026-09-17) — which is what the annualised figure extrapolates from. Book equity: each day's return is the equal-weighted mean of that day's closed trades, compounded across 88 trading days. Expectancy is a per-TRADE mean, so the two can differ in sign when trade counts vary sharply day to day.
Total return
+69.12%
gated book · 88d series
Annualised (turn-aware)
+22.18%
annualised on portfolio turns, not calendar time: an average 12.6-day holding period redeploys capital about 29.0 times a year, and the per-trade return is compounded over those turns across 143,727 closed trades
Avg return / trade
+0.692%
per-trade mean
Avg return / day
+0.612%
per-day book mean
How these reconcile. Per-trade mean +0.692% vs per-day (book) mean +0.612%. They agree in sign here, but they are still different weightings — per-day drives the curve and Total return, per-trade drives Expectancy.
-9.4%24.0%57.4%90.8%
2026-06-20book equity (dashed = break-even)2026-09-17
Max drawdown
-8.03%
peak-to-trough on this curve
Calmar
2.76
annualised return / max DD
Trades in book
143,727
closed trades
Trading days
88
2026-06-20 → 2026-09-17

By asset class

Stocks (Intl)
54.5%71,449 calls
Stocks (India)
46.5%16,683 calls
Stocks
45.5%14,580 calls
Commodities
67.1%1,177 calls
Crypto
60.7%728 calls
Forex
51.1%669 calls

By region

Asia-Pacific
56.0%44,904 calls
Europe
52.5%22,455 calls
India
46.5%16,683 calls
North America
45.5%14,580 calls
Latin America
51.3%2,541 calls
Middle East & Africa
48.2%1,549 calls
Commodities
67.1%1,177 calls
Crypto
60.7%728 calls

By sector

Unknown
51.6%84,843 calls
Financial Services
52.8%4,738 calls
Industrials
54.1%3,627 calls
Consumer Cyclical
57.0%2,459 calls
Technology
53.0%2,254 calls
Basic Materials
51.3%1,812 calls
Consumer Defensive
57.9%1,493 calls
Healthcare
55.4%1,308 calls

By model

v18
51.6%56,827 calls · 53.9% of book
modular
52.9%47,978 calls · 45.5% of book
news_causal_v2
54.1%529 calls · 0.5% of book

Accuracy by market × horizon

7 markets · 10 horizons · dimmed under n=30

Directional accuracy for every market at every horizon, against the 50% coin flip. This is the joint cut, not the two marginal breakdowns above: a market can look strong overall and still be at or below chance on the horizon you actually trade. Cohort: All markets · last 90 days.

Market1d5d7d10d15d20d25d30d45d60d
United States
43.2%
n=185
42.7%
n=5,243
38.3%
n=2,043
48.6%
n=2,567
50.6%
n=2,816
54.4%
n=182
50.5%
n=745
50.5%
n=542
46.3%
n=257
India
28.6%
n=227
47.6%
n=2,860
41.9%
n=5,059
46.9%
n=2,938
49.5%
n=3,128
51.8%
n=629
49.6%
n=716
55.7%
n=673
51.2%
n=453
Europe
52.2%
n=910
49.1%
n=4,944
48.9%
n=5,942
53.3%
n=3,861
54.9%
n=4,511
65.9%
n=478
62.1%
n=824
63.0%
n=635
60.0%
n=350
Asia-Pacific
51.4%
n=1,058
51.2%
n=7,754
51.1%
n=10,197
55.9%
n=10,098
57.4%
n=8,021
61.7%
n=1,614
64.2%
n=2,448
69.5%
n=2,609
71.2%
n=1,101
25.0%
n=4 · thin
Crypto
48.2%
n=85
56.6%
n=143
55.7%
n=334
76.9%
n=39
76.9%
n=39
91.7%
n=12 · thin
100.0%
n=22 · thin
74.5%
n=47
85.7%
n=7 · thin
Forex
53.3%
n=167
52.4%
n=42
42.1%
n=235
59.9%
n=132
68.9%
n=45
0.0%
n=11 · thin
54.2%
n=24 · thin
75.0%
n=12 · thin
0.0%
n=1 · thin
Commodities
53.1%
n=145
58.3%
n=300
65.7%
n=210
70.6%
n=204
75.8%
n=190
75.0%
n=36
90.2%
n=41
94.3%
n=35
93.8%
n=16 · thin

Green is above the 50% coin flip, red is below; intensity saturates at ±10 points so a thin outlier can't shout down a real one. 9 cells are dimmed for a sample under 30 — those numbers are shown as measured, but they are not yet evidence of anything. A cell with no scored predictions shows an em-dash rather than a zero.

Returns by market × horizon

10 markets · 10 horizons · dimmed under n=30

What following the calls actually returned, for every market at every horizon — the same joint cut as the accuracy grid above, in the unit that pays. These come from the trade ledger, not the scored prediction cohort, so the sample sizes are their own and will not match cell for cell. Cohort: All markets - last 90 days.

Metric
Average return per trade by market and horizon, up-calls only, All markets - last 90 days.
Market1d5d7d10d15d20d25d30d45d60d
United States
+0.13%
n=645 · 43% won
-0.30%
n=5,459 · 38% won
-0.19%
n=4,194 · 37% won
+0.01%
n=3,401 · 43% won
-0.04%
n=3,945 · 42% won
+0.22%
n=981 · 46% won
+0.31%
n=1,638 · 44% won
+1.40%
n=2,712 · 52% won
+2.09%
n=1,603 · 55% won
+3.07%
n=284 · 61% won
India
-0.40%
n=770 · 32% won
-0.15%
n=4,182 · 41% won
-0.19%
n=4,312 · 40% won
-0.11%
n=4,117 · 42% won
+0.11%
n=4,371 · 45% won
+0.81%
n=1,615 · 50% won
+0.58%
n=943 · 47% won
+1.00%
n=2,309 · 50% won
+1.89%
n=640 · 49% won
+4.90%
n=134 · 66% won
Europe
+0.13%
n=1,451 · 48% won
+0.01%
n=3,943 · 43% won
+0.14%
n=5,490 · 45% won
+0.32%
n=4,435 · 49% won
+0.39%
n=5,017 · 49% won
+1.73%
n=1,076 · 59% won
+1.38%
n=1,408 · 55% won
+1.33%
n=3,043 · 55% won
+1.85%
n=1,357 · 55% won
+1.33%
n=104 · 49% won
Asia-Pacific
+0.07%
n=1,094 · 43% won
+0.18%
n=7,008 · 44% won
+0.28%
n=10,262 · 43% won
+0.49%
n=10,595 · 45% won
+0.89%
n=11,020 · 50% won
+1.42%
n=4,236 · 53% won
+1.87%
n=3,738 · 54% won
+2.54%
n=7,018 · 56% won
+3.91%
n=3,411 · 63% won
+4.90%
n=1,121 · 67% won
Crypto
-0.85%
n=62 · 31% won
+0.43%
n=85 · 46% won
+0.49%
n=278 · 48% won
+2.20%
n=47 · 55% won
+1.24%
n=54 · 46% won
+4.82%
n=27 · 70% won · thin
+5.63%
n=42 · 64% won
+5.47%
n=85 · 64% won
+1.27%
n=20 · 50% won · thin
-6.50%
n=2 · 50% won · thin
Forex
+0.00%
n=188 · 54% won
+0.07%
n=48 · 58% won
+0.07%
n=169 · 58% won
+0.28%
n=143 · 69% won
+0.11%
n=56 · 63% won
-1.03%
n=21 · 24% won · thin
+0.08%
n=46 · 54% won
+0.14%
n=64 · 55% won
+1.32%
n=15 · 73% won · thin
-0.21%
n=3 · 33% won · thin
Commodities
+0.18%
n=193 · 50% won
+0.70%
n=286 · 52% won
+1.22%
n=297 · 57% won
+1.97%
n=276 · 62% won
+1.71%
n=241 · 58% won
+4.56%
n=81 · 73% won
+3.36%
n=67 · 69% won
+4.71%
n=112 · 71% won
+5.23%
n=46 · 70% won
-2.64%
n=4 · 25% won · thin
INDX
-0.17%
n=1 · 0% won · thin
-0.09%
n=4 · 25% won · thin
+0.25%
n=8 · 38% won · thin
+3.69%
n=6 · 100% won · thin
+2.23%
n=13 · 69% won · thin
-1.49%
n=3 · 33% won · thin
+3.35%
n=4 · 100% won · thin
+2.24%
n=11 · 91% won · thin
+2.02%
n=1 · 100% won · thin
Latin America
+0.57%
n=200 · 56% won
-0.35%
n=512 · 37% won
-0.11%
n=764 · 42% won
+0.78%
n=791 · 50% won
-0.42%
n=560 · 40% won
+1.65%
n=284 · 52% won
+0.24%
n=206 · 48% won
-0.66%
n=334 · 43% won
-1.16%
n=225 · 44% won
+6.54%
n=23 · 70% won · thin
Middle East & Africa
+0.19%
n=71 · 48% won
-0.17%
n=714 · 41% won
+0.73%
n=241 · 52% won
+1.18%
n=204 · 54% won
+1.97%
n=114 · 61% won
+0.22%
n=125 · 46% won
+1.66%
n=109 · 55% won
-4.07%
n=59 · 34% won

Green is a profit, red is a loss — the pivot is zero, not the 50% of the accuracy grid, because this is a return and the null result is “made nothing”. Intensity saturates at ±2% per trade, so a thin outlier can't shout down a real one. 17 cells are dimmed for a sample under 30 — those numbers are shown as measured, but they are not yet evidence of anything, and the annualized view magnifies them hardest. A market/horizon pair with no trades shows an em-dash rather than a zero.

What “annualized” means here. The average trade in a cell, compounded at that cell's own observed holding period — (1 + avg return)periods / avg hold days − 1. Periods is 252 trading sessions for equities and commodities and 365 days for crypto and forex, because a share cannot be traded on a weekend and a coin can. It assumes the capital is redeployed continuously into trades exactly like the ones measured. That is an extrapolation, not a realised return: nobody earned these figures, and a cell with a short hold and a thin sample can print a number in the hundreds of percent from a handful of trades. Each cell's tooltip states how many round trips a year its figure assumes — read that before believing the headline. A 1d trade is an intraday round trip: bought at the open, sold before the close. It is stored with the same entry and exit date, and is counted as a one-day hold — which is why the 1d column carries the largest annualized numbers on the grid, off some of the smallest per-trade averages.
This grid is long calls only. OQ no longer publishes short calls, so they are not counted here and not counted in the accuracy grid above. The reason is measured, not stylistic: following every published call — long and short together — returned −0.164% per trade after costs, while the long book alone returned +0.678% and won 49.8% of trades against 45.1%. The short book got worse the longer it was held, from roughly break-even at 1 day to −3.7% beyond 30 days. The models still generate short calls and we still score them privately; if that book earns its place back, it comes back. Neutral-direction trades belong to neither side and are excluded as well — 1 trade on that basis.
Cohort and basis. Realised returns from the oq_trades_v1 simulated ledger, scoped by exit_at inside the window; market bucket read from the trade's own exchange where the row carries one, otherwise inferred from the symbol via core.exchange_taxonomy.market_for_symbol; annualized = ((1 + mean(pnl_pct)/100) ** (periods / mean holding days) - 1), where periods is 252 trading sessions for equities and commodities and 365 calendar days for crypto and FX. A 1d trade is an intraday round trip (entry_at == exit_at) and is counted as a one-day hold. On multi-day rows the mean return covers EVERY trade in the cell while the holding period is measured only on those that ran past the first bar, so the two rest on different subsets — compare n with n_with_hold per cell. Market attribution: 134578 trade(s) carry their own exchange; 121827 predate venue stamping (2026-08-14) and are attributed from the symbol alone. A foreign listing stored as a bare ticker is indistinguishable from a US one, so those rows lean the US bucket high — measured at 13,860 of 133,473 rows (10.4%) on 2026-08-14. This affects WHICH market a trade is counted under, never its return, and never the all-markets total. Ledger: 256,405 trades over the requested 90-day window. On the multi-day rows, 27,717 trades closed inside their first session — a stop or target hit on day one. Their returns are in every average here, but they carry no holding period, so the annualized figure compounds a full-sample return over a holding period measured only on the trades that ran longer. Compare a cell’s n with its holding sample in the tooltip to see how far apart those two are. Cross-venue repair: separately, 3,818 trades were priced against a venue the prediction did not name — one bare ticker collapsing several venues and currencies into one record — and have been re-walked against the venue it did name. A further 590 could not be resolved and are excluded. These cells are computed on the repaired book, so they differ from figures published before that correction.

Net edge vs benchmark

Portfolio return minus what the benchmark did over the same window, per market, each against its own named local index. Positive means the calls beat simply holding that index; negative means they did not, and it is printed here either way.

Measuring net edge against each market's benchmark…

Performance by market regime

computed 2026-09-17

Not just accuracy per regime — the return, Sharpe and win rate of the calls made in each market state. This is where you find out whether an edge is real or is one regime carrying the whole record.

Tagged predictions
140,085
carry an ex-ante regime tag
Total in window
140,774
predictions considered
Tag coverage
99.5%
of the window is tagged
Regimes observed
2
volatility x trend

Volatility x trend (ex-ante tags)

The market state recorded at the moment each call was made — knowable in advance, so these numbers are the honest ones.

RegimenAccuracyAvg returnSharpeWin rateCoverage
high mean reverting71,289
53.6%
+1.65%2.1553.6%50.6%
medium mean reverting68,608
51.1%
+0.86%2.0351.1%48.7%

Realized move size (post-hoc diagnostic)

Not a regime, and not a performance claim. These rows bucket the predictions with no ex-ante tag by how large the move turned out to be — which conditions on the outcome and is not knowable when the call is placed. A high accuracy in the “small” bucket is expected and means nothing tradeable: a small realized move is easy to call directionally after the fact, and a near-0% average return on it is arithmetic, not a red flag. Read this as a diagnostic of where errors concentrate, never as an achievable edge.
Move sizenAccuracyAvg returnSharpeWin rateCoverage
small537
52.9%
+0.22%1.0852.9%0.4%
large152
57.9%
+2.54%3.9657.9%0.1%

Best tracked

SNR.L53/53100%
BEZ.L43/43100%
SOCI.JK34/34100%
9101.T32/32100%
TOTL.JK30/30100%
0019.HK29/29100%

Worst tracked

— shown honestly
GMRAIRPORT.BO0/370%
BLTZ.JK0/320%
5201.T0/270%
SMMA.JK0/250%
1812.T0/250%
VIE.PA0/230%

Discipline & coverage

Most attempts never become a call at all — the gate abstains upstream, before a prediction is issued, and those abstentions are counted against us in the two figures on the right. What remains is scored in full. Committed calls here are the published cohort, which applies no confidence floor — every scored prediction is included.

Committed calls
105,343
Abstained
1,657,535
59.1% of attempts
Accuracy (committed)
52.2%
Coverage rate
6.0%
of all attempts
Acc if abstain = wrong
3.1%
every abstention counted as a miss
Acc if abstain = 50/50
50.1%
every abstention counted as a coin flip

Abstention by horizon

platform-wide 59.1%

Share of attempted calls the engine declined to publish, per horizon. Higher means more was withheld.

HorizonAttemptedAbstainedAbstention rate
1d109,08779,532
72.9%
5d225,716123,891
54.9%
7d161,53369,762
43.2%
10d250,245138,898
55.5%
15d301,892173,229
57.4%
20d233,899133,922
57.3%
25d241,493141,403
58.5%
30d354,446225,143
63.5%
45d309,760192,066
62.0%
60d310,730193,074
62.1%
90d303,956186,615
61.4%

The intelligence stack

last promotion 2026-07-06

Which models are serving the predictions scored above, and the rules that govern replacing them.

Active models
1
serving live predictions
Auto-promotion
enabled
challenger can replace champion
Modular combo
active
per-symbol model blending
Swarm consensus
60%
agreement needed to commit

Currently serving

Cross-reference these against the “By model” breakdown above — a version with a small share of the book is a challenger being measured, not the champion.

DAILY_LIGHT_v20260622_1414_90d_AU

Top forecasters

No per-agent forecast records scored yet. This leaderboard fills once the swarm attributes settled outcomes to individual agents — until then there is nothing here to rank, and we would rather show that than an empty table.

Cohort definition & audit trail

canonical_trust_strict

Exactly which predictions the numbers above are computed over, and what was filtered out to get there. Published so the accuracy figure can be checked rather than taken on faith.

Strict directional accuracy — a prediction is scored correct only when the symbol moved beyond its noise band and the model called that direction. Band-indeterminate outcomes (see indeterminate.band_description for the exact rule in force) are excluded from both the numerator and the denominator — not counted for or against. Cohort also excludes archived, abstained, out-of-universe, and shadow-challenger predictions. No confidence floor is applied: every scored prediction is in the cohort, including the ones the model was least sure of. Short (DOWN) calls are excluded: the platform publishes long calls only, so the track record describes the product actually sold. Short predictions made before 2026-08-23 remain in the database but are not scored here.

Cohort size
267,375
predictions in scope
Snapshot age
82m
stale past 3h
Validator last run
2026-09-17
outcome scoring pass
Facet health
nominal
no failed facets

Cohort invariants

Every prediction counted above satisfies all of these. Each one removes a population that would otherwise flatter the number.

  • 1archived != True (exclude pre-2026-05-07 broken-pipeline cohort)
  • 2abstain != True (exclude model-declined predictions)
  • 3in_universe != False (exclude out-of-universe symbols)
  • 4is_neutral != True (exclude neutral predictions from directional accuracy)
  • 5is_shadow != True (exclude shadow-challenger predictions — the ~75% not shown to users)
  • 6confidence >= 0 (no confidence floor — every scored prediction is included, including the lowest-confidence ones)
  • 7direction in ("UP","up","Up") (long-only — short calls are neither published nor scored)

Filters applied

The query this snapshot was built with. null means unfiltered.

days back
90
exchange filter
model version
horizons

Excluded from the cohort

Printed exactly as the backend reports it — where a count is not separately tracked it says so rather than showing a zero.

archived count excluded
see #202 archive script — 11,153 docs archived 2026-05-16
abstained count
1,657,535
out of universe count
not separately surfaced
neutral count
23,290
low confidence count
not separately surfaced

Top-conviction ledger

All markets · last 90 days

The highest-conviction decile of settled calls, ranked on the calibrated probability — not the stated confidence. Held to the horizon with no stop and no target, which is why these rows reconcile with the hero number above and not with the traded book.

Sign in to inspect the raw ledger

This is the conviction slice itself. These are the individual predictions behind every aggregate number above — one row per call, with the outcome it actually settled to. The aggregates stay public; the row-level record is for signed-in members.

Sign in

Audit ledger

All markets · last 90 days · scored rows only

Every scored prediction in the canonical cohort, newest settlement first. This is the same set of rows the headline accuracy is computed over — the six invariants noted beneath the table are applied identically — so this is the panel to audit against. The recent-settled panel is a wider, non-canonical cohort and will not tie out.

Sign in to inspect the raw ledger

This is the canonical audit trail. These are the individual predictions behind every aggregate number above — one row per call, with the outcome it actually settled to. The aggregates stay public; the row-level record is for signed-in members.

Sign in

Recent settled predictions

All markets · last 100 returned

The latest calls that have already settled — pending ones are pulled from the same response and filtered out here, so this is scored rows only. This cohort is not the canonical one. It excludes archived and shadow-challenger rows only — it does not apply the six canonical invariants the headline accuracy above is computed on, so tallying these rows will not tie out to that number. The audit ledger panel is the cohort that does. Every market, because the scope above is set to Global.

Sign in to inspect the raw ledger

This is the recent settled record. These are the individual predictions behind every aggregate number above — one row per call, with the outcome it actually settled to. The aggregates stay public; the row-level record is for signed-in members.

Sign in

How we measure

Directional

Right if the price moves the way we said by the horizon. Unambiguous, checkable.

Close-to-close

Official closes at the horizon date — not intraday ticks we could pick to flatter.

No look-ahead

Locked before the outcome window opens. Out-of-sample, every time.

Probability, then scored

Every call ships a stated probability and we publish how far it lands from the realised hit-rate — including when that gap is bad.

The record is the proof. Now see the picks.

Today’s high-conviction calls for Commodities.

View picks