OPEN RESEARCH · EVERY RESULT PUBLISHED · INCLUDING THE FAILURES

What 15,000+ AI Chart Reads Taught Us

We run AutoProphets as a research program, not a black box. Below is every major experiment: what we asked, how we tested it, what we found — and what we got wrong. Each card starts with a TL;DR. Expand it for the full story.

Public ledger updates as new research is published.
Start Free 7-Day Trial Explore the Public Ledger →
28
AI models benchmarked head-to-head
200
blind charts per fleet run
15k+
scored AI predictions
23
research laws derived
ADoes the fleet actually work?

Live production ledger — timestamped outlooks remain public as the research program continues.

Live data
The public record: 718 timestamped calls
The complete ledger preserves each call and its outcome metadata. Definitions, limitations, and the research process remain open for inspection.
718 calls retained0 calls removedLIVE public ledger

The question

Can language models reading live market context predict BTC's next 24–48h direction better than a coin flip — and can we PROVE it rather than claim it?

How we test

  • Each outlook is written to an immutable ledger with a timestamp before the forecast window begins.
  • 48 hours later, an automated checker fetches the exact Binance BTCUSDT hourly candles for that window.
  • Predicted direction (Bullish / Bearish / Neutral) is compared to the realized move; a move within ±3% counts as Neutral.
  • Scoring has run hourly since Dec 2025 with zero manual intervention — nobody touches the grades.

What we found

The research program records model behaviour across time and keeps the underlying ledger available for review. No performance headline is presented here.

What it changed

Definitions and source data stay linked to the public research record so readers can inspect how each entry is created and evaluated.

Running
Is it the news, the indicators, or the search grounding? A forward A/B is deciding
Backtested without live search, article-only and indicator-only contexts perform similarly (~60%), and price-only looks deceptively strong. Our hypothesis: the real edge is live search grounding. A daily shadow test is now collecting the evidence.
62.5% articles-only62.5% indicators-only83.3% price-only (artifact)24 weeks backtested

The question

Which context stack deserves the production outlook engine: our own article pipeline, technical indicators, both, or nothing but a price snapshot? (Popular wisdom says indicators; our own earlier season said news.)

How we tested (backtest)

  • 24 weekly cutoffs spanning 6 months; for each, four prompt variants saw ONLY data timestamped before the cutoff (no future leak, no search).
  • Scored with the identical ±3% band logic as production, against the real next-48h candles.

What we found

Articles ≈ indicators ≈ both — all mediocre without grounding. "Price-only" scored highest (83% lenient) but 19 of 24 windows moved less than ±3%, so "Neutral" was nearly always right: a quiet-regime artifact, not skill. Conclusion: static context is not where the alpha lives.

What it changed

We built a forward shadow test: every day, a no-grounding twin of the production outlook is generated and scored by the same automated checker. After ~2 weeks of production data we'll have a statistically honest verdict on what search grounding contributes. This page will publish it either way.

This research isn't theoretical — it runs a live fleet, in public, right now. Every finding above is applied to the 10 prophets trading today.
Watch the Research Work — Free Trial
BTeaching AIs to read charts

Multimodal benchmark series V1–V17: models look at rendered candlestick charts and must decide. Blind datasets, balanced regimes, z-scored results.

Twenty-eight AI model pods undergoing a blind Bitcoin chart-reading benchmark in a monitored research lab
Surprise finding
The EMA Trap: adding one popular indicator made AIs worse
Candles alone: 63% accuracy. Add only EMA lines: 55% — the extra lines induced a systematic SELL bias. Add the full indicator suite: back up to 70–80%. More information is not more signal.
63.3% candles only55.0% EMAs added70–80% full TA

The question

Does visual indicator clutter help or hurt a vision-language model reading a chart image?

How we tested

Identical chart windows rendered at 8 indicator levels (none → candles → EMAs → full TA → mega-clutter), shown to the same models under identical prompts. 3 models × 3 levels × 20 charts, then a 160-call expansion on one model.

What we found

A U-shaped accuracy curve. EMAs alone anchor the model bearish (falling-price frames dominate its attention). The full suite — RSI + MACD + Bollinger in dedicated sub-panels — calibrates the bias back and beyond. Beyond full TA, accuracy falls again (information overload).

What it changed

The production chart renderer draws exactly the "Goldilocks set": EMA20/50, BB, RSI and MACD in dedicated panels. This finding is white-paper material — nobody else we know of has quantified indicator-induced bias in vision models.

Benchmark
Timeframe is King: daily charts win, hourly charts lose money
On daily candles AIs hit 70% win rate. On 1-hour charts they managed 36.7% — statistically worse than random. The same models, the same prompts: only the candle interval changed.
70% 1D win rate36.7% 1H win rate50% 4H (volatile)

The question

Which candle timeframe plays to AI vision strengths?

How we tested

7 experiments across 210 simulated trades (V5 series): identical models and prompts, charts rendered from 15-minute to 1-week candles.

What we found

1D is the sweet spot: enough candles to show structure, little noise. 1H is noise-dominated — the models' 36.7% means their reads were actively wrong. 4H is volatile middle ground; 1W too slow to matter for a 48h horizon.

What it changed

All chart prophets render daily-context charts. We also learned the meta-lesson: a "smart" model on the right timeframe beats a "genius" model on the wrong one.

Benchmark
Model economics: cheap models can win — until you scale
Gemini Flash Lite hit 85% on full-TA charts in one small run. But at fleet scale (200 charts), the best models settle at 57–58% — a few points above coin flip. Good can be cheap; consistent is hard.
85% best single run58.0% best at N=200 (z=2.26)

The question

How expensive must a model be to read charts well — and do small-sample hero runs survive scale?

How we tested

Model races from 10 to 28 models, 20 to 200 blind charts each, with per-model z-scores against the 50% null. V17 then re-ran the full fleet on stripped-down prompts to isolate engineering effects.

What we found

  • Small cheap models genuinely compete (mistral-small-2603: 57.5%, z=2.12 at N=200).
  • Headline accuracies from small samples evaporate: a V11 star at 73% on 30 charts collapsed to 48% at scale (Law 17).
  • Model choice matters more than prompt engineering (Law 8) — but engineering still helps 68% of models (Law 20).

What it changed

Production runs on the cheapest model that passes benchmarks, with quarterly re-checks. Every model swap must beat the incumbent in a blind race first — never on vibes.

CWhat to feed the AI (and what poisons it)

Context ablations V11–V14: personality, memory, crowd sentiment, commentary — each tested as an injected variable against blind chart performance.

Surprise finding
The personality premium: "Poker Champion" beat analysts by 23 points
We tried 15 trading personalities on the same charts. Game-theory minds (Poker Champion +23.3pp, Chess Grandmaster +16.7pp) beat every analytical persona, and all 15 beat or matched the baseline. How an AI thinks matters more than what it knows.
+23.3pp Poker Champion+16.7pp Chess Grandmaster15 personas tested15/15 ≥ baseline

The question

Do personas change decisions, or just tone?

How we tested

V12: 15 personality prompts × 30 blind charts on the cheapest model, then V12.5 hybrid fusions to check we weren't cherry-picking.

What we found

Strategic/game-theory framings (odds, bluffing, opponent modeling) produced decisively better calls than analytical/scientific framings. Direct prompt concatenation did nothing; novel thinking styles created new alpha. A Contrarian Sage independently matched Poker Champion.

What it changed

The ten production prophets are built on the winning archetypes (Poker Champion, Chess Grandmaster, Contrarian Sage, Trend Follower, EV Maximizer…). Their "personalities" are not marketing garnish — they are the measured alpha.

Negative result
Popular sentiment data made the AIs worse: Fear & Greed is toxic
Injecting the Fear & Greed Index cost −8.2% alpha. Injecting our own AI-generated market commentary cost 10 points of win rate. Both are crowd anchors — the models herd instead of analyze.
−8.2% alpha w/ F&G−10pp win rate w/ commentary−3.3pp win rate w/ trade memory

The question

Everyone feeds sentiment data to trading AIs. Should we?

How we tested

V13/V14 ablations: 30 sequential blind trades per condition, then a 5-model × 6-condition suite (900 trades) covering commentary, F&G, open interest, funding, identity and kitchen-sink combinations.

What we found

  • Fear & Greed: toxic. The model anchors to the crowd exactly when it should fade it.
  • Market commentary: anchors the vision model to someone else's directional opinion.
  • Trade memory: the AI anchors on its own past P&L (Law 6: memory is poison).
  • Open interest helped (+2.5% alpha); a structured identity helped more (see next card). Kitchen-sink combos were worst of all.

What it changed

Chart prophets receive zero commentary and zero sentiment indices — a deliberate, benchmark-proven starvation diet. This is the opposite of how most "AI trading" products are built, and it came from publishing our negative results.

Benchmark
Consensus works — until the crowd gets too big
A 3-model majority vote beat the best individual model (+16.2% alpha). A curated 5-model vote is free alpha. But the full 28-model swarm scored 47–48% — worse than a coin flip. Diversity has an optimum, not a maximum.
+16.2% alpha (3-model vote)56.0% consensus at N=200 (p=0.045)47–48% 28-model swarm

The question

Does the wisdom of crowds apply to AI chart readers?

How we tested

V9–V10 (fleet trading + consensus votes), V15 (consensus at N=200 with z-scores), V16/V17 (full 28-model swarm head-to-head).

What we found

Small curated consensus: reliable, statistically significant, and free (reuses existing runs). Unfiltered swarm: errors correlate, weak models drag the vote below chance. The skill is in curating voters, not counting them.

What it changed

The platform's consensus prophet uses a small curated voter set with published per-model stats — and the voter list changes only through benchmark evidence.

Benchmark
Identity beats data: a named agent with a structured self outperforms
Swarm agents given a persistent identity scored 63.3% vs 60.0% anonymous, with better P&L and profit factor. A structured identity prompt outperformed feeding the same models extra sentiment data.
63.3% identity swarm60.0% anonymous swarm+75.7% vs +63.4% P&L

The question

V11 asked whether giving agents a persistent identity (name, style, memory-of-self) changes decision quality — and whether identity substitutes for more data.

How we tested

28-model swarm × 30 charts × two conditions (identity vs anonymous), 1,680 API calls.

What we found

Identity won on every metric: accuracy, P&L, profit factor (2.58 vs 2.17). High-conviction identity agents (>70% confidence margin) reached 66.7%. In V14, structured identity also beat injecting extra sentiment data — coherence of self > volume of inputs.

What it changed

Every prophet has a persistent name, persona, and self-consistent memory policy. That's why the leaderboard has characters, not "model_7". The characters ARE the experiment.

DHow we keep ourselves honest

Methodology and integrity — the parts most operations skip, and the reason the numbers above mean anything.

Method
Blind mode, balanced datasets, and z-scores: the anti-cheat toolkit
Models see charts with no asset names or dates. Bull and bear windows are balanced 50/50. Every headline result carries a z-score against the 50% null. A July 2026 integrity audit found 6 weaknesses in our own earlier benchmarks — so we rebuilt them.
blind chart rendering52/48 balanced biasz≥2.0 significance bar6 flaws self-audited & fixed

Why this exists

Our V1–V3 benchmarks produced spectacular numbers (+1,127% returns!) that turned out to be phantom profits from a leverage accounting bug. We invalidated them publicly, listed the 12 methodological flaws, and rebuilt as V6 with production-replay mechanics and anti-cheat hardening (direction-neutral prompts, overtrading penalties, fee-adjusted scoring).

The standing rules

  • No asset labels or dates on chart images (blind mode).
  • Balanced bull/bear chart sets; seeds published for reproducibility.
  • N≥200 for any "leader" claim; N=30 results are treated as flukes until proven otherwise.
  • Negative results are published as laws, not buried (Laws 6, 12, 13, 23 are all "don't do this").

Full lab notebooks: quantitative series · multimodal series V1–V17 · exact prompts used.

Reference
The 23 research laws — one page of everything we learned
Every experiment series ends by promoting its sturdiest finding to a numbered "law". These are the load-bearing conclusions of the whole program.
23 laws17 benchmark versions4 laws are negative results
Law 1Timeframe is King. Daily candles: 70% win rate. Shorter timeframes decay toward worse-than-random.
Law 2Full Technical Analysis wins. EMA + BB + RSI + MACD together beat any subset.
Law 3Less is More. Beyond the full-TA set, extra indicators reduce accuracy (overload).
Law 4Prompt identity is model-specific. The best persona for one model isn't the best for another.
Law 5Regime context is the real winner. Trend/volatility framing beats most other injected text.
Law 6Memory is poison. Trade history anchoring degrades decisions.
Law 7The Bear Detector effect. Models carry an inherent SELL bias; full TA calibrates it.
Law 8Model choice matters more than prompt. Pick the model first, then engineer.
Law 9Accuracy ≠ alpha. The accuracy champion wasn't the P&L champion.
Law 10Consensus beats individuals. A 3-model majority vote out-earned the best single model.
Law 11Personality is free alpha. Game-theory personas add up to +23pp.
Law 12Market commentary is noise. Injecting outlook text costs ~10pp win rate.
Law 13Crowd sentiment is toxic. Fear & Greed injection: −8.2% alpha.
Law 14Identity beats sentiment data. Structured self > more inputs.
Law 15Scale crowns new champions. At N=200, llama-4-scout led (58%, z=2.26).
Law 16Best-value consensus is curated and free. 5-model vote from existing runs.
Law 17N=30 results are flukes. A 73% star collapsed to 48% at scale.
Law 18Models absorb market regime. Bear-period training data bleeds into SELL bias.
Law 19Specialists exist. nova-lite: 84% SELL bias, +267% P&L as a short specialist.
Law 20Engineering helps 68% of models — not all. Test per model.
Law 21New fleet leaders at N=200. mistral-small-2603 57.5% (z=2.12), ministral-14b 57.1%.
Law 22Simple prompts break some models (0% parse rates). Robustness matters.
Law 2328-model consensus is worse than top individuals. Curate voters; never max them.
Method
What this research can't tell you (yet)
Single asset. Nine months of scored history. Paper execution. We publish the limits alongside the results because the limits are part of the results.
BTC only so far9 months verified historypaper execution only

Known limits — on the record

  • One asset. ETH/SOL fleets are the natural next experiment; multi-asset runs exist in the lab but not yet in production scoring.
  • One macro regime mix. The verified span includes a downtrend and recovery, but no full bull mania. The forward record keeps accumulating through whatever comes.
  • ±3% neutral band. Reasonable, but a choice. Tighter bands make the task harder; looser, easier. The band is published so anyone can re-score.
  • Paper only, by design. We never execute real money — that's both a compliance posture and the point of the product: transparent simulation you can audit.
  • Model providers can change under us. Model-era analysis shows accuracy tracks model capability; quarterly re-benchmarks guard this, and any swap requires a blind race win.
EDoes utility beat hype in a bear market?

The Utility Index — a published, rules-based SIMULATED portfolio built at maximum fear on Aug 14, 2026. Six public criteria picked the components; an automated referee grades it against Bitcoin daily. Paper research only — never advice, never executed.

Live data
The Utility Index vs Bitcoin — every day since publication, in public
15 components across five sleeves (bedrock networks, exchange rails, AI & data, one gold anchor, tiny new-gen), fixed weights, pre-registered rebalances. The ledger is append-only — bad days stay published.
15 components6 public criteriadaily gradingpaper simulation only

The question

In a market where everything fell, did tokens with real revenue and real usage hold up better than the reference asset — and does a transparent rulebook capture that without any discretion?

How it runs

  • Components rated 1–10 on six published criteria; weighted score set the sleeve allocations.
  • Fixed weights between pre-registered rebalance events (first scheduled review: Oct 2026). NAV recomputed daily from CoinGecko end-of-day marks; Bitcoin benchmark uses the identical method.
  • History is append-only. If a source fails, the day is flagged stale — never silently rewritten.

Where to see it

The full live tracker — chart, component contributions, sleeve rollups, rulebook, per-asset theses and risks — is on the Utility Index page.

Watch the Research Run Live — Start Free Trial Explore the Public Ledger

🔓 Pro Deep Data — take the research home

Members don't just read the findings — they get the data behind them. Every Pro account can mint API keys for programmatic access.

/v1/track-record full scored ledger /v1/outlooks every call + verdict /v1/leaderboard live fleet returns 60 req/min included

What's included with Pro: the complete scored-outlook ledger (not just the summary charts), per-month model-era breakdowns, benchmark result archives, and your own API key for bots, sheets, or dashboards. An MCP server (let AI assistants like Claude query our data directly) is on the roadmap for Desk-tier members.

Unlock With Pro — $29/mo

Educational & entertainment research only. AutoProphets runs simulated paper portfolios. Nothing here is investment advice, a solicitation, or a promise of performance. Crypto assets are volatile and you can lose money. Past accuracy does not guarantee future results.