When a large M&A deal is announced, which other stocks move — and can a vector-association engine find them better than a frontier LLM? 154 announced US-relevant deals, Jan 2024 – Jul 2026. Research only — not investment advice.
This report was generated and validated by Claude Fable 5 and Claude Opus 5.
Note what that means for reading it: two of the three selectors graded here are the same model family that
produced and checked this report, and both are scored below our own Tuatara model. Every number is
recomputable from the raw CSVs and the verify.py script below — do not take the attribution as
assurance, run the script.
| Hidden-leg profile | Tuatara | Claude Fable 5 | Claude Opus 5 |
|---|---|---|---|
| Median leg float | $0.20B | $12.96B | $8.21B |
| Legs under $250M float | 52.5% | 3.6% | 4.1% |
| Qualifying events (of 117 big deals) | 44 | 15 | 19 |
| Hidden legs generated | 971 | 1,031 | 1,115 |
Flagship configuration: deals > $5B, leg float < $250M, entry price ≤ $20, long only, hold 1 trading day, excess return vs SPY, per event.
| Selector | Per event | Hit rate | t-stat | n |
|---|---|---|---|---|
| Tuatara | +3.77% | 61.4% | +1.5 | 44 |
| Claude Fable 5 | −1.26% | 33.3% | −1.6 | 15 |
| Claude Opus 5 | −0.95% | 26.3% | −0.8 | 19 |
Same events as the table above, taken in date order, with the full balance rolled into each successive event. Raw basket return (entry to exit), no SPY subtraction.
| Selector | Events | Compounded | $1,000,000 becomes |
|---|---|---|---|
| Tuatara | 44 | +202.99% | $3,029,893 |
| Claude Fable 5 | 15 | −20.18% | $798,243 |
| Claude Opus 5 | 19 | −18.92% | $810,832 |
The compounded spread (+202.99% vs −20.18% / −18.92%) is far wider than the per-event spread (+3.77% vs −1.26% / −0.95%). That gap is a property of compounding, not a bigger edge.
| Raw basket returns | n | Mean | Median | Max | Min | SD | Skew |
|---|---|---|---|---|---|---|---|
| Tuatara | 44 | +3.53% | +0.01% | +79.28% | −22.38% | 16.38 | +3.13 |
| Claude Fable 5 | 15 | −1.43% | +0.18% | +2.86% | −8.29% | 3.59 | −0.73 |
| Claude Opus 5 | 19 | −0.98% | −1.07% | +14.89% | −8.70% | 4.99 | +1.28 |
| Selector | Compounded | Best single event | Without that one event |
|---|---|---|---|
| Tuatara | +202.99% | +79.28% | +69.01% |
| Claude Fable 5 | −20.18% | +2.86% | −22.39% |
| Claude Opus 5 | −18.92% | +14.89% | −29.43% |
| Float bucket | n | Mean leg return | Median | Max |
|---|---|---|---|---|
| Under $10M | 110 | +2.49% | +0.00% | +102.90% |
| $10M – $100M | 84 | +5.63% | +0.00% | +519.05% |
| $100M – $250M | 98 | +0.26% | +0.00% | +17.94% |
| Horizon | Tuatara | Claude Fable 5 | Claude Opus 5 |
|---|---|---|---|
| 1 trading day | +0.93% | −0.09% | −0.10% |
| 2 trading days | +1.17% | +0.13% | +0.20% |
| 3 trading days | +0.43% | +0.26% | +0.25% |
A published sub-claim failed on re-run. The earlier report said the attention gate inverts on LLM picks (at a $2B float ceiling: high-attention −1.19%, t = −2.4, vs low-attention +0.75%). On Opus 5 the same cut shows high −0.15% vs low −0.31% — no separation in either direction. Treat that inversion as noise at n≈16. The headline null result for LLM picks does replicate; the inversion does not.
Everything below still applies to every number on this page:
A companion finding from the live platform. Event cards trade as weighted stock baskets; the same seven-stock biotech basket ("Bispecific antibodies sweep into the clinic", launched 2026-08-05) was scored under five weighting schemes. Weights change the story more than the stocks do.
| Weighting scheme | CNTX weight | Index return |
|---|---|---|
| Price-weighted (the live index definition) | 1.65% | +3.32% |
| Equal-weighted | 14.29% | +1.04% |
| Tuatara Fibonacci ladder (φ−rank by relevance score) | 15.11% | +0.33% |
| Weighting | 90d return | Sharpe | Sortino | Max DD | OOS return | OOS Sharpe |
|---|---|---|---|---|---|---|
| Price-weighted (as submitted) | +27.2% | 2.63 | 5.05 | 15.2% | +14.0% | 3.63 |
| Tuatara golden-softmax | +8.7% | 1.13 | 1.92 | 21.9% | +7.1% | 2.24 |
| Tuatara Fibonacci ladder | +7.1% | 0.97 | 1.63 | 24.5% | +7.1% | 2.27 |
| Inverse-volatility | −3.1% | −0.09 | −0.13 | 20.7% | +0.1% | 0.27 |
| Equal weight | −9.5% | −0.71 | — | — | — | — |
A reader made the right objection: asking a frontier model whether our numbers look fine is not verification — a model that never saw the price data cannot audit anything. So here is every qualifying event, with tickers, dates and entry/exit prices.
| Target (deal) | Announced | Entry t0 | Exit t1 | Attn × | Legs | Basket | SPY | Excess |
|---|
| Cut (Tuatara, high-attention, flagship) | n | avg | median | t | hit |
|---|---|---|---|---|---|
| as published | 44 | +3.77% | +0.32% | +1.50 | 61.4% |
| require ≥3 priced legs | 28 | +6.76% | +0.32% | +1.81 | 60.7% |
| drop baskets with any leg < $0.01 | 22 | +5.16% | +0.13% | +1.09 | 54.5% |
| drop baskets with any leg < $1.00 | 10 | −1.55% | +0.29% | −0.61 | 70.0% |
| drop the single best event (Roku, +79.87%) | 43 | +2.00% | +0.32% | +1.09 | 60.5% |
What does survive every cut above is the selector comparison — it rests on the float profile and basket overlap of ~1,000 legs per selector, not on a handful of tail events. "Tuatara finds different names than an LLM does" holds. "Those names make +3.77%" does not.
What this comparison was. A capability test — a domain-specific vector engine against frontier LLMs from multi-trillion-dollar labs, on the LLMs' best terms. The durable finding is the selection asymmetry: the frontier models never surface the low-float universe at all. It was never a strategy pitch; the stress-cut section above already says nobody should trade the paper configuration.
Tuatara ran un-optimized and uncombined — the floor, not the ceiling. Raw association output: no task-specific tuning, no ensemble, no execution layer; the only fitted element (the flagship filters) is disclosed as multiple-testing risk. The selectors were deliberately kept separate for clean attribution — and in our separately published equity headline study, the Combined (Tuatara ∪ Claude) book posts the best risk-adjusted result on the page (per-trade Sharpe 0.79 vs 0.70 and 0.27). Precision, after a reader pushed on this: that Combined book is a mechanical union — complementary coverage plus diversification, not proof of model synergy — and the actual production architecture (an LLM independently ranking and vetting vector-surfaced candidates) remains unmeasured. See §13.3–13.4 of the full report for the input-parity and correctness-labeling tests that would settle the mechanism and accuracy questions.
The validation bar (reader-contributed, adopted). What would establish the stronger, executable-alpha claim — now adopted as the forward-test spec in §13.4 of the full report: timestamp-safe entries at the first post-announcement price; as-of-date float; strictly pre-entry information sets with a fixed ex-ante attention threshold; execution modeled from quotes/spreads/volume with capacity limits (or real logged fills); matched random microcap controls; pre-specified, logged human-in-the-loop rules; and blind validation of the surfaced relationships. Until that run exists, the claim is and remains: a differentiated long-tail discovery engine — with the executable-alpha question open in both directions.
Published before results, not after. The executable-alpha question is settled by a prospective forward test, not by more reanalysis of the historical sample. The full protocol — the validation bar above, instantiated with concrete non-negotiable values — is published at this URL and timestamped before the first forward event is scored. If it changes before the start, the prior version stays published alongside it with the reason. Nothing is amended once results begin arriving.
Fixed in advance: entry rule and fill convention; holding period; the ex-ante attention threshold as a fixed number; candidate universe and screens as of the event date; benchmark; matched-random control construction; sizing and any human-override rules with their logging format; the primary endpoint and the stopping rule; and the blind relationship-labeling procedure with rater instructions.
One primary endpoint, declared up front. Every other cut is labeled secondary/exploratory, whatever it shows. We will not promote a secondary cut to the headline because the primary one disappointed.
The protocol is now instantiated, not just described (§15.6, locked 2026-08-02). Concrete values, fixed before scoring: entry at the first regular close strictly after the captured announcement timestamp (after-hours deals enter the next session); 1-day hold; excess vs SPY; as-of-date float < $250M and price ≤ $20, unknown float excluded; a fixed attention threshold of 5.0 — a constant, not a per-sample median, since the in-sample median was itself a lookahead — measured over a window ending at entry so no post-entry data enters the gate. One primary endpoint: HIGH-attention legs minus matched-random controls, judged net of modeled spread/impact, evaluated once at n=100 scored events with no interim peeking. Forward capture began 2026-07-28; 36 deals captured, none scored at the time this was locked.
Negative results get published the same way. The forward result appears at this URL with the same prominence whichever way it falls — including the case where it kills the executable-alpha claim outright. If forward high-attention legs do not beat matched random controls on the declared primary endpoint, the claim is rejected and this report will say so. The discovery claim is separable and stands or falls independently on the blind-labeling results. The stress cuts above already show we publish numbers that go negative (−0.89%/event).
Why frontier LLMs overweight the known. An LLM's factual recall scales with how often its pretraining data mentions an entity — Kandpal et al. (ICML 2023) measure it directly, and Mallen et al. (ACL 2023) show parametric memory is reliable mainly for popular entities. The head of that frequency curve — famous, big-float, heavily-covered names — is exactly the region markets have already priced. This report measures the consequence: LLM median leg float $8–13B vs Tuatara's $0.20B, basket overlap ~0.02.
Where alpha is made. Execution, with a human in the loop: Perold (1988, implementation shortfall); Frazzini, Israel & Moskowitz (2018, ~$1.7T of live orders — costs vary several-fold with execution style); Novy-Marx & Velikov (2016, implementation decides anomaly survival); Kaminski & Lo (2014, stop-rules help or hurt by regime); Cao, Jiang, Wang & Yang (2024, JFE — human+AI beats AI alone, strongest in small illiquid names); López de Prado (2018, meta-labeling). The data-engineering pipeline: Sculley et al. (2015, NeurIPS) and López de Prado (2018) — and this report's own §8/§12.3, where pipeline choices move the measured result more than the strategy's mean return. Model optimization headroom: every number here is the un-optimized, un-combined engine (§13.2–13.3); the headroom stays unclaimed until the §13.4 forward test measures it.
Discovery comes from the unknown. Swanson (1986, "Undiscovered Public Knowledge" — fish oil → Raynaud's, found by connection-mining two disconnected literatures) and Tshitoyan et al. (Nature, 2019) — word embeddings trained only on past materials-science papers prospectively predicted future thermoelectric discoveries, later confirmed by experiment. That is the epistemology of the Lawrence Berkeley National Lab patent lineage this platform descends from: in science and markets alike, everything new comes from the hidden structure of the data — never from the fat head everyone has already read. The backtest was designed to test exactly that capability in multi-trillion-dollar frontier models (Claude Fable 5 / Opus 5): measured, they surface the well-known and do not reach the hidden (§6.2/§11/§12.3), with the causal-mechanism test preregistered in §13.4(h).
Everything above is a retrospective study, and the caveats section says so plainly: float is measured as-of today, the attention window peeks two days past entry, announcement times are known only to the day. Those are lookahead problems, and no amount of re-running fixes them — they are baked into studying the past.
There is a separate body of evidence that has the opposite shape. Between February 2022 and January 2024, the Tuatara model generated thematic basket indices (AIBs) in response to live news catalysts and they were published publicly and timestamped at the moment of generation — the basket, its constituents and its entry prices posted to a public channel before anyone knew how it would resolve. The closing prices were filled in later, once the market had decided.
The mechanism is the same one this page is about. Those baskets were built by the same vector-association method that produces the Tuatara column in the tables above: a news catalyst goes in, and a basket of associated but non-obvious names comes out. The SIVB short published 10 March 2023 is the clearest case — the model named regional banks around the failure rather than the failure itself, and the basket resolved at +16.58% five days later while the S&P moved −0.61%. That is the hidden-leg thesis of this report, executed in public, two years before this study was written.
The record starts earlier than the archive above. The first widely-followed demonstration of the same engine was the "coronavirus" thematic smart basket, generated in early 2020 as the COVID-19 catalyst emerged — the vector-association method surfaced candidate vehicles related to the catalyst before they were consensus names, most famously NVAX, which it flagged as a coronavirus-associated candidate before the stock's 2020 run. Per the results set provided to the editor of the study above , the to-date return of that basket exceeds 3,000%. The same era produced the "Earthquake in Taiwan" basket, presented at a Morningstar conference — months before an actual Taiwan earthquake made it topical. The live publication of new thematic baskets has continued since in the public Discord channel #thematicbaskets-tier1, where anyone — or anyone's AI agent — can walk years of timestamped posts and check the entries against subsequent prices.
The live forward cohort now has two weeks of deals. Wire capture (running since 2026-07-26) recorded 110 deal announcements; collapsing repeat wires of the same transaction leaves 69 distinct deals. On this cohort we ran a target-based variant of the engine — the basket is the model's top hidden relateds to the acquired company (the deal's own parties excluded) — and then let a genetic optimizer (~1,700 parameter combinations, 40 generations) search the strategy's knobs: basket size, conviction floor, deal-size floor, holding period, and an hourly-MACD oversold/overbought weighting. Fitness rewards mean return and penalizes downside deviation only (Sortino-style) — a symmetric penalty would select against exactly the outsized winners that carry real books.
Champion parameters: top 2 legs per deal · conviction floor 0.53× the deal's top score · acquisitions ≥ $1.5B · hold 5 trading days · 3× overweight on legs entering deeply oversold (hourly MACD histogram z ≤ −1.27 over a 28-day window).
| View | 2 day (n=22) | 3 day (n=20) | 5 day (n=8) | 5d median | 5d hit rate |
|---|---|---|---|---|---|
| Raw — unhedged, unadulterated | +2.88% (t 2.60) | +4.18% (t 2.39) | +10.34% (t 2.40) | +5.03% | 87.5% |
| Excess vs S&P 500 | +1.56% (t 1.59) | +2.42% (t 1.50) | +6.40% (t 1.54) | +1.16% | 75.0% |
| Excess vs deal's sector ETF | +2.11% (t 2.06) | +3.28% (t 1.97) | +8.57% (t 2.01) | +5.86% | 75.0% |
| Combined with the FTA-10 short basket | +2.32% (t 1.97) | +3.70% (t 2.23) | +10.33% (t 2.56) | +5.11% | 87.5% |
| Deal (announced, size) | Basket | Basket return, 5 days |
|---|---|---|
| 07-29 · $7.7B | CMS −3.7% · FET +47.8% | +34.35% |
| 07-31 · $8.6B | LBRX | +19.76% |
| 07-31 · $1.5B | BNKK | +15.77% |
| 07-28 · $2.2B | AEP −3.5% · AGM +13.7% | +5.09% |
| 07-30 · $6.0B | VALU | +4.97% |
| 07-30 · $5.0B | PJT +3.1% · LAZ +4.5% | +3.73% |
| 07-28 · $2.1B | HSBC +2.3% · HBCYF +1.0% | +1.61% |
| 07-31 · $25.0B | HSBC | −2.53% |
Concentration is the mechanism, not a blemish. One deal contributes about a third of the 5-day mean, and its big leg (+47.8%) was produced by the rules — it entered deeply oversold and got the 3× overweight. That is the expected shape of strategy returns: Bessembinder (2018) shows 4.3% of US stocks account for all net equity wealth creation since 1926, and trend-following CTAs run 30–40% hit rates carried by a few outsized winners. What distinguishes this cut from a lottery ticket is the rest of the distribution: 7 of 8 five-day baskets positive, the worst at −2.5%, and — for the first time in this forward window — positive medians in every view, not only winner-carried means.
This page publishes results evidence: the study's event set, leg prices, recompute script, and the timestamped public basket record. It does not publish the machinery that produced them, and that line is drawn on purpose. Cymetica operates as a proprietary fund with internal information barriers ("Chinese walls"), the same posture the established quantitative funds take with their models and methodology — the industry standard is that algorithms, model internals and methodology are protected as trade secrets, and ours are no exception.