# M&A Ripple Effects — Tuatara Vectors vs Claude Fable 5 & Opus 5 — authored by Opus 5 ## Full Research Report (LLM-validation edition) Cymetica Research / EventTrader — published 2026-07-24 Authored by Claude Opus 5. GENERATED AND VALIDATED BY CLAUDE FABLE 5 AND CLAUDE OPUS 5. Note what that means for reading it: two of the three selectors graded here are the same model family that produced and checked this report, and both are scored BELOW our own Tuatara model. Every number is recomputable from the raw CSVs and the verify.py script published in §12 — do not take the attribution as assurance, run the script. UPDATED 2026-08-02 (later still): §15 records the PREREGISTRATION COMMITMENT — the forward-test protocol is published, timestamped, before the first forward event is scored; one primary endpoint declared in advance; negative results published at the same URL with the same prominence. UPDATED 2026-08-02 (§15.6): the protocol is now INSTANTIATED with fixed values — entry strictly after the captured announcement timestamp, a FIXED 5.0 attention threshold (not a per-sample median), one primary endpoint vs matched-random controls judged NET of costs at n=100, evaluated once. Locked before any of the 36 captured forward deals were scored. Status: research only — NOT investment advice. This document intentionally includes every caveat, negative result, and statistical weakness we know of, so that any third-party model or analyst can independently judge the claims. UPDATED 2026-07-24: §11 adds a full re-run of the LLM selector on Claude Opus 5 — including one §6.2 sub-claim that did NOT replicate. §12 publishes the RAW per-event and per-leg data (CSV + a recompute script) plus adversarial stress cuts. READ §12.2 BEFORE QUOTING THE HEADLINE RETURN — under the most obvious robustness cut it goes negative. UPDATED 2026-08-02: §8 adds caveats 9-10 — an attention-window lookahead (the d..d+2 pageview window extends past the t0-close entry) and announcement-date granularity (after-hours deals can price at a pre-announcement close) — both raised by a reader; §4 corrected to match. UPDATED 2026-08-02 (later): §13 adds study-design context (capability comparison; Tuatara un-optimized and uncombined — the floor, not the ceiling) and adopts the reader-contributed validation protocol as the forward-test spec. UPDATED 2026-08-02 (later still): §12.3 publishes the MEASURED variants for the §8.9/§8.10 lookaheads — timestamp-safe entry (worst case goes NEGATIVE: -0.89%/event) and fully-causal attention gating (~zero drift) — plus a no-float-filter cut and a repeat-pick artifact cut with popularity-concentration data across all three selectors. UPDATED 2026-08-02 (final): §14 adds the full literature stack with citations and links — LLM long-tail/fat-head evidence (Kandpal; Mallen), execution and pipeline alpha (Perold; Frazzini-Israel-Moskowitz; Novy-Marx-Velikov; Kaminski-Lo; Cao-Jiang-Wang-Yang; López de Prado; Sculley), and discovery-from-the-unknown (Swanson; Tshitoyan; Huberman- Regev) — so third-party AI systems can verify and validate every claim. --------------------------------------------------------------------------- 1. RESEARCH QUESTION --------------------------------------------------------------------------- When a large M&A deal is announced, which OTHER stocks move — and can a vector-association engine (Tuatara) find them better than a frontier general-purpose LLM (Anthropic's Claude Fable 5)? Theoretical basis: Huberman & Regev, "Contagious Speculation and a Cure for Cancer: A Nonevent that Made Stock Prices Soar," Journal of Finance 56(1), 2001. Their finding: ATTENTION, not information, moves prices — and it moves THIN, obscure names violently while barely moving mega-caps. Our design tests exactly that: post-announcement drift concentrated in small, low-float, high-attention "ripple" names. --------------------------------------------------------------------------- 2. UNIVERSE --------------------------------------------------------------------------- - 154 announced US-relevant M&A deals, announcement dates 2024-01-02 through 2026-07-22. Cohorts: 2024 n=31, 2025 n=45, 2026 n=78. - Deals curated from public announcement coverage; each record: acquirer, target, tickers (as of announcement), deal value USD, announcement date. - 2024-25 deals were restricted to $1B+ (verified value); 2026 cohort includes smaller deals; the headline strategy below filters to >$5B in all cohorts, so the comparison set is symmetric. --------------------------------------------------------------------------- 3. BASKET CONSTRUCTION (the two selectors) --------------------------------------------------------------------------- For every deal, two "hidden" (non-obvious) baskets were built from the SAME deal facts (acquirer, target, value, announcement description): A) TUATARA — our vector-association engine. The deal announcement text is the query; the engine returns associated equities ranked by association score; the acquirer's and target's own tickers are removed and stored separately as the "obvious" control basket; top 8 remaining legs kept, score-weighted. B) CLAUDE FABLE 5 — Anthropic's newest frontier model (Claude 5 family), prompted with identical deal facts and asked for 3-8 NON-OBVIOUS US-listed ripple stocks (suppliers, customers, competitors, sector peers, next-target candidates), explicitly excluding acquirer/target. Picks are free-form (whole market), equal-weighted. 154/154 deals returned valid baskets. Tradability filter (both selectors, identical): a leg with no Yahoo Finance daily price history is dropped at pricing time (weight renormalized over surviving legs). No other symbol screening — both selectors face the same market. CONTROL basket: acquirer + target tickers ("obvious"). BENCHMARK: SPY. --------------------------------------------------------------------------- 4. PRICING METHODOLOGY --------------------------------------------------------------------------- - Prices: Yahoo Finance daily adjusted closes, ~3 years of history. - Trading calendar: SPY's own daily series. - t0 = first trading CLOSE on/after the announcement date (entry at t0 close). CORRECTION (2026-08-02): an earlier revision called this entry "causal — never enters before the announcement is public." That is only true when the announcement lands during or before market hours. Deal announcements are DATE-granular in our source; a deal announced after the close on date d still gets t0 = d's close, which printed BEFORE the news. For such deals the measured 1-day return captures the announcement-day pop from an unrealizable pre-news fill — a lookahead that INFLATES returns. See §8.10. - N-day return = weight-renormalized basket return from t0 close to the Nth trading close after t0. Excess return = basket return minus SPY over the identical window. - Baskets require >= 3 priced legs to count in aggregate statistics (degenerate 1-2 leg baskets excluded — early unfiltered runs showed single-microcap baskets producing +200% artifacts). --------------------------------------------------------------------------- 5. STRATEGY FILTERS (the "flagship" configuration) --------------------------------------------------------------------------- Applied on top of the hidden baskets, long-only, hold = 1 trading day: - Deal size > $5B (117 of 154 deals qualify). - Leg float < $250M USD (float = floatShares x price from Yahoo; legs with UNKNOWN float are excluded, not assumed small). - Leg price <= $20 at t0 close. - ATTENTION GATE: Wikipedia pageviews of the TARGET company's article. attention ratio = (avg daily views over announcement day..+2 + 1) / (avg daily views over day -30..-4 + 1). Known for 114/117 deals; median ratio 5.0x. "HIGH attention" = above the median, "low" = below. - A MACD overbought gate (EMA12-EMA26 vs 9-signal at t0) was ALSO tested and REJECTED — it subtracted return in every cohort (see §8). --------------------------------------------------------------------------- 6. HEADLINE RESULTS --------------------------------------------------------------------------- 6.1 Flagship configuration (Tuatara hidden legs, >$5B, float<$250M, <=$20, long 1 trading day, excess vs SPY): HIGH attention: +3.77% per event | 61.4% hit rate | t = 1.5 | n = 44 by cohort: 2024 +1.27% | 2025 +2.78% (80% hit) | 2026 +6.25% low attention: -0.01% per event | 42.9% hit rate (2026 low: t = -2.11) ALL of the strategy's return comes from the high-attention half — exactly the Huberman-Regev prediction. Median event return is near zero; the mean is carried by right-tail "poppers" (this is a tail strategy). 6.2 Selector head-to-head, same flagship configuration on CLAUDE baskets: Claude flagship: -1.26% per event | 33% hit | t = -1.6 | n = 15 ...and it barely fires: only ~17 qualifying low-float legs across 117 big deals, vs ~44 tradable events for Tuatara. 6.2a COMPOUNDED SEQUENTIALLY, $1,000,000 REINVESTED (same events as above, taken in date order, full balance rolled into each successive event; raw basket return entry-to-exit, no SPY subtraction): Selector Events Compounded $1,000,000 becomes Tuatara (high-attention half) 44 +202.99% $3,029,893 Claude Fable 5 15 -20.18% $798,243 Claude Opus 5 19 -18.92% $810,832 READ THIS BEFORE QUOTING +202.99%. It is an arithmetic roll-up of the published prices, NOT a track record and NOT a live or paper equity curve. It assumes every fill at the printed entry and exit, full reinvestment with no capital constraint, and no position sizing. It is pre-cost like everything else in this report. - Exposure is 53 leg-days (holds are 1 trading day; three at 3 days, one at 4) spread over 2024-01-11 to 2026-07-17. There is no meaningful annualised figure and none is claimed. - ONE basket (Roku, entered 2026-06-15, exited the next day) is 51% of Tuatara's total profit. Excluding it, the other 43 baskets compound to a small fraction of the headline. - Sequential compounding MAGNIFIES the right-tail concentration documented in 12.2. Read it next to the +0.32% median event. Basis note: Tuatara's row is the high-attention half (n=44), matching 6.1. The Fable 5 and Opus 5 rows are all qualifying events for those selectors (n=15, n=19), matching 6.2 — their high-attention subsets are n=7 and n=9, too small to report separately. 6.2b WHY THE CLAUDE RUNS COMPOUND SO MUCH WORSE. The compounded spread (+202.99% vs -20.18% / -18.92%) is far wider than the per-event spread (+3.77% vs -1.26% / -0.95%). That gap is a property of compounding, not evidence of a bigger edge. Raw basket returns: Selector n mean median max min sd skew Tuatara 44 +3.53% +0.01% +79.28% -22.38% 16.38 +3.13 Fable 5 15 -1.43% +0.18% +2.86% -8.29% 3.59 -0.73 Opus 5 19 -0.98% -1.07% +14.89% -8.70% 4.99 +1.28 IT IS NOT THAT THE CLAUDE PICKS LOSE MONEY — THEY NEVER PRODUCE A BIG WINNER. Fable 5's median event is +0.18%, the HIGHEST of the three: its typical trade beats Tuatara's typical trade (+0.01%). What it never does is hit. Its best result in 15 events is +2.86%, against five Tuatara events over +10% and three over +25%. Compounding is multiplicative and asymmetric — losses are floored at -100%, gains are unbounded — so a book of small wins and small losses grinds down while one large winner carries an entire book: Selector compounded best event without that one event Tuatara +202.99% +79.28% +69.01% Fable 5 -20.18% +2.86% -22.39% Opus 5 -18.92% +14.89% -29.43% Opus 5 has the SAME single-event dependency as Tuatara — strip its one +14.89% and it falls to -29.43%. It simply drew a much smaller winner. THE MECHANISM IS FLOAT (same finding as section 4). Every leg in all three books is under $250M float — that is the screen. But the distribution WITHIN that ceiling differs by roughly 8x: Median leg float: Tuatara $17.2M (253 legs) Fable 5 $118.5M (17 legs) Opus 5 $137.3M (22 legs) Float bucket (all selectors pooled, per leg): under $10M n=110 mean +2.49% median +0.00% max +102.90% $10M - $100M n= 84 mean +5.63% median +0.00% max +519.05% $100M - $250M n= 98 mean +0.26% median +0.00% max +17.94% The Claude median leg sits in the $100-250M bucket, which returns +0.26% — indistinguishable from nothing. Tuatara's median leg sits in the buckets that produce the +102% and +519% prints. The frontier models are not picking bad companies; they are picking companies too large to move 80% overnight. TWO CAVEATS THAT CUT BOTH WAYS. First, the Claude samples are tiny (n=15 and n=19 events; 17 and 22 qualifying legs), so their compounded figures are barely more than noise — and that scarcity is itself the headline finding: frontier models rarely name a company small enough to clear the screen at all. Second, this is not a clean win for Tuatara. "Produces large winners in illiquid names" and "would face crushing spread and impact costs" are the same sentence. The +519% leg is a $0.021 stock. The selector divergence is the durable result; the return spread is not. Why (the float profile of the two selectors' picks): Median hidden-leg float: Tuatara $0.20B vs Claude $12.96B Legs under $250M float: Tuatara 52.5% vs Claude 3.6% Claude names liquid mega-cap peers (e.g. for Paramount-Skydance/WBD it picked FOXA, CMCSA, NFLX, AMC, CNK) — reasonable names, but already efficiently priced. The attention gate INVERTS on Claude's picks (high attention on an already-famous large cap = news already priced): at a $2B float ceiling, Claude HIGH-attention = -1.19%/event (t = -2.4) vs low-attention +0.75%. 6.3 Unfiltered short-horizon aggregates (ALL deals, >=3 priced legs, excess vs SPY, per-event): horizon | Tuatara avg (t) med hit | Claude avg (t) med hit 1 day | +0.93% (1.50) -0.05% 47.2% | -0.09% (-0.77) -0.16% 44.4% 2 day | +1.17% (1.77) -0.13% 48.8% | +0.13% (0.66) +0.23% 54.7% 3 day | +0.43% (0.60) -0.42% 44.4% | +0.26% (1.13) +0.39% 56.4% (Tuatara n=126-127, Claude n=149-151 per horizon) Note the shape difference: Tuatara's means are larger but tail-driven (negative medians); Claude's returns are small but more uniform. The edge being tested is specifically the low-float tail — which only Tuatara's picks can express. Longer horizons (Tuatara hidden, calendar-window engine): 7d -0.07%, 30d -0.27%, 90d +2.67% (t=1.0) — no reliable edge beyond ~1-3 days. The obvious (acquirer+target) control underperforms SPY at every horizon tested. --------------------------------------------------------------------------- 7. NEGATIVE / FALSIFIED RESULTS (kept deliberately) --------------------------------------------------------------------------- These hypotheses were tested and FAILED. We publish them because a report you can hand to an adversarial validator must show the search path, not just the survivor: a) "Bigger deals -> bigger ripple returns": FALSE. Mega-deals (>= $10B) hidden baskets were NEGATIVE from day 1. Returns did not rise with deal size. b) Long/short (long $5-10B deals / short >=$10B deals): failed out-of-sample. Derived on 2026 data, it was negative at 1-3 days on 2024-25 (-4.03% at 21 days, t = -2.2, wrong sign). Rejected. c) Volume-based salience (abnormal deal-ticker volume) does NOT separate returns — it measures arbitrage flow, not public attention. Only the Wikipedia-pageview attention ratio separated. d) MACD overbought gate: subtracted return in EVERY cohort (consistent with post-announcement continuation/underreaction — momentum should not be filtered out here). Rejected. e) The same Wikipedia attention gate does NOT transfer to general market-news headlines (non-M&A): on those, attention ratios compress to ~1.0 (famous subjects have huge stable baselines) and the split carries no signal. Attention works as a SHOCK detector on obscure names, not a LEVEL filter on famous ones. --------------------------------------------------------------------------- 8. KNOWN WEAKNESSES & CAVEATS (read before trusting §6) --------------------------------------------------------------------------- 1. STATISTICAL STRENGTH: the flagship result is t ≈ 1.5 (n = 44). That is NOT conventionally significant (t >= 2). It is "promising and consistent across three yearly cohorts," not "proven." 2. PRE-COST: returns exclude commissions, borrow, and — critically — spread/impact. The strategy trades LOW-FLOAT names where spreads are wide; realized returns will be materially lower. Median event return ~0 means costs eat the median trade; the strategy lives on the tail. 3. FLOAT LOOKAHEAD: float values are CURRENT (as-of 2026-07), not as-of the announcement date. Floats change (buybacks, offerings, lockups). This is a real lookahead bias in the float filter. 4. ATTENTION THRESHOLD IN-SAMPLE: the HIGH/low split uses the sample median (5.0x). A live implementation must fix the threshold ex-ante. The attention window also includes the announcement day itself — coincident, not strictly pre-entry. 5. SURVIVORSHIP: Yahoo removes full price history for delisted tickers. Targets of COMPLETED deals disappear; ~108/845 symbols in the 2024-25 expansion were unrecoverable. This mainly damages the "obvious" control basket but is a structural bias in all legs. 6. HINDSIGHT CONTAMINATION (Claude leg): Fable 5's training data may include what actually happened after 2024-25 announcements. If anything this should INFLATE Claude's measured performance; its null result is therefore conservative. Tuatara's corpus is also current-day, so neither selector is strictly point-in-time. 7. SELECTION OF THE FLAGSHIP CONFIG: the low-float/price/attention combination was reached through sequential hypothesis testing on the same dataset (see §7 for everything that failed). Multiple-comparison risk is real; treat the config as a hypothesis for forward validation, not a validated system. 8. WEIGHTING ASYMMETRY: Tuatara legs are association-score weighted; Claude legs are equal-weighted (the model emits no scores). Both were also compared under identical renormalized pricing; the float-profile gap (§6.2) dominates any weighting effect. 9. ATTENTION WINDOW LOOKAHEAD (added 2026-08-02, raised by a reader): the attention ratio's numerator averages pageviews over announcement day..+2, but entry is at the t0 close — so the HIGH/low classification uses pageview data from one and two days AFTER entry. Caveat 4's earlier wording ("coincident, not strictly pre-entry") understated this: d+1 and d+2 views are future information at entry time, full stop. A live implementation must gate on data available at entry (e.g. views through d only, or intraday views before the close); we have NOT measured that variant, and the HIGH-half returns in §6 cannot be claimed as achievable until someone does. MEASURED 2026-08-02 (§12.3 variant C): entering strictly AFTER the full d..d+2 window leaves ~zero drift (+0.13%/event, t=0.14, n=52). The gate as parameterized is a DIAGNOSTIC of where returns concentrated, not a live entry signal. 10. ANNOUNCEMENT TIMESTAMP GRANULARITY (added 2026-08-02, raised by a reader): deal announcement times are known only to the DAY. Deals announced after market close on date d get t0 = d's close — a price printed before the news existed. The measured next-day return then includes the announcement pop from a fill no trader could have had, which inflates measured returns. We have not partitioned the sample by announcement time (the data source does not carry it); an independent replication should use timestamped announcements and enter strictly at the first post-announcement price. BOUNDED 2026-08-02 (§12.3 variant B): assuming the WORST case — every trading-day deal announced after that day's close, entry at the first close strictly after the announcement date — flips the flagship HIGH-attention arm to -0.89%/event (t=-1.11, n=52). On the current universe the truth lies between +3.03% and -0.89% and cannot be allocated retrospectively; forward events are timestamp-exact (deals captured live from the wire since 2026-07-28). --------------------------------------------------------------------------- 9. REPRODUCTION GUIDE (for an independent implementation) --------------------------------------------------------------------------- 1. Collect announced M&A deals (date, acquirer, target, value) for 2024-2026; restrict to >$5B. 2. For each deal, generate candidate ripple equities two ways: (a) any association/vector engine over news text; (b) prompt an LLM with the deal facts for 3-8 non-obvious US-listed ripple names excluding acquirer/target. 3. Price daily adjusted closes (any survivorship-aware source is BETTER than Yahoo); t0 = first close on/after announcement; hold 1 trading day; excess vs SPY. 4. Filter legs: float < $250M, price <= $20. 5. Attention: Wikipedia REST pageviews API, target company article, ratio of announcement-window (d..d+2) to baseline (d-30..d-4) daily averages; split at a FIXED ex-ante threshold (we measured median 5x). 6. Compare HIGH-attention vs low-attention event returns; compare the two selectors' float profiles and qualifying-event counts. Expected qualitative findings if our result is real: (i) returns concentrate in the high-attention, low-float half; (ii) the LLM's picks skew to mega-caps and cannot express the strategy; (iii) median event return ~0 with a positive right tail. --------------------------------------------------------------------------- 10. RELATED PUBLISHED WORK --------------------------------------------------------------------------- - Huberman & Regev (2001), J. Finance 56(1): attention-driven repricing of thin biotech names on a "nonevent" (EntreMed) — +330% day move on months-old news, contagion to peer names on 50x volume, mega-cap partner (BMY) barely moved (+3.12%). - Our earlier benchmark (2026-06-30): Tuatara vs Claude vs MiniLM on general US-equity news baskets — cymetica.com/static/research/ comparison/index.html. NOTE: we have since identified that ALL trades in those runs share a single entry session (2026-05-01) due to the price table's calendar coverage; treat those per-trade statistics as a cross-sectional selector RANKING (Tuatara > Combined > Claude > MiniLM held across every variant), not a time-diversified return estimate. A rerun with true per-headline entry sessions is in progress and will be published the same way. --------------------------------------------------------------------------- 11. ADDENDUM — MODEL-UPGRADE REPLICATION (Claude Opus 5), 2026-07-24 --------------------------------------------------------------------------- The obvious objection to §6.2 is "your LLM leg was just a weak model." We re-ran the ENTIRE LLM selector on Anthropic's Claude Opus 5 — same 154 deals, byte-identical prompt, same pricing, same tradability rule, same filters. Only the model changed. 11.1 Coverage: 152/154 deals returned a basket (2 empty), 1,115 hidden legs, 851 distinct symbols. Price ingest: 755 priced, 97 Yahoo MISS (11.4%, vs 14.3% on the earlier run — no data-coverage handicap either way). 11.2 Short-horizon aggregates (>=3 priced legs, excess vs SPY, per event): horizon | Fable 5 | OPUS 5 | Tuatara 1 day | -0.09% (t -0.77) n=151 | -0.10% (t -0.72) n=152| +0.93% (t 1.50) 2 day | +0.13% (t 0.66) n=150 | +0.20% (t 1.00) n=151| +1.17% (t 1.77) 3 day | +0.26% (t 1.13) n=149 | +0.25% (t 1.08) n=150| +0.43% (t 0.60) 11.3 Flagship configuration (>$5B, float<$250M, price<=$20, long 1 day): Tuatara HIGH attention: +3.77%/event | 61.4% hit | t = 1.5 | n = 44 Claude Fable 5: -1.26%/event | 33.3% hit | t = -1.6 | n = 15 OPUS 5: -0.95%/event | 26.3% hit | t = -0.8 | n = 19 (Opus 5 HIGH attention -0.09% n=9 | low -2.24% n=9) 11.4 Why it does not move: the float profile is a property of how an LLM picks, not of how good the LLM is. Median hidden-leg float: Tuatara $0.20B | Fable 5 $12.96B | Opus 5 $8.21B Legs under $250M float: Tuatara 52.5% | Fable 5 3.6% | Opus 5 4.1% Opus 5 names somewhat smaller companies than the earlier run, but still liquid large caps. It qualifies for 19 of 117 big-deal events instead of 15 — it still cannot express a low-float strategy. 11.5 Basket overlap (mean per-deal Jaccard on hidden legs): Opus 5 vs Fable 5 0.401 | Opus 5 vs Tuatara 0.018 | Fable 5 vs Tuatara 0.019. The two LLM runs agree with each other ~20x more than either agrees with the vector engine. The selector gap is STRUCTURAL (a class-of- method difference), not a model-capability difference. 11.6 A §6.2 CLAIM THAT DID NOT REPLICATE — the "attention gate inverts on LLM picks" result. At the $2B float ceiling the earlier run showed HIGH attention -1.19% (t = -2.4) vs low +0.75%. On Opus 5 the same cut shows HIGH -0.15% (t = -0.19, n=24) vs low -0.31% (t = -0.49, n=25) — no separation in either direction. Treat the inversion as noise at n~16, not as an established effect. The NULL result for LLM picks (§6.2 headline) does replicate; the inversion does not. 11.7 HONEST LIMITS OF THIS ADDENDUM: a) n=19 vs n=15 qualifying events. Opus 5 "beating" the earlier run by +0.31%/event is statistically indistinguishable from zero. The finding is "no material difference," not "Opus 5 is better." b) Provenance hardening: the basket table now stores the concrete model id per row (Opus 5 rows carry claude-opus-5). The Fable 5 rows predate that column, so their model id lives in the run record rather than in the dataset itself — a reproducibility gap we closed rather than one that affects the numbers. c) Hindsight contamination (§8.6) applies at least as strongly to a newer model with a later training cutoff. The null result stays conservative. d) Every §8 caveat — pre-cost, float lookahead, in-sample thresholds, survivorship, multiple comparisons — applies unchanged here. Reproduce: scripts/ma_backtest/score_deals_claude.py (MA_CLAUDE_TABLE + ENTITLEMENT_MODEL_MAP pin the destination table and model), run_short_horizon.py / run_lowfloat_long.py (MA_BASKET_TABLE), compare_selectors_report.py (MA_SELECTORS side-by-side). --------------------------------------------------------------------------- 12. RAW EVENT DATA + ADVERSARIAL STRESS CUTS (added 2026-07-24) --------------------------------------------------------------------------- A reader correctly pointed out that "we asked a frontier model and it did not object" is NOT verification — a model that never saw the price data cannot audit the result. The only real answer is the raw data. Published: /static/research/ma-ripples/flagship_events.csv (130 events, all 3 selectors) /static/research/ma-ripples/flagship_legs.csv (292 legs) /static/research/ma-ripples/verify.py (recompute script, no deps) flagship_events.csv: selector, deal_id, acquirer, target, deal value, announcement date, entry date (t0), exit date (t1), Wikipedia attention ratio, attention bucket, n_legs, SPY entry/exit, basket return, SPY return, excess. flagship_legs.csv: every leg's ticker, raw + renormalized weight, float, entry price, exit price, leg return. Prices at FULL precision — some legs trade at $0.0004, where rounding alone changes the answer. `python3 verify.py` (in that directory) rebuilds every basket return from the raw leg prices, re-derives excess vs SPY from the published SPY prices, and regroups the headline. Worst rebuild mismatch: 0.00005 pp. 12.1 STRESS CUTS ON THE HEADLINE (Tuatara, HIGH attention, flagship). Run them yourself from the CSVs — this is us running the attacks first: cut n avg med t hit as published 44 +3.77% +0.32% +1.50 61.4% require >= 3 priced legs 28 +6.76% +0.32% +1.81 60.7% drop baskets w/ any leg < $0.01 22 +5.16% +0.13% +1.09 54.5% drop baskets w/ any leg < $1.00 10 -1.55% +0.29% -0.61 70.0% drop the single best event 43 +2.00% +0.32% +1.09 60.5% (that event is Roku, +79.87% excess) 12.2 WHAT THOSE CUTS MEAN — read this before quoting +3.77%: a) THE EDGE IS CONCENTRATED IN SUB-$1 STOCKS. Restrict to baskets whose legs all trade above $1.00 and the mean goes NEGATIVE (-1.55%, n=10). Every headline number here depends on names priced in cents. b) IT IS A THREE-EVENT RESULT. The top 3 events contribute 175.7pp of a 165.8pp total — i.e. more than 100%, so the other 41 events are net NEGATIVE in aggregate. Removing just the best one cuts the mean nearly in half and drops t to 1.09. c) THE MEDIAN IS +0.32%, not +3.77%. The typical event does approximately nothing. This was always stated (§6.1) but the raw data makes it concrete. d) Therefore §8.2 (pre-cost) is not a footnote, it is decisive: a strategy whose return lives in sub-$1, sub-$250M-float names, carried by three events, will not survive realistic spread and impact. We do not trade this configuration and nobody should on this evidence. The SELECTOR comparison (§6.2/§11) is more robust than the return claim: it rests on the float profile and basket overlap of ~1,000 legs per selector, not on a handful of tail events. "Tuatara finds different names than an LLM does" survives all of the cuts above. "Those names make +3.77%" does not survive cut (a). 12.3 MEASURED VARIANTS FOR THE §8.9/§8.10 LOOKAHEADS + TWO MORE CUTS (added 2026-08-02, run in response to the same reader exchange) §8.9 and §8.10 name two lookaheads and say the corrected variants were unmeasured. We measured them. Universe is the CURRENT deal set (128 deals >$5B vs 117 at publication — the live wire keeps adding deals), so n differs from the published 44. Tuatara flagship legs, HIGH attention, 1-trading-day hold, excess vs SPY: variant n avg t A published entry (first close on/after ann.) 52 +3.03% +1.41 B timestamp-safe (first close STRICTLY after) 52 -0.89% -1.11 C fully-causal (entry after the d..d+2 window) 52 +0.13% +0.14 A without the float filter 59 +2.02% +1.42 A no float filter AND all legs >= $1.00 20 +0.70% +1.77 a) Variant B is the worst-case bound for §8.10 (it assumes EVERY trading-day deal was announced post-close; weekend/holiday deals are unaffected since their first close already follows the news). It forfeits the entire day-1 move: the flagship arm goes NEGATIVE. The realizable figure lies between A and B and cannot be allocated without per-deal timestamps. b) Variant C closes §8.9: with the attention window fully pre-entry, the remaining drift is ~zero — consistent with §6.3's decay by day 3. The gate diagnoses WHERE returns concentrated; it is not, as parameterized, a tradable signal. c) The attention SPLIT does not depend on the float-lookahead filter (§8.3): with no float filter at all (price <= $20 at entry is the only leg gate, point-in-time) HIGH +2.02% vs low +0.87%; excluding sub-$1 legs as well, HIGH +0.70% (t=1.77) vs low +0.20%. Small, not significant — but not an artifact of the contaminated input. A further cut we ran on ourselves, unprompted. Repeated-pick concentration per selector (hidden legs, all 154 deals): Tuatara has 11 symbols appearing in >= 5 deals' baskets, covering 27.1% of its legs (top repeats: RSVR/BACQ x40, DFPH/AFJK/HYAC/MLAC x39) — vs 3.8% (Fable 5) and 2.4% (Opus 5), whose top repeats are TMUS/CMCSA/FI/PYPL-class mega-caps. Two readings, both published: the LLMs' repeats make popularity weighting visible (famous names, median pick float $8-13B vs Tuatara's $0.20B — first controlled data on §13.4(h)'s popularity hypothesis, not yet its proof); and Tuatara's repeats are a candidate "generic microcap attractor" artifact — names associated with M&A language generally rather than with the specific deal. Dropping every repeat-picked leg and rerunning the flagship: n avg med t hit as published (current univ.) 52 +3.03% +0.32% +1.41 59.6% drop repeat-picked legs 33 +3.78% +0.38% +1.18 66.7% (low-attention arm 34 -0.08% -0.65% -0.06 38.2%) The attention split survives its strongest artifact candidate: the mean holds, the hit rate rises, and the low-attention control stays at zero. Validation of the remaining open claims proceeds under the §13.4 protocol (collecting, timestamp-exact, since 2026-07-28). --------------------------------------------------------------------------- 13. STUDY DESIGN CONTEXT + THE VALIDATION BAR (added 2026-08-02, from a public reader exchange) --------------------------------------------------------------------------- This section was added after a public back-and-forth with a reader whose critique (delivered via two frontier LLMs) sharpened the report. Three pieces of design context first, then the validation protocol that emerged. 13.1 WHAT THIS COMPARISON WAS. The study is a CAPABILITY comparison: a domain-specific vector engine against frontier LLMs from multi-trillion-dollar labs, on the LLMs' best terms (prompted directly for non-obvious ripple names). The durable finding is the selection asymmetry (§6.2/§11): the frontier models never surface the low-float universe at all. It is NOT a strategy pitch — §12.2 already says nobody should trade the paper configuration. 13.2 TUATARA RAN UN-OPTIMIZED — THIS IS THE FLOOR, NOT THE CEILING. Raw association output from the corpus: no hyperparameter tuning for this task, no per-deal feature engineering, no ensemble, no execution layer. The only fitted element is the flagship filter set, disclosed as multiple-testing risk in §8.7. Measured at daily-close granularity, standalone. Every one of those choices understates a production configuration. 13.3 THE SELECTORS WERE DELIBERATELY NOT COMBINED — AND WHAT THE PUBLISHED "COMBINED" RESULT DOES AND DOES NOT SHOW. The M&A legs are pure-Tuatara vs pure-LLM for clean attribution. In our separately published equity headline study, the Combined (Tuatara ∪ Claude) book posts the best risk-adjusted result on the page (per-trade Sharpe 0.79 vs 0.70 Tuatara-alone and 0.27 Claude-alone, with the highest win rate): cymetica.com/static/research/comparison/index.html. PRECISION (2026-08-02, after a reader pushed on this): that Combined book is a mechanical UNION of the two selectors' baskets — broader coverage plus diversification can lift win rate and per-trade Sharpe without demonstrating model synergy. It is evidence the books are complementary, NOT a test of the actual production architecture (an LLM independently ranking, vetting and sizing vector-surfaced candidates), which remains unbuilt-in-public and unmeasured. Two further disclosed asymmetries in that study: Tuatara's query was the keyword-condensed article body while Claude's was the title only, and its Sharpe is per-trade (mean/stdev), not an equity-curve Sharpe. 13.4 THE VALIDATION BAR (reader-contributed, adopted). The reader's protocol for what WOULD establish the stronger, executable-alpha claim — which we adopt as the §10 forward-test spec: a) Timestamp-safe entries: exact announcement timestamps; enter strictly at the first post-announcement price (fixes §8.10). b) As-of-date float, not current float (fixes §8.3). c) Information sets strictly pre-entry, including the attention gate (fixes §8.9) and a FIXED ex-ante threshold (fixes §8.4). d) Execution modeled from quotes/spreads/volume with explicit capacity limits — or better, real logged fills. e) Matched random microcap controls (same float/price screens, random selection) to separate "low-float names move a lot" from "THESE low-float names move." f) Human-in-the-loop discretion PRE-SPECIFIED: decision rules, information sets and fill policies fixed before events, every override logged — so execution judgment is testable, not hindsight. g) Blind validation of alleged hidden relationships: does a reviewer who cannot see returns judge the surfaced connection real? h) INPUT PARITY for any selector-mechanism claim: both systems get the same point-in-time corpus, full documents, candidate universe and top-K ranking task. Until then, "compression / alignment / popularity weighting" as the CAUSE of the LLM gap is a hypothesis consistent with the data, not a demonstrated mechanism. i) Relationship-correctness labeling against BOTH matched-random and retrieval baselines — "different and more obscure" is not automatically "more accurate"; accuracy has to be scored directly. NOTE on 13.2, same reader: "un-optimized" does not make the measured result a floor — optimization can improve, do nothing, or overfit. Correct. The floor language properly attaches only to the ABSENT components (no execution layer, no combination), not to tuning. Until a run like that exists, the published claim is and remains: Tuatara is a differentiated long-tail DISCOVERY engine. The executable-alpha claim is open — in both directions (§8.2 costs could kill it; §13.3 headroom could enlarge it). --------------------------------------------------------------------------- 14. LITERATURE: THE FAT HEAD vs THE LONG TAIL, AND WHERE ALPHA IS MADE (added 2026-08-02 — full citations so third-party AI systems and agents can verify and validate every claim independently) --------------------------------------------------------------------------- 14.1 WHY FRONTIER LLMs OVERWEIGHT THE FAT HEAD (the known). An LLM's factual recall of an entity or relationship scales with how many pretraining documents mention it — measured, not asserted: - Kandpal, Deng, Roberts, Wallace & Raffel (2023), "Large Language Models Struggle to Learn Long-Tail Knowledge," ICML 2023. https://proceedings.mlr.press/v202/kandpal23a.html - Mallen, Asai, Zhong, Das, Khashabi & Hajishirzi (2023), "When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories," ACL 2023. https://aclanthology.org/2023.acl-long.546/ Consequence for markets: the head of the frequency curve — famous, heavily-covered, big-float names — is precisely the region already priced by analyst coverage and index flows. A model whose recall is frequency-weighted has structural trouble even REPRESENTING the long tail, which is where hidden relationships and (if it exists) edge live. This report's §6.2/§11/§12.3 measure that signature directly: LLM median leg float $8-13B vs Tuatara $0.20B; basket overlap (Jaccard) ~0.02 across ~1,000 legs; LLM repeated picks are mega-caps. 14.2 WHERE ALPHA IS MADE. A) HUMAN-IN-THE-LOOP TRADE EXECUTION. - Perold (1988), "The Implementation Shortfall: Paper Versus Reality," Journal of Portfolio Management 14(3) — the paper-vs-real gap IS execution. - Frazzini, Israel & Moskowitz (2018), "Trading Costs," SSRN 3229719 (~$1.7T of live executed orders) — realized costs vary several-fold with execution style; execution skill, not the signal alone, decides the net result. - Novy-Marx & Velikov (2016), "A Taxonomy of Anomalies and Their Trading Costs," Review of Financial Studies 29(1) — implementation choices decide whether a published anomaly survives at all. - Kaminski & Lo (2014), "When Do Stop-Loss Rules Stop Losses?," Journal of Financial Markets 18 — mechanical stops help or destroy value by regime; day-high/day-low context judgment decides. - Cao, Jiang, Wang & Yang (2024), "From Man vs. Machine to Man + Machine: The Art and AI of Stock Analyses," Journal of Financial Economics 160 — human+AI beats AI alone, with the human edge concentrated in small, illiquid, thinly-covered firms. - López de Prado (2018), Advances in Financial Machine Learning, Wiley — meta-labeling: a judgment layer on sizing/veto over a primary signal measurably improves precision. B) THE DATA ENGINEERING PIPELINE. - Sculley et al. (2015), "Hidden Technical Debt in Machine Learning Systems," NeurIPS 2015 — the model is a small box in a system dominated by data plumbing. - López de Prado (2018), op. cit. — most quant failure happens before any model runs: curation, point-in-time correctness, labeling. - This report is its own demonstration: §8.3/§8.9/§8.10 pipeline choices (float vintage, timestamps, attention window) each move the measured result by more than the strategy's mean return (§12.3). C) MODEL OPTIMIZATION HEADROOM (Tuatara). - Every number in this report is the UN-optimized, UN-combined engine (§13.2-13.3): raw association output, no task tuning, no ensemble, no execution layer. The optimization/combination headroom is real and deliberately UNCLAIMED until the §13.4 forward test measures it. 14.3 DISCOVERY COMES FROM THE UNKNOWN, NOT THE KNOWN. Scientific discoveries are made not from the known — the known is the literature everyone already read — but from hidden relationships latent in data: - Swanson (1986), "Undiscovered Public Knowledge," Library Quarterly 56(2) — the fish-oil -> blood-viscosity -> Raynaud's connection sat unnoticed across two disconnected literatures until literature-based connection mining surfaced it; later clinically supported. - Tshitoyan et al. (2019), "Unsupervised word embeddings capture latent knowledge from materials science literature," Nature 571, 95-98 — embeddings trained ONLY on past papers prospectively predicted future thermoelectric discoveries years before humans made them; hidden relationships extracted from vector geometry, later confirmed by experiment. - Huberman & Regev (2001), §1 — attention, not information, moves prices; the market analogue of the same asymmetry. This is the epistemology of the Lawrence Berkeley National Lab patent lineage this platform descends from (X-Mine 2001 -> LBNL/DOE -> SeeqPod -> Vectorspace AI -> Cymetica): in science and in markets alike, everything new comes from the hidden structure of the data — never from the fat head everyone has already read. 14.4 WHAT THE BACKTEST WAS DESIGNED TO TEST. Whether multi-trillion-dollar frontier models (Claude Fable 5 and Opus 5) can extract hidden relationships/connections from data — financial language here, the same class of task as biological language modeling and novel-discovery mining — when prompted on their best terms. Measured answer (§6.2/§11/§12.3): they surface the well-known; they do not reach the hidden. Whether that gap's CAUSE is the frequency weighting of 14.1 is the input-parity test of §13.4(h); the capability gap itself is measured. --------------------------------------------------------------------------- 15. PREREGISTRATION COMMITMENT (added 2026-08-02) --------------------------------------------------------------------------- The executable-alpha claim (§13.4) is open and will be settled by a prospective forward test, not by any further reanalysis of the historical sample. To make that test informative rather than persuasive, we commit to the following BEFORE any forward results are known: 15.1 PUBLISH-BEFORE-RESULTS. The full forward-test protocol — the §13.4(a)-(i) spec instantiated with concrete, non-negotiable values — is published at this URL, timestamped, BEFORE the first forward event is scored. Nothing in the protocol is amended after results start arriving. If it is amended before the start, the prior version stays published alongside, with the reason. 15.2 WHAT THE PROTOCOL FIXES IN ADVANCE. Entry rule and exact fill convention; holding period; the ex-ante attention threshold as a FIXED number (§13.4c); the candidate universe and screens as of the event date; the benchmark; the matched-random control construction (§13.4e); position sizing and any human-override rules with the logging format (§13.4f); the primary endpoint and the sample size / stopping rule; and the blind relationship-labeling procedure and rater instructions (§13.4g). 15.3 PRIMARY ENDPOINT DECLARED IN ADVANCE. One primary endpoint, stated before the start. Every other cut published alongside it is explicitly secondary/exploratory and will be labeled as such, whatever it shows. We will not promote a secondary cut to the headline because the primary one disappointed. 15.4 NEGATIVE RESULTS ARE PUBLISHED. The forward result is published whichever way it falls, at the same URL and with the same prominence as this report — including the case where it contradicts §6/§11 or kills the executable-alpha claim outright. A protocol that can only produce good news is not a test, and the §12.3 variants already show we publish cuts that go negative (-0.89%/event). 15.5 WHAT WOULD FALSIFY US. If forward high-attention legs do not beat matched random controls on the declared primary endpoint, the executable-alpha claim is rejected and this report will say so. The DISCOVERY claim (a differentiated long-tail universe, §13.4 close) is separable and survives or falls on the §13.4(g)/(i) labeling results independently. 15.6 THE INSTANTIATED PROTOCOL (locked 2026-08-02, before scoring). §15.1 requires concrete values, not a description of values. These are they. Forward capture began 2026-07-28 (live wire, tier R); as of this writing 36 forward deals have been captured and NONE has been scored. Everything below is fixed as of publication of this section. a) EVENT UNIVERSE. Announced M&A deals with deal value > $5B, captured live from the wire with a capture timestamp (ma_wire realtime feed). No retrospective additions: a deal enters only via live capture. b) ENTRY. The first regular-session close STRICTLY AFTER the captured announcement timestamp. After-hours announcements therefore enter the NEXT session — this is the §8.10/§10 fix and it is non-negotiable. c) HOLDING PERIOD. 1 trading day. Exit at the next regular close. d) BENCHMARK. Excess return vs SPY over the identical window. e) LEG SCREENS, as of the event date (never current-date): float < $250M USD; price <= $20 at entry. Legs with UNKNOWN float are EXCLUDED, never assumed small. f) ATTENTION GATE — FIXED EX-ANTE THRESHOLD. attention ratio = (avg daily target-article pageviews over the 3 days ENDING at the entry session + 1) / (avg daily views over day -30..-4 + 1). HIGH attention = ratio >= 5.0. This 5.0 is now a FIXED CONSTANT, not a per-sample median. Fixing it is the §13.4(c) requirement: the historical study used the in-sample median, which is itself a lookahead. The window ends at entry, so no post-entry pageview data enters the gate (the §8.9 fix). g) PRIMARY ENDPOINT (one, per §15.3). Mean 1-day excess-vs-SPY return per event of HIGH-attention Tuatara legs, MINUS the same quantity for matched-random controls. Controls are drawn from the same universe under the same float/price screens, matched per event, same entry and exit convention. Success = this difference is positive with p < 0.05, two-sided. Everything else we report is secondary/exploratory. h) SAMPLE SIZE / STOPPING RULE. The primary endpoint is evaluated ONCE, at n = 100 scored forward events. No interim peeking is used to stop early or to keep going; if we report anything before n = 100 it is labeled interim and is not the test. i) BLIND LABELING (§13.4g/i). Relationship correctness is scored by a rater who sees the deal and the surfaced leg but NOT the return, and the same rater scores matched-random and retrieval-baseline legs in the same blinded pass. j) EXECUTION REALISM. Returns are reported both gross and net of a modeled spread/impact charge from quoted spread and volume at entry. The primary endpoint is judged on the NET series. If any value above changes before scoring begins, the prior version stays published alongside the new one with the reason (§15.1). After scoring begins, nothing here changes. This section exists because a reader's close made the point exactly right: if the methodology is locked in advance, both positive and negative outcomes become informative. Locked in advance is the only version of this test worth running. Contact: contact@cymetica.com | Get your own Tuatara model: cymetica.com/contact