Skip to content

Latest commit

 

History

History
387 lines (341 loc) · 71.6 KB

File metadata and controls

387 lines (341 loc) · 71.6 KB

What the teardowns taught us

Famous trading ideas — anomalies, folk strategies, vendor backtests, things people swear by — each put through the same protocol and stamped twice: is the signal real? and does it survive real execution and scale? This page is the view from above. It aggregates; it doesn't re-judge — every verdict below links back to the study that earned it.

The bench map — every study on a Signal × Tradability grid

(Regenerate with python tools/make_bench_figures.py — it parses the ledger, so it's always in sync. A cell shows its studies by number while they still fit and just counts them when they don't; the live map is where you click into one and read what's inside.)

Counting happens in one place. This page describes shape, not totals. The ledger is the only complete list, the map above is redrawn from it, and anything that needs an exact number should be read off one of those two rather than quoted here — a count copied into prose is a count that goes stale the next time a study lands, which is how this page once opened with a total that was hundreds out of date.


The score

Where the bench ends up, as a share of it. Every study carries both stamps except 14, which is pre-registered with its verdict still pending:

Investable Fragile Mirage
Real 1% 8% 4%
Weak 11% 23%
None 1% 53%

(Shares of the bench, rounded to the nearest point, so they don't sum to exactly 100.)

Read it the way the colours tell you to:

  • About one famous idea in eight is statistically real. The other seven-eighths are weak (Mixed folded in) or plain noise.
  • Four in five are mirages once you try to trade them. Costs, capacity, decay, or the discovery that the "edge" was beta all along.
  • The green column is a rounding error, and it holds four distinct flavours. The originals — Storm-Shy, All-Weather, Balancing-Act — are risk-managers, not forecasters. Those added by the "green hunt" lot (591–640) are harvestable return engines, and tellingly none is a crystal ball either: Fallen-Angels is a forced-seller risk premium, Currency-Hedged-Carry is a covered-interest-parity rate-differential identity, and Starting-Yield is duration arithmetic (the yield on your ticket ≈ your next decade). Those added by the "plumbing" lot (913–962) are a third kind again — not a return you go and get, but a cost you stop paying: Total-Cost-of-Ownership prices the fee-versus-spread trade-off by holding period, Hidden-Financing backs out what a leveraged wrapper really charges you to borrow, and Which-Gold shows the cheapest wrapper for one identical metal genuinely wins. The one from the measurement lot (963–1012) is a fourth kind: a discipline that pays for itself. Confirmation — wait k days before acting on a signal — cuts whipsaw round trips from 78% to 30% [981], and the study prices both sides of the trade-off instead of netting them, then admits the winning k is only knowable afterwards. You can bank a premium, an identity, a saving, or a discipline; you still can't see the future.

The single most important cell isn't the green one — it's Real × Mirage: effects that are genuinely there in the data and still can't pay you, and there are far more of them than there are greens. That gap between "true" and "tradable" is the bench's whole thesis, measured — and it has now widened twice: first under the green hunt (chase likely-real premia and most still die at the trading desk — borrow fees, roll drag, one-way bounds you can't stand on), then under the measurement lot, where an effect can be real because it is arithmetic and still unbankable — the 1% bitcoin sleeve [1003] is the cleanest case.

A quieter cell worth a look: None × Fragile — gold [69] and bitcoin [70] flunk the claims made for them (inflation hedge, digital haven) yet keep a Fragile stamp as plain diversifiers. The story dies; the asset survives. The measurement lot added three of exactly that shape: mismatched holiday calendars [973], silver as "gold with a beta" [987] and the hunt for cycles in a Fourier spectrum [1000] — a folk claim that fails, sitting on top of a real thing you still have to handle.


Where ideas go to die — mortality by family

We sorted the bench into rough families. The boundaries are judgement calls (is the 52-week high a chart pattern or a momentum factor? we said factor) — the shares below are honest, the taxonomy is approximate. This table is the taxonomy: tools/make_bench_figures.py parses the family rows below rather than carrying a family map of its own, so a study missing from here is a study missing from every per-family count.

Family Real Survived costs* Investable
Equity factors & fundamentals — 18 34 38 43 44 45 46 50 51 52 53 54 57 58 64 65 88 94 121 122 123 124 153 154 155 177 198 199 200 201 206 211 212 218 219 220 228 229 230 231 232 233 236 238 239 240 242 243 244 250 262 263 327 330 331 332 356 362 363 365 368 369 390 391 395 396 400 563 568 569 570 571 572 573 574 601 623 627 628 629 630 631 652 653 654 664 718 721 745 746 747 748 749 789 790 791 798 799 801 802 853 854 855 856 857 858 859 860 861 862 900 901 902 904 4% 18% none
Academic factors — pointed replications — 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 869 871 872 873 874 875 876 903 905 6% 12% none
Vol, hedges & allocation — 3 6 16 30 61 62 63 68 69 70 83 86 92 97 100 101 102 110 111 112 117 130 131 133 134 144 145 156 157 167 168 171 172 173 174 175 203 204 205 207 209 210 221 222 273 274 275 276 277 291 292 293 294 295 323 324 325 326 337 338 339 340 341 342 354 357 358 359 361 370 371 372 373 374 375 578 579 582 583 584 585 586 587 591 592 593 594 595 596 597 598 599 600 615 617 624 626 633 655 656 657 659 663 711 712 713 714 715 716 763 764 766 767 768 883 890 891 893 894 895 896 897 898 899 911 912 21% 34% 16 68 97
Technical & chart patterns — 2 7 8 13 15 17 19 21 22 72 73 74 75 76 77 78 87 91 93 98 99 104 106 107 108 109 116 125 126 127 128 129 137 176 178 179 180 181 182 183 184 185 186 187 188 189 190 202 301 352 353 397 398 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 541 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 4% 11% none
Momentum & trend — 20 24 25 28 31 40 103 105 146 225 237 632 638 15% 77% none
Carry, curves & commodities — 27 29 35 36 59 60 113 132 147 208 302 303 304 306 308 309 310 364 380 382 576 580 581 610 611 612 613 614 619 658 660 661 662 725 727 792 793 794 795 796 797 863 864 865 867 868 884 885 886 887 888 889 892 906 907 908 909 910 17% 29% 610 613
Calendar & seasonal — 1 41 42 48 55 67 79 80 81 82 89 90 95 96 135 136 142 143 148 149 150 158 159 161 162 163 164 165 166 191 192 193 194 195 223 224 226 227 234 235 247 248 278 279 280 281 282 283 284 285 286 287 288 289 290 296 297 298 299 300 307 314 319 320 321 322 542 543 544 545 546 547 602 603 604 605 606 609 616 637 639 640 641 642 643 644 645 646 647 648 649 650 651 707 708 709 710 719 720 723 724 726 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 770 775 776 777 779 780 781 786 787 849 9% 8% none
Mean reversion & stat-arb — 5 23 26 32 33 71 196 241 251 329 333 366 367 618 620 621 800 41% 29% none
Macro & valuation timing — 37 47 49 56 66 85 114 115 118 119 120 151 152 160 197 214 215 216 217 245 246 255 257 258 260 261 264 265 266 267 268 269 270 271 272 305 311 312 313 315 316 317 318 360 381 383 384 385 386 387 388 394 548 556 575 577 607 625 717 728 729 752 753 754 755 756 757 760 762 823 824 825 826 827 828 829 830 831 832 848 866 877 878 879 880 881 882 6% 7% 625
Wrapper mechanics, plumbing & implementation — 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 24% 34% 920 945 961
ML & model forecasting — 10 12 39 84 138 139 765 0% 0% none
Microstructure & crowds — 4 9 11 140 141 169 170 213 249 252 253 254 256 259 328 334 335 336 351 376 377 378 379 389 392 549 550 551 552 553 554 555 557 558 559 560 561 562 564 565 566 567 588 608 622 634 635 636 722 750 751 758 759 761 769 771 772 773 774 778 782 783 784 785 788 843 844 845 846 847 850 851 852 870 8% 19% none
Research-method demos — 343 344 345 346 347 348 349 350 355 393 399 401 589 590 833 834 835 836 837 838 839 840 841 842 0% 0% none
Measurement, estimation & portfolio construction — 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 60% 80% 981
Pre-registered — 14

* "Survived costs" = stamped Investable or Fragile (alive on paper, even if thin). The complement is Mirage.

Four patterns jump out:

  • ML & forecasting is the deadest corner of the bench — nothing in it survives. Every model-driven forecaster — Markov pipeline [10], ARIMA+GARCH [12], neural net [39], the Stock-to-Flow model [84], a Random Forest [138] and the 'AI-powered' ETF [139] — produced an in-sample story and an out-of-sample coin flip.
  • Calendar effects are the opposite failure mode: among the most real per capita and almost none tradable. The pattern is genuinely in the data; the trade built on it forfeits more than it captures (42, 55) — or dies the moment it's published (67).
  • Momentum, trend and carry don't die — they limp. These families collect Fragile stamps, not Mirage ones — mostly alive-but-thin: premia with a century of literature that one tape can't certify and costs nearly erase.
  • The measurement lot inverts the whole table: most of it is real, most of it survives costs — and one entry is investable. Every other family asks does this edge exist? and mostly hears no. This one asks what does the choice of estimator, window, benchmark or rebalance date do to the number you publish? — and the answer is almost always "something real", because these effects are properties of the arithmetic, not of the market. The tradability column is the punchline: knowing that rebalance dates are a lottery, that a correlation matrix is mostly noise or that beta has a half-life makes your measurement honest. It does not make you money. Real is cheap here; the costs line is still where it ends.

Six lessons the bench keeps teaching

These aren't opinions — each one fell out of multiple studies independently.

1 · The edge dies at the costs line, not the signal line. Of the signals that clear the statistical bar, all but a handful still failed or barely survived tradability. The overnight drift is real and untradable [01]; intraday reversal is real with a 3.31 bp break-even that lives in the least-liquid names [33]; the turn-of-the-month premium is real at t = 5.1 and a window-only book — even with its cash leg paid the T-bill — still compounds half of buy-and-hold [42]; IBS snap-back is real and gone at the spread [19]. Beat 6 — could you trade it? — is where almost everything dies.

2 · Survivorship doesn't just flatter results — it manufactures and even inverts them. On a survivor panel of large caps, the lottery effect ran backwards (−10.4%/yr, t = −2.5) [53], the idiosyncratic-vol puzzle inverted decisively [54], the 52-week-high premium came out negative [50], the net-issuance hedge flipped because the decade's diluters were the growth winners [64], and asset-growth showed nothing where the literature's premium hides in micro-caps [44]. When we found a positive result on a survivor panel, we capped it as an upper bound [48] — the bias cuts both ways and we say which.

3 · Post-publication decay is the norm, not the exception. The pre-FOMC drift is the most spectacular case on the bench: 3% of sessions carried 11.5% of SPY's entire cumulative return — until Lucca-Moench published it in 2011 and the drift collapsed from +0.24%/day to +0.09% [67]. The size premium never showed at all on a 39-year tradable proxy [45]; turn-of-the-month faded from 13.8 to 4.8 bp/day after 2008 [42]; betting-against-beta decayed 0.70 → 0.28 [43]; textbook pairs stopped paying once everyone copied them [05]; the vendor's dual-momentum edge thinned right after the publication that sells it [40]; oil-predicts-stocks reads exactly zero out of sample [49]. An anomaly's discovery date is the start of its obituary.

4 · Leverage is never free — the "free lunch" is usually the financing bill, in disguise. Betting against beta needs 2.78× leverage to be market-neutral, and realistic financing drags its Sharpe from 0.47 to 0.02 [43]. The retail CFD markup — charged on the whole notional, not the borrowed slice — costs a levered dip-buyer 2.65 pts/yr [30]. The 3× ETF "free amplifier" tripled the drawdown, not the Sharpe (0.90 vs 0.98, −82% trough) [61]; extending duration for term premium lowers the Sharpe [59]; and vol-targeting the carry trade makes its crash worse [27]. When a strategy's appeal is "same return, just levered," the lender has already priced your idea.

5 · What's green on the bench you earn by managing risk, banking a premium, or reading an identity — never by forecasting. For most of the bench's life the green column was pure risk machinery: scale exposure down when markets get loud [16], balance risk across assets and win on Sharpe not return [68], or hold the plain 60/40 [97]. The "green hunt" lot (591–640) — hand-picked because they should be real — finally added three greens you buy for yield, and the lesson survives them intact: a forced-seller risk premium you get paid to absorb [610], a rate-differential identity a currency hedge hands you mechanically [613], and duration arithmetic where the yield on the ticket is the return [625]. None forecasts anything — you collect a premium or read an identity. And the hunt's own body count proves the rule: chase 50 likely-real edges and most still die at the desk — borrow fees [617], roll drag [619], one-way bounds you can't stand on [621], the most famous alpha on Earth already spent [628]. Nothing on this bench forecasts returns and pays. Several things manage risk, harvest a premium, or bank an identity — and those do.

6 · Before you ask whether the edge is real, ask what the number is made of. The measurement lot (963–1012) went looking not for edges but for the choices buried in every backtest — which volatility estimator, which window, which benchmark, which rebalance date, which bootstrap — and found that those choices move the published number more than most of the "anomalies" on this bench ever did. A rebalance date chosen a week apart changes the result [997]; a correlation matrix estimated from a normal sample is mostly sampling noise [1010]; beta has a half-life and the Blume slope everyone uses to "correct" it turns out to measure the signal-to-noise ratio rather than the stability [1005]; a purged-CV fold boundary quietly invents a negative IC out of signal-free data [1001]; and your Sharpe ratio depends on the currency you happen to bank in [995].

Several of them rejected their own pre-registered hypothesis, and those write-ups were kept rather than rewritten — the data can tell 1% from 5% bitcoin and wants 16.5% [1003]; the "most stocks underperform cash" result is absent on a survivor panel [1006]; time does not diversify, but the small-sample bias in measuring that is larger than the effect [1007]; Sortino and Sharpe rank identically (Spearman 1.000) [1009]. The lot's own measurement errors were caught mid-build and are now pinned by tests that fail if the mistake comes back. That is the lesson in one line: the most reliable way to find a signal that isn't there is to measure carelessly, and these hold up precisely because they are properties of the arithmetic. Exactly one of them is investable.


The podium

🟩 The ones that made it. The risk-managers came first; the "green hunt" lot (591–640) added harvestable return engines — the first time the bench's green column contained anything you buy for its yield rather than its calm; the "plumbing" lot (913–962) added savings; and the measurement lot (963–1012) added a discipline.

The risk-managers — they predict nothing and win on risk-adjusted terms: 16 · Storm-Shy — Real × Investable. Scale exposure down when realized vol spikes. It survives robust inference, real costs, capacity, a parameter sweep and a third tape — and note what it is: not an alpha, a risk overlay. 68 · All-Weather — Real × Investable. Risk parity earns the best Sharpe of anything we tested (0.92) with a third of equities' drawdown — by predicting nothing and balancing everything. Half the return of stocks, though: the green is risk-adjusted, not absolute. 97 · Balancing-Act — Real × Investable. The plain 60/40 lifts the excess-of-cash Sharpe over 100% stocks (HAC t = 2.3, a bootstrap CI clear of zero) and halves the drawdown — but it forfeits ~2.6 pts/yr of return, leans on the historic bond bull, and the bonds did not cushion 2022. Risk-adjusted, not absolute.

The return engines — premia and identities you can actually bank, and still not a crystal ball: 610 · Fallen-Angels — Real × Investable. Bonds kicked out of investment grade get dumped by forced sellers; catching them (ANGL over HYG) pays +18.3 bp/mo, HAC t = 2.44 — and it's not duration (β to Treasuries ≈ 0). It survives dropping the entire 2020 wave, a single-year jackknife, and a block bootstrap (CI clear of zero), and it clears costs with ~1.4 bp/yr of drag. A genuine forced-seller risk premium. 613 · Currency-Hedged-Carry — Real × Investable. Two funds hold the same Japanese basket; the hedged one (HEWJ) quietly out-earns the unhedged (EWJ) by the whole US–Japan rate gap — +1.2%/yr at HAC t = 2.7 even before the 2022 hiking cycle, pass-through slope ≈ 1.0. A covered-interest-parity mechanical identity, not a signal — no forecasting, just collecting the differential the wrapper hands you for free. 625 · Starting-Yield — Real × Investable. The 10-year yield on your buy ticket predicts your next decade of bond returns at R² = 0.92, slope t = 12.5, identity slope ≈ 1 — pure duration arithmetic (Bogle/Leibowitz), robust across pre/post-1950 and a drop-one-decade jackknife. One entry per decade, ~$0 cost, effectively unlimited capacity. You don't forecast the yield; you read it off the ticket.

The savings — not a return you go and get, but a cost you stop paying: 920 · Total-Cost-of-Ownership — Real × Investable. The cheapest fund is not the one with the lowest fee: expense ratio and spread trade off against each other, and which wins is decided entirely by how long you hold. The crossover is computable in advance, per holding period, and it is the rare bench result you can act on before you place the trade rather than after. 945 · Hidden-Financing — Real × Investable. Back out what a leveraged wrapper actually charges you to borrow, rather than what the factsheet says. The implied rate is recoverable from the fund's own tape, it is materially above the headline, and it is a cost you can decline by financing the leverage yourself. 961 · Which-Gold — Real × Investable. Several wrappers hold one identical metal, so the return difference between them is pure cost with nothing else in it — the cleanest natural experiment on the bench. The cheapest wrapper wins by exactly the fee gap, persistently, with no forecast involved.

The one discipline — the only green on the bench that is about how you act, not what you hold: 981 · The Price of Waiting — Real × Investable. Requiring k consecutive days of agreement before acting cuts whipsaw round trips from 78% to 30% and trades from 7.1 to 1.9 a year, across all 12 tape × signal cells. What makes it green rather than folklore is that the study prices both sides separately instead of netting them: 3,636 sessions spent in cash while the raw signal was already right, worth −216,670 bps forgone, against +253,703 bps avoided by exiting late. Some k beat the unconfirmed rule on Sharpe in 92% of cells by +0.110 — and the study says plainly that the winning k differs in almost every cell, which is what choosing it in hindsight looks like. Every arm carries the same one-day execution lag, so the comparison is about confirmation and not about being late.

🟨 The honest fragiles — Real signals that survive on paper but are thin, decaying, or capacity-starved. Worth knowing; not worth quitting your job for:

Study What's real Why only fragile
48 Groundhog Month-of-year seasonality, t = 4.1, undecayed Survivor-panel upper bound; breaks even near ~19 bp
52 Smoke-Screen Accruals: cash-backed earnings win, Sharpe 0.64 Short-side costs; documented post-2000 fade
56 Tide-Table CAPE forecasts 10-year returns (R² 0.28) A tide table, not a stopwatch — useless at 1 year
59 Downhill Term premium, +2.2%/yr over cash Sharpe 0.32 vs cash's 1.82; 2022 took −23%
63 Free-Fall Short-vol carry, +12%/yr (SVXY) Skew −4.8, one −83% day; five crash days wiped 95%
66 Inverted Curve inversion → +1% next 18m vs +16% normal ~5% of months, a year of melt-up first — no sell button
67 Fed-Drift Pre-FOMC drift carried 11.5% of SPY's return Publication killed it: +0.24%/day → +0.09% after 2011
71 Ambush Confluence of four dead-net edges: +19.6 bp/day at K≥3 (HAC t = 3.1), undecayed, costs defeated by rarity ~15 trades/yr → +1.2%/yr excess; OOS Sharpe +0.28 under the frozen 0.30 bar
75 Knee-Jerk Connors RSI(2) oversold bounce: pooled HAC t = 10.7, beats a coin by +57 bp/trade Decayed 35% since the 2008 book; long-only beta in a bull market; the 200-SMA filter hurts
103 Turtle-Trader The Turtles' Donchian breakout: a real long-side trend premium, HAC t = +11 Shorts are a structural trap on up-drifting markets; edge ~halved post-publication; years-long drawdowns
106 Supertrend The ATR(10,3) daily flip beats a coin (HAC t = +3.3, bootstrap CI clear of zero) Only the canonical multiplier works (2 and 4 are noise); ~2.5%/yr gross at ~8 flips/yr
110 Faber-Timing The 200-day timing rule lifts SPY's Sharpe 0.55→0.73 and halves the drawdown (−55%→−22%) Pure risk reduction, not alpha; lags in bull markets; whipsaw + switching/tax drag
144 Permanent-Portfolio Browne's 25/25/25/25 stocks/long-bonds/gold/cash: a genuinely low-drawdown all-weather mix Forfeits much of equities' return; leans on the gold + bond bull; risk reduction, not alpha
151 Stocks-For-Long-Run The real equity premium holds in every long rolling window (Siegel) The unit of time is the decade; 20-yr windows can still trail bonds; useless as a timer
173 Four-Percent-Rule Bengen's 4% withdrawal survived every historical 30-yr US retirement cohort Sequence-of-returns risk; today's valuations/yields; non-US history has failed it (Pfau)
203 Golden-Butterfly Tyler's 20x5 mix beats SPY on Sharpe (0.68 vs 0.53) with a third of the drawdown Loses to its simpler parent (Permanent Portfolio); forfeits ~3 pts/yr of return; risk reduction, not alpha
209 ETH-BTC-Ratio The 20-day ETH/BTC momentum rotation beats 50/50 (HAC t=+3.7, alpha t=+9.4) Strapped to a single crypto cycle, -68% drawdown; one regime, not a law
210 Crypto-Trend A 200-day timing rule cuts Bitcoin's -83% crash to -70% and beats buy-and-hold on Sharpe (0.90 vs 0.63) Whipsaws; ~7 switches/yr; ~12-year history; a drawdown shield, not alpha
223 Same-Month Seasonality Same-calendar-month return persists: top-bottom decile spread HAC t = +5.57 over 330 months, bootstrap Sharpe CI clear of zero Survivor-inflated on ~8-stock deciles; ~78% monthly turnover; short leg hard-to-borrow; a live Russell-1000 build dilutes the gross edge
301 Triple-RSI The viral "90% win-rate" RSI(5) bounce is genuinely real: SPY +132.8 bp/trade, HAC t = +5.07, survives an honest next-open fill and post-2010, beats a coin by +85 bp ~3.5 trades/yr, ~7% of the time in the market → ~+4.7%/yr vs the index's +10.8%; the 90% win-rate is the exit's shape, not the edge (a coin clears 62% at a loss)
302 Lithium-Boom A 200-day trend overlay on lithium (LIT) is real — net HAC t = +2.09, and trend-timing beats same-exposure random timing at the 100th percentile Net excess-Sharpe ~0.51 (bootstrap CI [0.02, 0.99] barely clears zero), trails SPY buy-and-hold; the only real benefit is a halved drawdown (−66%→−41%) — crash-dodging, not skill
340 Bank-Loans Floating-rate loans (BKLN) really do dodge duration: beta to long Treasuries = −0.055 (HAC t = −2.95), and BKLN gained +13.9% through the 2020–23 bond repricing The risk just moved duration→credit (equity beta +0.20, t = +4.48); a thin sleeve (CAGR 3.7%) that fell in 7/7 equity crashes and gapped −24% in March 2020
363 PEAD-Drift Post-earnings drift sorted on the EPS surprise: +1.34% at 20 days (t = 2.96), survives a quarter-block placebo Nothing in week 1; net edge only at 20–60-day holds on a 30-name survivor basket; thin and long-leg-dependent
367 CEF-Discount Widest-discount closed-end funds beat the narrowest by +6.6%/yr, market-neutral (Welch t = 3.72), placebo-clean A NAV proxy on a survivor basket; fades to t = 1.56 post-2010; hard-to-borrow shorts, ~18 tiny funds
375 VXX-Roll-Decay Shorting a VIX-futures ETP's contango is a real carry: +0.171%/day, HAC t = 2.51, +35%/yr net Skew −1.65, a −43% day and a −92% drawdown make a constant-notional short un-allocatable
601 Factor-ETF-Live-Test The factor ETFs do deliver their exposure live: USMV cuts vol to 0.80× (t vs 1 = −7.5), MTUM/VLUE/QUAL load their factor at HAC t = +7.5/+10.0/+8.8 Exposure is real, alpha isn't — none reliably beats SPY once you pay for it; a delivered risk profile, not free return
618 GBTC-Premium-Cycle All three regimes verified to the digit — +36% premium (t = 9.0), −24% discount (t = −7.4), then par; the 2023 discount→par convergence was a real, dated trade A one-off wrapper lifecycle that ended at the Jan-2024 ETF conversion; not repeatable now
622 Thematic-ETF-Curse A calendar-time book of 48 thematic ETFs loses −16.3%/yr CAPM alpha (HAC t = −3.27) in their first 36 months; the broad-index-launch placebo is clean It's a short-the-hype signal: costly-to-borrow small names, clustered 2020–21 launches, survivorship flatters the long-only escape
626 Unemployment-Trend-Timing Gating Faber's 200-day rule on rising unemployment beats the pure rule by +12.3 bp/mo (HAC t = +2.32) and cuts whipsaw spells 64% over 928 months Most of the edge is trading less, not seeing more; current-vintage unemployment (revisions unmodeled); a smoother Faber, not alpha
628 Buffett's-Alpha The most famous alpha on Earth replicates: +9.5%/yr CAPM alpha, HAC t = 3.47 over 46 years (the FKP t > 3 holds to 2011) The fade is itself significant (−11 pp/yr post-2010, t = −2.38) — quality + low-beta + cheap leverage, mostly spent; little left for today's buyer
632 Crypto-XS-Momentum Last week's crypto winners keep winning: +164 bp/wk on the winners-minus-losers quintile, HAC t = 3.56 over 449 weeks on a 44-coin panel including dead pairs (LUNA shorted) Rides the 2017/2021 bulls; ~55% hit rate drowned by 10–50 bp crypto spreads; paid roughly nothing through the LUNA/FTX year
806 Prospect-Theory Value The Barberis-Mukherjee-Wang cumulative-prospect-theory value predicts returns negatively: long low-TK / short high-TK earns +138.7 bp/mo (HAC t = +3.15 over 137 months, positive in both eras, correct BMW sign) Survivorship flatters the short (blown-up lottery names absent) and a 50-name universe concentrates it into hard-to-borrow lottery mega-caps; +124 bp/mo net at 5 bps but real borrow/squeeze exceeds the charge
812 Corwin-Schultz Spread The high-low bid-ask illiquidity estimator earns a premium — high-spread names out-earn by +4.45 bp/day (HAC t = +3.24, both eras, +4.88σ placebo) The illiquid long leg pays the very spread it earns: net +2.31 bp/day at 1 bp (t = +1.59, insignificant), −5.69 at 5 bps
861 Debt-Maturity Rollover Firms funded with a high share of SHORT-TERM debt under-earn the long-funded — low-minus-high-rollover-share tercile +45 bps/mo (NW t = +3.22), right-signed, both regimes (post-2017 +2.83, post-2019 +2.51), ~2.6× stronger in the 2022+ hiking era Thin ~26-name survivor cross-section (no delisted rollover-wall casualties), monthly turnover + hard-to-borrow shorts; a real balance-sheet risk premium too small and capacity-starved to bank cleanly
888 CLO AAA Carry The senior AAA CLO tranche (JAAA) pays a genuine low-vol carry over cash — +1.38%/yr on 1.63% vol, excess Sharpe +0.84 (HAC t +2.33, block-bootstrap CI clear of zero), topping the excess-Sharpe race and beating the un-tranched loans it's carved from ~5.7y stress-free sample (misses the 2020 CLO mark-down), and all the carry lives in the high-rate era (ZIRP excess ~0) — mechanically a short-rate+spread, so regime-bound and short-history; thin and not yet crisis-tested
889 Broad Dollar-Hedge Overlay Generalising 613: on broad developed-international (HEFA/DBEF vs EFA, the same basket) the hedged-minus-unhedged return IS the US–EAFE rate differential mechanically (β on −fx ≈ 1, high R², HAC t clears the bar, bootstrap CI clear of zero) The identity is robust but the whole tape sits in one dollar-regime (US out-yielded EAFE); the tradable overlay is regime-limited and the pickup is small — a mechanical identity, not a forecast

Challenge the bench

This page will be wrong eventually — that's the design. Every verdict on the bench is a falsifiable claim (bar 14, still pre-registered), each with reproducible code, pinned data fingerprints, and the exact line where we think the dream dies.

  • Think a Mirage is tradable? Fork the study, change the cost model or the venue, and show the break-even. Beat 7 of every notebook says what we'd consider convincing.
  • Think a None is real? The inference stack (HAC, Lo, bootstrap, Reality Check) is in quantlab/ — run it on your variant.
  • Got a candidate for the queue? Open an issue. The ideas that look most embarrassing to test are usually the best ones.

The map gets a new chip every time. python tools/make_bench_figures.py redraws it.


Part of Open-Alpha-Lab. Counts generated from the ledger by tools/make_bench_figures.py, and checked against the studies themselves by tools/check_reference_table.py. Not investment advice — research and education. See LICENSE.