Everything the funnel does, shown step by step — including the test that decided it doesn't predict anything, and, at the end, a replay you can scrub and a playground where you can try to tune the model yourself. This page exists because a ranking you can't inspect is a ranking you can't argue with, and arguing with it is the point.
Each month-end, the universe is every US-listed stock that existed on that date — including companies that later died or were bought. That matters more than it sounds: build the list from stocks that exist todayand the dead ones silently vanish, and any backtest will "discover" that stocks go up. Names are then screened for tradability — at least $5 (unadjusted) per share, at least $1M/day of volume, real book value, and enough trading history to measure a year of returns. Roughly 2,000 names survive to be scored.
Four standard factors, chosen for being the most-studied in finance — deliberately nothing clever, nothing fitted. Next to each is its measured correlation with the next month's returns (the "IC") over the full test: a value near zero means it told us nothing.
Total return over the last 12 months, skipping the most recent month (which tends to reverse).
The hope: Winners keep winning for a while.
Book value and earnings, each divided by market cap — from SEC filings, using only what was filed by the scoring date.
The hope: Cheap relative to assets and earnings gets repriced upward.
Return on equity: net income over book equity.
The hope: Profitable businesses stay profitable.
Negative 12-month realised volatility — calmer stocks score higher.
The hope: Boring compounds better than exciting.
Raw factor values are wild. A data quirk once gave a thinly-traded stock a "12-month momentum" of +5,459% — and in a raw average that one number owns the whole composite. So each factor is converted to a rankwithin the month (best, second-best, …), centred and scaled to roughly ±1.73. The most extreme value in the universe can never be more than "first".
The composite is the plain average of the four rank scores — equalweights, because weights you tune are a backtest you can't believe. Sort all ~2,000 names by composite: the top 20 become
Being top of this list means the four dials point its way this month. Given the test below, that is a reading list, not a forecast.
Go back to every month-end since 2017 — 113 of them. At each one, score every stock using only information knowable that day(a filing counts from the day it was filed, not the period it describes), then watch what actually happened over the following month. Group each month's names into five buckets by score. If the model works, the best bucket (Q5) should beat the worst (Q1) — a staircase. Here is the staircase:
Mean forward 1-month return per composite quintile, 2017-01-31 to 2026-05-29. The buckets aren't even in order — and the "best" bucket sits on the dashed line, which is what you got for free by holding everything.
Top-minus-bottom spread: +0.173%/mo (t = 0.35) — statistical noise. The top bucket against simply holding everything: +0.029%/mo (t = 0.14) — a rounding error from zero, before any trading costs. And the honest other half: 113months is too short a sample to detect a realistic edge even if one were there, which is why the claim is "we can't find one and the estimates sit within a hair of zero", never "we proved there is nothing".
The staircase above is a snapshot. Here is the same experiment as motion: each of those 113month-ends, every bucket's monthly mean compounded into growth-of-a-dollar on a log scale, next to the dashed line you got for free by holding everything. Press play, or scrub.
The signal, month by month: rank correlation of the composite with the next month's returns
A working model would braid apart into a staircase — Q5 compounding away on top, Q1 sinking below, the gap widening month after month. This braid never does: the buckets swap places the whole way and finish in a knot around the hold-everything line. The top-minus-bottom spread you just watched accumulate is +0.173%/mo (t = 0.35) — noise.
And the tuning trap, made playable. The model weights its four factors equally on purpose — weights you tune are a backtest you can't believe. Here you may briefly believe whatever you like: drag a dial and every one of the 113 months is re-scored, every quintile reassigned, every t-stat recomputed, in your browser, from the same graded panel. (The copy it runs on is quantized and carries no tickers — statistics, not names — and a test asserts equal weights through this engine reproduce the published headline within the quantization tolerance.)
fetching the graded panel…
If you find weights that clear the significance bar in here, you have rediscovered the reason this site leads with a negative result: given four dials and a fixed past, something always fits.
Every hypothesis we tested on the way here — and the two data bugs that briefly made the model look like it worked — is documented on What didn't work. The scan keeps running monthly because the data layer is solid and a disciplined shortlist is still a useful way to decide what to read. That is all it is.