1880Q
September 16, 2026

The Scorecard, Tested

We spent a week putting our own Scorecard on trial. Here is what held up, what did not, and what changes on the site because of it.

We spent a week putting our own Scorecard on trial. Here is what held up, what did not, and what changes on the site because of it.

Everything below was run the hard way: a hypothesis written down before the data was touched, tested on two separate eras of market history, and only counted if it held in both. Where our own guesses were wrong, we say so.

Finding one: the grade describes a company. It does not predict its stock.

We measured whether the Scorecard's composite, the 0–100 number and its Strong Buy to Sell tier, ranks names by what they went on to do against the market over the next one, three and six months. Across the S&P 500 from 2021 to today, it does not. The rank correlation is zero. Ranking within a sector instead of across the whole index does not change that.

That is not a tuning problem, and it is not fixed by fitting the weights to recent history, which would only teach the Scorecard the last five years. The grade is an honest profile of where a company stands on five fixed yardsticks. It was never a forecast, and from now on the site says so.

Finding two: the inputs change sign by sector, and that is why.

The same number means opposite things in different industries. This is the central result, and it survived the earlier era of history as well as the recent one.

Input Helps in Hurts in
Revenue growth (Demand) Technology, Staples, Industrials Energy, Utilities
Margin change Nearly everywhere Energy
Cheapness (free-cash-flow yield) Unstable, recent era only Technology, Communication Services
Revenue acceleration Energy only Financials, Consumer Discretionary
Price momentum Better than the moving-average stack Fails the bar everywhere

A single global weighting averages across those signs and lands near zero. In Energy and Utilities the fastest-growing companies went on to lag their sector in both eras. Growth there marks the top of a cycle or a build-out the market has already paid for, not demand. In Technology the cheapest names on cash flow were the broken ones. The Scorecard was rewarding the wrong thing in some sectors and the right thing in others, and the two cancelled.

Finding three: one sector earned a live test.

In Energy, one pair of signals ranked names in both eras: quarterly revenue accelerating while debt falls. It cleared our bar on 2021–26, cleared it again on 2015–20, and held on a third, independent look across every sector.

It now runs every night as a pre-registered live read across 82 Energy names, with its rule frozen and every outcome recorded. The live test began on September 16; a verdict takes about nine months. If it fails in the real world, it comes off the site and the failure is published. That is the only way a sector-specific read gets onto 1880Q.

Finding four: everything else failed, including our own best guess.

  • Technology (growth plus widening margins) looked good in the recent era and vanished in 2015–20.
  • Consumer Discretionary (cash yield, margin change, momentum) did the same. The sector is a dozen unrelated industries, and a sector-wide score mostly sorts industries.
  • Utilities was positive in both eras but short of the bar. The one piece that replicated was the negative: revenue growth is bad for a utility, in both eras.
  • Our own hypothesis that revenue acceleration works everywhere was wrong. Nine of eleven sectors said no. The written-down prediction caught it, which is what it is for.

In every sector, the "best" combination a computer could mine from the data claimed more than the pre-registered one delivered. That is why only pre-registered results count here.

Three changes on the site

  1. Demand is now read by sector. In Energy and Utilities, trailing revenue growth counts against a name. The card marks it: "growth inverted." Everywhere else growth keeps its sign. This is the one finding strong enough to act on without a further test.
  2. The grade is presented as a profile, not a forecast. Every card now says so under the tier, and always prints the base rate for its grade: how often names with that grade actually beat the market, with the sample size beside it. The full table is public on The Record, along with a dated log of every change to the method.
  3. The 10th Man knows the sector. Our written-down dissent is now told what the tests found in that industry, so it knows that cheap in Technology and growth in Utilities are warnings, and argues accordingly.

What we will not do

We will not fit the Scorecard's weights to the current library. We will not add a sector read to the site without a pre-registered pass on two eras of history and a live trial. And we will not present a number as a forecast until the ledger shows that it is one.

The receipts start arriving in October

Every score the site has ever produced is being measured against what the stock then did. The first one-month outcomes land in mid-October, and the cards begin printing real base rates as each grade reaches twenty resolved cases. The Energy read's first outcomes arrive at the same time; its verdict comes next summer.

Two leads remain, and both can only be tested live now that the history has been used: acceleration without cheapness in Technology, and inverted growth alone in Utilities. Neither is on the site.

How the tests were run

S&P 500 constituents by date added, month-end scores, forward excess return against the S&P (or the sector ETF for sector tests) at 21, 63 and 126 trading days. Financial statements count only from their filing date; prices only up to the score date. Two eras: 2021–2026 and a 2015–2020 holdout drawn from deeper history that was fetched only after each hypothesis was written. A test passes when the rank correlation is positive in both eras, positive in most months of each, and clears a t-statistic of two on the overlap-adjusted sample. Only the pre-registered composite counts. Every pre-registration and results file is kept in the codebase.

One known limitation: companies removed from the index during the sample are not fully reconstructed, so all results carry a survivorship caveat. It biases toward flattering the signals, not against them.

1880Q · Quantitative research, written down before it is believed. This is research, not investment advice.

1880Q

The tools behind these notes — the Scorecard, the daily Top 10, Receivables, Commentary and Price Range — are open to members. Start a 7-day free trial or read more posts.