II
Research·

Systems Travel, Edges Don't

Christopher CottiSelf-published

A pre-registered falsification study of coach–quarterback fit as an unpriced edge in college football betting, and the four bounded nulls it returned.


The hypothesis had an attractive shape. Betting lines are built from additive power ratings: a coach's record is in the number and a quarterback's pedigree is in the number, but fit between them is an interaction, and an additive model has no term for an interaction. If mismatch costs real points, that cost sits outside the price by construction.

It does not survive contact with the data. This is the record of how it failed, and of three further constructs that failed after it.

What held

A coach's system does travel with him. Decomposing play-calling identity across 62 same-coach school changes gives 55.8% coach, 9.1% school, 35.1% year, with a cross-move correlation of 0.558. The premise of the whole project — that a system lives in the coach's head and arrives intact at a new job — is a durable empirical result, and it is the one asset the project produced that outlived its own thesis.

What died, and how

Everything downstream of that premise.

The quarterback side of the fit measure had to be a contrast — passing-down efficiency minus early-down efficiency, within player — because any non-contrast measure is quality, and quality is a main effect the market already prices. The contrast differences quality out. It differences out everything else too: split-half reliability −0.008. Split one quarterback's own season at random and the halves do not agree.

That collapsed the design to adaptation cost — distance between systems — which returned +0.009 points per standard deviation, 95% CI [−0.33, +0.35], across 5,022 team-games. The interval excludes any effect above 0.95 percentage points of cover probability, against the 2.38 needed to clear vig at −110. The market does not price system distance, and system distance does not predict what the market missed. That is a more interesting failure than being front-run: the interaction does not exist at measurable magnitude.

Two further constructs were specified and refused. A tempo hypothesis — that totals under-adjust when a coaching change brings a predictable pace shift — was registered, powered, and not run, because its own power calculation killed it. The signal is real: a coach's pace estimate keeps 28% of its variance after being residualised against what the market plainly already knows, and what survives still predicts realised pace at p < 0.0001. But the transmission chain, measured link by link, ends at 0.50 percentage points against 2.38 needed. The weak link is pace to drive count: −0.406 drives per game per standard deviation of pace, not the −1.51 a naive conversion assumes. Game clock is fixed, so a slower snap is absorbed by plays per drive rather than by the number of possessions.

The fourth, on whether line movement carries information the recorded price has not absorbed, returned a bounded null on both spread and total.

Why a null is the deliverable

Four constructs, four intervals that exclude a tradeable effect rather than merely failing to reject one. The distinction is the point. A study that cannot distinguish "no effect" from "not enough data" has produced nothing; each of these puts a ceiling on what could be hiding.

The methodology is the argument. Decision rules were fixed before the data was touched, and the pre-registration records the tests that were refused as well as the ones that ran — a decision not to run is precisely what a file drawer swallows. Cluster-robust standard errors were re-estimated under a wild cluster bootstrap after simulation showed the conventional estimator rejecting a true null 9.5% of the time at a nominal 5%. Three source defects were found that are plausible-looking, fully populated, and wrong, none of which throws an error. Two errors of my own are documented in place rather than quietly corrected: an overclaim about what the quarterback result established, and a bias estimate that was wrong by a factor of thirteen because it was computed from a median where the variance demanded a mean of squares.

Alongside the nulls sits a working piece: a log-optimal bankroll allocator, convex under log utility, correlation carried by scenario sampling rather than a covariance matrix, validated against ground truth it cannot fake — closed-form Kelly, brute-force search, and an optimality certificate shown to reject a perturbed allocation. Its most useful output is a warning. When edge estimates are noisier than the spread of true edges, betting on them destroys wealth; shrinkage reduces the damage without reversing its sign. With no measurable edge the correct allocation is zero, and the optimiser finds it.

The full study, including the instrumentation notes and every interval, is in the PDF above. Code and data pipeline: github.com/ctcotti/cfb-falsification-harness.

Cite this

Cotti, C. (2026). Systems Travel, Edges Don't — A Pre-Registered Falsification of Coach–Quarterback Fit as a Betting Edge. Self-published. https://github.com/ctcotti/cfb-falsification-harness

© 2026 Christopher Cotti

In the margins