This is the methods section: substrate, encoding, estimation, calibration, validation protocol and the machinery that makes every published pick auditable after the fact. It stops short of the parameterisation, which is not public. Everything else is here.
Live record bound at build time to the QC-signed canonical store · as of 2026-10-10 · 35 graded cards
In one paragraph. Every UFC bout is encoded as an orientation-invariant differential over point-in-time fighter state and, where a market exists, a de-vigged implied probability. An ensemble of independently trained estimators — gradient-boosted decision trees and regularised linear models, some anchored to the market price and some blind to it entirely — is combined by a logistic meta-learner fitted on out-of-sample outputs. The result is monotonically calibrated, scored under a proper scoring rule, validated walk-forward with a standing leakage battery, published before the card, and graded in public afterwards whether it was right or wrong.
How well that can possibly work is itself a question we researched rather than assumed. Our working paper, The Predictability Ceiling of UFC Fight Outcomes, measures the irreducible uncertainty of a UFC bout and finds the ceiling to be a property of the data rather than of any model — §7 states the finding and what follows from accepting it.
Across every pick we hold on file, 2023–2026 inclusive, that process has graded 1307-646 — 66.9% on n=1953. The 2026 live season, the only cohort published prospectively, stands at 292-129 (69.36%, n=421). Both are historical measurements, not forward promises.
The estimation set is the public professional record — bout outcomes and per-fight statistical lines reaching back to the early 1990s — joined to historical betting-market prices for the subset of bouts where a market existed. No private data, no paywalled feed, nothing a determined researcher could not assemble.
The substrate is assembled in a single chronological pass. Every state variable a fighter carries into a bout — rating, form, activity, layoff, durability, strength of schedule — is advanced only after every bout on that date has been featurised. A fight's own result, and every result after it, is therefore structurally unavailable to the row that predicts it. This is an as-of(D) construction rather than a filtered one: the guarantee comes from the order of the pass, not from a date predicate somebody has to remember to write.
Prices are taken from the latest pre-event snapshot strictly earlier than the event date, then converted to de-vigged (no-vig) implied probabilities — the bookmaker's margin removed so the two sides sum to one. That de-vigged probability is the market prior; it is never treated as ground truth.
One of the four voices behind every published probability is the market price. Until 2026-09-16 we took that price from the best line on the board, and the board we read included two prediction exchanges (Kalshi and Polymarket) alongside five sportsbooks. An exchange price is a different quantity from a sportsbook line: it is what other bettors were willing to trade at, not a price a bookmaker offered.
Reconstructing which venue actually priced each bout, we find that at least 88 of the 219 bouts we served and graded between 9 May and 5 September 2026 had at least one side priced only by an exchange. An independent recount by our QC lane, using a looser matching rule over a narrower set of cards, put the figure higher still — 114 of 192. We publish the lower number because it is the one we can floor, not because it is the likelier one.
We cannot make this exact, and the reason is worth stating. None of those served cards recorded which book its price came from — 0 of 218 of them — so the venue has to be reconstructed after the fact by matching the served price back to a historical snapshot of the board. Different matching rules give different counts. Treat the total as a floor and do not read any per-card figure as precise.
Two gaps we cannot speak to at all: the cards before 7 May 2026, for which no per-book history exists, and the 12 September card, which was priced from a median of sportsbook quotes because the usual board carried only 2 of its 12 bouts. Those are unknown, not clean.
Since 2026-09-16 the market voice must come from a sportsbook or it is absent. A bout no sportsbook priced is now withheld and named on the card rather than served on an exchange price, and every served bout records the book that priced it.
The published results of those earlier cards stand as published. We correct the process, not the record.
One of the four models behind every published pick, used to substitute a league-median fighter profile whenever it had no career record for a fighter. 55 of the 387 graded 2026 picks (14.2%, on 17 of the 32 cards) were produced with one corner's stats invented that way. Those 55 went 42-13 (76.4%) — above our 69.25% headline, so the published record is not inflated by them: strip them out and it reads 226-106 from 332 picks, 68.07%. The fallback was removed on 2026-09-21. Bouts involving a UFC debutant, where our models have no read of their own, carry a pick taken from the betting market, shown without a confidence tier; any other bout without a real read is published with no prediction. Grades stand as published.
Encoding is symmetric by construction, because the raw record is not. The eventual winner is listed in the first corner far more often than chance would put them there — an artefact of how results are recorded, and a label sitting in the column order waiting to be learned.
Three properties remove it. Every feature is a differential — corner A minus corner B — plus context that is invariant to which corner is which. Every bout is presented to the estimators in both orientations during fitting. Predictions are averaged over the two orientations at inference. Corner order therefore carries no information any estimator can exploit, and the positional-inheritance class of defect — a value attached to the wrong fighter — is a build failure rather than a silent bias.
Stated at category level. The feature set itself, its weighting and its parameterisation are not published.
The published probability is not one model's opinion. It is an ensemble of independently trained estimators — gradient-boosted decision trees and regularised linear models — combined by a logistic meta-learner: stacked generalization, with the meta-learner fitted on the members' out-of-sample outputs on a walk-forward schedule so it never sees a member's in-sample optimism.
Ensembling pays only to the extent members' errors are decorrelated, and on this problem they largely are not: measured error correlation across the members is high, because they read the same public record through similar lenses. The honest consequence is that combination buys robustness — insulation from any single estimator degrading, drifting or breaking — considerably more than it buys accuracy. We say that rather than sell the blend as the source of the number.
Bouts with no tradeable price — short-notice bookings, for one — are served by the price-blind path, so a missing price does not by itself leave a bout without a probability. Wide coverage is a product decision, not an accuracy claim: those bouts are genuinely harder to call, and the confidence tier reflects it rather than disguising it. Two cases fall outside that: a bout with a UFC debutant carries the market's pick and no tier, and a bout with no real read at all is published with no prediction.
A pick is a direction. A probability is a claim about frequency, and it is graded on a different axis.
Estimator outputs are mapped to probabilities by a monotone (isotonic-family) transform fitted strictly on past data. Monotone calibration is accuracy-invariant by construction — it cannot reorder picks, only restate their confidence — so a calibration change is a claim about the number, never about who we think wins. Fitting a calibrator on the full sample is a textbook subtle leak, invisible to any feature scan, and the walk-forward schedule below is what precludes it.
Our primary metric is the Brier score — a strictly proper scoring rule, meaning it is minimised only by reporting your true belief, so it cannot be gamed by shading a number toward the side you want to look good on. Accuracy is reported because readers ask for it, and it is the second-order metric: it grades only the sign of the deviation from evenness, and is completely indifferent to whether a genuine coin flip was published at 52% or at 88%. A proper scoring rule is not indifferent, which is why it is the one we optimise and the one we lead with internally.
Never as one number. We take it in its calibration–refinement form — uncertainty − resolution + reliability, after Murphy and DeGroot. Uncertainty is the base-rate variance of the sport (0.245 on the historical cohort in §7); it is fixed, and no forecaster touches it. Resolution is discriminating power — how far a forecaster's conditional outcome rates depart from that base rate. Reliability is calibration error. A forecaster improves by driving reliability toward zero and pushing resolution as far as the available information allows, which is a shorter distance than the literature tends to assume — §7.
The full picture — a reliability diagram, the decile table, Brier and expected calibration error for every graded year and for all of them together, with the live season and the back-tested years labelled apart — is on the calibration page, recomputed from the same graded rows at every build.
Do not read that figure against a floor or a benchmark computed on some other cohort. A Brier score is only interpretable against its own sample: the base rate, the favourite–underdog mix and the price distribution all move it, so two cohorts a season apart are not the same test. The floors quoted in §7 belong to the cohorts they were estimated on, and are not a bar this number is being held against.
A model is only as honest as the test that produced its number. Five things keep ours straight, and each of them is run as a control that is expected to be able to fail.
A forecast that can be edited after the event is not a forecast. Most of the engineering effort here is not in the model at all — it is in making the prediction impossible to quietly revise, and the grading impossible to quietly flatter.
tier provenance · write-once tier freeze · 34 card(s) provisional · 0 with post-freeze drift
Published UFC winner models cluster in the 65–70% band whatever the architecture — logistic regression, boosted ensemble or neural network, market-aware or market-blind. We did not treat that as a plateau waiting for a better feature. We treated the convergence itself as the object of study, and measured it: The Predictability Ceiling of UFC Fight Outcomes, our working paper.
The question is not how do we beat 70%? but why 70? — because the two answers imply opposite research programmes. If the band is a modelling limit, the response is better features and architectures. If it is an aleatoric limit — set by the irreducible randomness of a single fight, where one landed strike ends a bout regardless of who was the better fighter — then further effort spent on winner accuracy is misdirected, and the honest contribution is to characterise the bound.
Two things follow, and both are already on this page. Calibration over accuracy — when accuracy is bounded and nearly saturated it is a weak discriminator between models, and reliability and resolution are the informative axes, which is why §4 leads on a proper scoring rule. And a standing prior: a model reporting materially higher walk-forward accuracy on real fights is extraordinary, and should be treated as a leak until proven otherwise — including when it is ours. That prior is what the battery in §5 exists to serve.
The same prior applies to short-window betting results, and the literature has a clean cautionary case: the one prominent live-betting result in this field was withdrawn by its own author in 2026, after a longer sample showed substantially weaker performance than the eight-week window it was announced on. A short window against a single book is the highest-variance, highest-mirage quantity in this domain. It is exactly what a bound-limited process produces when you look at it briefly, and exactly why nothing on this site is presented as a betting result.
The paper's out-of-sample protocol — freeze the weights, publish before the card, never tune on the forward sample — was a promise kept by discipline. It is now kept by mechanism. Pre-event artifacts are write-once and witnessed by an externally timestamped copy we do not control; every number rendered on this site asserts a content hash against a canonical manifest before the page is written; and grading is entity-keyed and re-derived adversarially from the raw artifacts rather than trusted. That machinery is §6, and it is the substantive difference between a protocol described in a paper and a protocol a reader can check.
If the point prediction is capped, the quantity worth optimising is the honesty of the confidence attached to it. The paper's corollary shows that split-conformal prediction turns a fight model's raw probabilities into nested confidence sets carrying a distribution-free, finite-sample coverage guarantee: thresholds fitted on an earlier block of the frozen season (n=151) and evaluated on a later one (n=102), split chronologically at a card boundary so no future fight can inform a past threshold. Realised accuracy landed inside its target interval at every level, and — unlike a conventional fixed-threshold tiering fitted on the same data — the tiers came out monotone out-of-sample. Each step up the ladder really was a step up.
The confidence ladder in the next section is the product expression of that corollary — the method, applied, with a fifth step the paper's four-set construction does not have: an explicit TOSS-UP floor for bouts no confidence set can separate. A tier is meant to be a confidence set with a number we are held to, not a label, which is the whole reason this site publishes per-tier hit rates instead of a single headline. The rates you see there are the live ones, recomputed from the graded record every build — not the paper's held-out figures, which belong to its own cohort. The paper's two caveats travel with the method regardless: coverage is marginal over the score distribution rather than conditional on weight class or favourite status, and the top tiers are thinly populated, so their floors are not yet provably tight at the sample sizes involved. The method is sound; the high tiers need more fights. The live numbers below are how you watch that happen instead of taking it on faith.
Finally, the bound is conditional on public pre-fight information, and says nothing beyond it. Two directions escape it in principle — signals genuinely orthogonal to the closing price, and prediction during the fight itself, where the relevant market is a different one. Both are open problems. Neither is solved by a better pre-fight winner model, which is precisely the point.
source: our working paper, The Predictability Ceiling of UFC Fight Outcomes · calibration–refinement (Murphy/DeGroot) decomposition, information accounting in bits, split-conformal corollary, nonparametric bootstrap intervals · figures above are the paper's own cohorts, are not comparable with the live-season figures elsewhere on this page, and are superseded for the current season by the public record — which has since grown and been re-graded. The record page is authoritative.
The rest of this page is in plain English, because a well-calibrated probability that nobody reads correctly is a wasted probability. Every pick our models make lands in one of five tiers by how confident the ensemble is. The exception is a bout with a UFC debutant, where our models have no read of their own: that pick is taken from the betting market and carries no tier, and it has its own row in the table below. Higher tier, higher conviction — that is what a tier is: a statement about how tightly the estimators agree, not a promise about the hit rate. LOCK (80.56%, 29-7) leads the live ladder, but the rungs below it are not a clean staircase: MED (77.5%, 93-27) currently out-hits HIGH (73.4%, 69-25). At these sample sizes that is what to expect, and we would rather say so than round it into a story: the 95% Wilson intervals of every neighbouring pair overlap, so no two adjacent tiers are statistically distinguishable yet. Read the gaps between neighbours as noise, not as a ranking. What the season does separate is the two ends of the ladder: LOCK (80.56%, 29-7) against TOSS-UP (48.81%, 41-43) is a real gap, and it is the gap the tier is for. These are the real live numbers (2026-10-10, 35 cards), not a target.
Read it top to bottom: our LOCK and HIGH calls are where we have the most conviction. A TOSS-UP is exactly that — a coin flip, and we label it one. We never say to bet a toss-up; it is shown for honesty, and it stays in the record.
A tiered pick carries both a number and a tier, and they can disagree — so it is worth being explicit that the tier is not simply a band cut out of the percentage. The percentage is the ensemble's point estimate that this fighter wins. The tier is how much conviction that estimate carries, and it also reflects how tightly the independent estimators agree with each other and where the conformal thresholds of §7 fall.
The practical consequence, which surprises people: a pick shown in the mid-70s that the estimators disagree about can sit a tier below a pick in the mid-60s they are unanimous on. That is not a display bug and it is not a typo — it is the ladder doing its job, marking down a confident-looking number that only one read supports. The percentage is the estimate; the tier is the confidence attached to it. If you only look at one of them, look at the tier.
Each bout where both fighters have a record runs through the whole ensemble — several independent reads of the same fight, built on the full public record and the live market.
The meta-learner combines them into a single probability. Where the estimators agree, confidence rises; where they disagree, it falls — honestly.
You get one pick per fight and one number. For every pick our models make, you also get a tier, from a coin flip up to a lock: the number is the estimate and the tier is the conviction behind it. A bout with a UFC debutant carries the market's pick and no tier.
The pick is frozen before the cage door shuts and graded after. Wins, losses and toss-ups all go on the public record.
The headline is the every-fight number, because it is the whole truth: every fight we called in 2026, toss-ups included, nothing dropped.
Not every call is equal. When the independent estimators line up on the same fighter, the pick sharpens — and the record shows it.
Those cohorts answer different questions and must never be blended into a single figure. The 2026 season is the only one published prospectively — frozen before the card, graded after. Earlier seasons are reconstructions of what the pipeline would have said, and are labelled as such wherever they appear.
The parts a methods section is obliged to state plainly. Every number on this page carries its sample size and its cohort; here is what they do — and do not — mean.