How We Score: The JMM Forecast Audit
The audit chain behind every JMM probability: the frozen issue record, the reference comparison, the official settlement source, and the Brier scoreboard.
JMM scores a forecast only after it passes a ten-part publication contract. A binary question, a probability strictly between zero and one, an immutable issue timestamp, a registered model with reproducible inputs, an exact resolution rule, and a named official resolution source are all frozen before the outcome. Accuracy is then measured with proper probability scores: Brier score first, supported by log loss, calibration, and threshold accuracy, and never with invented account P&L.
The publication contract, field by field
Every public JMM forecast must carry ten fields before it can appear: a stable record ID and slug; a binary question and category; a probability strictly between 0 and 1; an immutable issue timestamp; a model name and registered version; reproducible point-in-time inputs, summary, and calculation; an exact resolution rule with deadline, comparison operator, unit, and revision policy; an official resolution source with evidence URLs; an issuance publication proof created before resolution; and result fields left empty until the named source contains the required value, followed by a separate settlement proof.
A refresh may add source observations, issue a forecast for a new input period, or settle an existing call. It may never rewrite the probability, issue time, model inputs, or rule of an already published forecast. That separation is what makes a track record possible.
- Confirm the question is binary, measurable, and time-bounded.
- Check the probability is strictly between zero and one.
- Verify the model name, version, inputs, and calculation are frozen.
- Read the exact comparison operator, unit, deadline, and revision policy.
- Confirm the named official source and the proof that publication preceded resolution.
Four objects every forecast keeps separate
Each forecast distinguishes the JMM probability, the independent estimate produced from the evidence available at the frozen issue time; the reference probability, the strongest approved contemporaneous comparator such as a predeclared base rate or state-persistence model; the disagreement, JMM probability minus the reference in percentage points; and the strength of that disagreement, inspectable diagnostics about sample relevance, uncertainty, source freshness, and model stability.
Disagreement is not automatically an edge, an expected return, or demonstrated skill. The public interface calls it difference until prospective, event-deduplicated scoring establishes that the registered model adds skill against that exact frozen reference. The product remains valuable when JMM agrees with the reference, because its durable value is converting source material into a precise question and preserving the contemporaneous evidence.
- Find the JMM probability and the reference probability.
- Subtract one from the other to get the disagreement in points.
- Read the diagnostics: sample relevance, uncertainty, and source freshness.
- Treat a large disagreement as a question, not a conclusion.
The Brier-first scoreboard
Accuracy is evaluated with proper probability scores, led by Brier score and supported by log loss, calibration, and threshold accuracy. A Brier score measures the squared difference between the probability and the binary outcome, so it rewards calibrated confidence rather than confident hits alone.
Probability forecasts are evaluated across a set, not by celebrating one confident call. JMM publishes performance only after eligible public calls have settled, and none has yet: every record on the ledger is open. Historical replays never backfill it. The evaluation protocol behind the models was fixed on August 3, 2026 and opens its prospective window on September 1, 2026, and a case counts toward that window only if its issuance carries the protocol hash, its issue proof precedes resolution, and its settlement carries a separate proof. Everything replayed before that date is development evidence and is labelled as such where it appears.
- Wait for settlement evidence from the named official source.
- Score the set with Brier, log loss, and calibration.
- Compare against the predeclared reference, not a post-hoc benchmark.
- Check the correction record before trusting a headline.
What can never enter the record
Generation time never substitutes for publication time, and a JMM record keeps the two apart along with four other clocks: when the publisher released a figure, which period it describes, when JMM first received it, when the probability was frozen, and when the outcome became provable. A projection checked by an automated pipeline is not an admitted forecast: admission requires the issue bytes committed at a content address, with the publication instant falling at or after issuance and strictly before the cutoff. Both deadlines are rules rather than judgement calls. A company question is deadlined at the baseline quarter end plus two hundred calendar days; a statistical question is deadlined at the release time the publisher itself put on its calendar.
Retrospective replay is development evidence, not a public track record. A later data revision does not rewrite an earlier score; the initial published source value controls when the rule says so, and revisions stay visible as context. The same discipline applies to people: attribution is documented without becoming a reputation score, and a claim with no measurable rule stays under review.
- Confirm the issue proof predates the cutoff.
- Confirm the settlement proof is separate from the issue proof.
- Check whether the rule scores the initial release or a later revision.
- Treat any invented P&L or retroactive backfill as outside the record.
Get the Pro launch email
This guide stays free. Join the list for one email when the forecast feed, machine-readable research files, alerts, and saved research workflows open.