Kalshi's CFTC-regulated weather markets settle daily on official National Weather Service station observations, making them the first retail-accessible, continuously priced instrument written directly on local U.S. weather.[1] The paper examines a precise question: whether these contracts can transfer weather risk rather than predict it — a materially different use case than the one most retail participants assume.
The dataset covers 287,909 price observations from 15,994 resolved contracts.[1] The first finding is a null result that reframes everything that follows: simple univariate post-hoc recalibration — both isotonic regression and Platt scaling — does not improve probabilistic accuracy of raw prices out of sample.[1] The market's prices are not systematically correctable by standard machine learning calibration methods. That null result shifts the research question from price correction toward hedging effectiveness — which is where the structurally significant findings are.
The binding constraint on contract tradability is not market calibration — it is geography.[1] Residual basis risk between each contract's fixed settlement-station degree-day index and a customer-proxy local exposure index was measured across 30 station pairs and two trigger types over 24 years. Mean residual basis risk is 0.386 — large in absolute terms — but the key finding is that it is predictable: it rises systematically with station distance.[1]
Kalshi settles each market on one specific weather station, not the city. A forecast pulled for the city center can sit approximately 1°F off the settlement station — and because the brackets settle to an integer °F, approximately 1°F is often enough to flip which side of the contract wins.[3] The settlement source is defined in the market rules. Daily high and low temperature markets settle on the final climate report issued by the National Weather Service, typically released the following morning.[2]
Basis risk is systematically lower for heating-degree-day triggers than cooling-degree-day triggers.[1] Multi-station portfolios beat the best single station out of sample in only 43% of cases, and ridge shrinkage does not change that result.[1] Static hedge constructions do not reliably reduce the dominant source of error.
The implied-temperature bias — market-implied minus realized settlement value — is approximately zero across the full dataset.[1] Although the realized temperature escapes Kalshi's 8°F bin ladder on 17% of days, the binarization cost of synthesizing a degree-day claim from binary bins is approximately 1% of exposure variance — an order of magnitude below geographic basis risk.[1] The market is not meaningfully mispriced relative to its own settlement structure. The error is in the gap between the settlement station and the exposure being hedged.
A separate analysis of 8,494 historical settled weather markets from Kalshi's KXHIGHNY series — NYC daily high temperature — examined systematic miscalibration in market prices directly.[4] The finding is consistent with Luo's null: standard recalibration approaches do not produce reliable out-of-sample improvement. What the market prices encode is not a correctable bias — it is collective uncertainty about a geographically fixed measurement.
Weather markets dominate Kalshi's settlement volume. In an analysis benchmarking AI models on prediction market performance, weather markets accounted for 71–97% of each model's settled contracts, reflecting the high frequency of short-duration temperature contracts.[5] Performance on weather markets is the primary driver of overall settlement win rates across all model types tested.
The paper establishes that the dominant source of pricing error in Kalshi weather contracts is not market inefficiency in the traditional sense — it is the structural gap between settlement station geography and real-world exposure geography. That gap is measurable, predictable, and correlated with station distance. It is not eliminated by portfolio construction or shrinkage estimators under static hedge assumptions.[1]
For a systematic strategy, this produces a testable hypothesis: if geographic basis risk is the dominant and predictable source of variance, then stations with the highest basis risk relative to nearby population centers are the contracts where combination asymmetry — between NOAA observation distributions and Kalshi implied probability — is most likely to persist. The edge is not in recalibrating prices. It is in selecting the right station pairs and trigger types before entry.
The retail accessibility of these contracts — regulated by the CFTC, continuously priced, settled on official NWS data — means the infrastructure for systematic participation exists. The constraint is the analytical layer between settlement station topology and exposure mapping, not market access.[1]
ugly problem.
Direct intake for difficult information problems. Physical markets, procurement, pricing, supply chain, entity resolution, litigation support. Describe the problem. We scope the engagement.