Data Playground / Experiment 01
Experiment 01 · running since 28 Sep 2026

Which weather forecast is right most often?

DMI, MET Norway, OpenWeatherMap and the pilots’ own airport forecast, for five Danish cities. Four times a day every forecast is saved and locked in a hash chain before the weather happens. Then it is scored against what DMI’s weather stations actually measured.

  • Rules published before the first forecast
  • Anchored on GitHub every day
  • No results until there is something real to show
LIVE STATUS · FROM THE PUBLIC ANCHORSLOADING…
–
day of the test
–
downloads locked in the chain, failed ones kept as gaps
–
chain, recomputed from the first download
– days
to the first scoreboard, around 28 October
Latest public anchor– · –
Code the server runs–

Reading the public anchor file on GitHub…

01

What is being locked

Four sources, five cities, four times a day, up to 72 hours ahead. Measurements come in once a night.

THE SOURCES

Three apps and a forecaster

  • DMI, the Danish Meteorological Institute: its own forecast model for each city.
  • MET Norway: the Norwegian Meteorological Institute’s forecast, used by many weather apps.
  • OpenWeatherMap: the free 3-hour forecast used by thousands of sites and apps.
  • TAF: the airport forecast pilots use, written by a forecaster on duty. Scored at Aalborg and Copenhagen, where DMI measures at the airport itself.

Every source also has to beat a lazy guess: the most recent measurement at the same time of day, known when the forecast was made. “Tomorrow will be like today.”

THE CITIES · MEASURED AT

The nearest active DMI station

CityTemperatureRain
Hjørring06033 Hirtshals05005 Uggerby
Aalborg TAF EKYT06030 Aalborg airportsame
Aarhus06072 Ødumsame
Odense06126 Årslevsame
Copenhagen TAF EKCH06180 Copenhagen airportsame

A station counts only if it has delivered data in the last 24 hours. When no station within 20 km measures both, rain comes from the gauge nearest the temperature station.

02

The rules, fixed in advance

Published at 07:57 UTC on 28 September, four hours before the first forecast was collected. They don’t change while the test runs.

TEMPERATURE

How many degrees off?

Mean absolute error and warm/cold bias every third hour, split by how far ahead the forecast looked: 0–24, 24–48 and 48–72 hours. Only at moments where every source had a forecast.

RAIN OR NO RAIN

Did it see it coming?

Six-hour periods. “Rain” means at least 0.2 mm was measured. Hits, misses and false alarms per source, each summed over its own time steps.

“30 % CHANCE OF RAIN”

Does it rain 3 times in 10?

Calibration and Brier score for each source’s rain probability, next 48 hours. The pilots’ TAF is translated to probabilities by the ICAO rulebook, fixed before the start.

FAIR BY DESIGN

Same moments, same yardstick

A failed download stays in the chain as a gap and is never filled in. If DMI corrects a measurement later, the first value is kept and the correction logged. DMI’s trace code (−0.1 mm) counts as no rain.

HONEST ABOUT LIMITS

What to keep in mind

DMI is both a participant and the source of the measurements. OpenWeatherMap is scored on its free forecast, not a paid product. The first 30 days are marked preliminary. Nothing here is a verdict on anyone’s forecasters.

BEFORE THE SCOREBOARD · YOUR TURN

Which one do you think wins?

DMI at home, the Norwegians, the free app, or the pilots’ forecaster? Seal your guess now with Lock your prediction: it’s sealed in your own browser, and you post the seal wherever you like. When the first scoreboard appears around 28 October, reveal it.

  1. 28 Sep 2026 · 07:57 UTCMethod and code published. Scoring rules, thresholds and baseline fixed in public.
  2. 28 Sep 2026 · 12:08 UTCFirst forecasts locked. All sources answered.
  3. 29 Sep 2026First public anchor. The chain head committed to GitHub, then every day.
  4. Now · day –Collecting. Forecasts four times a day, measurements every night, the chain checked every night.
  5. Around 28 Oct 2026First scoreboard, marked preliminary.
  6. November–December 2026Results after 60–90 days.

This is the same method we use to validate trading signals: lock it before the outcome, measure it against data we don’t control, and compare it with a lazy guess.