Which weather forecast is right most often?
DMI, MET Norway, OpenWeatherMap and the pilots’ own airport forecast, for five Danish cities. Four times a day every forecast is saved and locked in a hash chain before the weather happens. Then it is scored against what DMI’s weather stations actually measured.
- Rules published before the first forecast
- Anchored on GitHub every day
- No results until there is something real to show
Reading the public anchor file on GitHub…
What is being locked
Four sources, five cities, four times a day, up to 72 hours ahead. Measurements come in once a night.
Three apps and a forecaster
- DMI, the Danish Meteorological Institute: its own forecast model for each city.
- MET Norway: the Norwegian Meteorological Institute’s forecast, used by many weather apps.
- OpenWeatherMap: the free 3-hour forecast used by thousands of sites and apps.
- TAF: the airport forecast pilots use, written by a forecaster on duty. Scored at Aalborg and Copenhagen, where DMI measures at the airport itself.
Every source also has to beat a lazy guess: the most recent measurement at the same time of day, known when the forecast was made. “Tomorrow will be like today.”
The nearest active DMI station
| City | Temperature | Rain |
|---|---|---|
| Hjørring | 06033 Hirtshals | 05005 Uggerby |
| Aalborg TAF EKYT | 06030 Aalborg airport | same |
| Aarhus | 06072 Ødum | same |
| Odense | 06126 Årslev | same |
| Copenhagen TAF EKCH | 06180 Copenhagen airport | same |
A station counts only if it has delivered data in the last 24 hours. When no station within 20 km measures both, rain comes from the gauge nearest the temperature station.
The rules, fixed in advance
Published at 07:57 UTC on 28 September, four hours before the first forecast was collected. They don’t change while the test runs.
How many degrees off?
Mean absolute error and warm/cold bias every third hour, split by how far ahead the forecast looked: 0–24, 24–48 and 48–72 hours. Only at moments where every source had a forecast.
Did it see it coming?
Six-hour periods. “Rain” means at least 0.2 mm was measured. Hits, misses and false alarms per source, each summed over its own time steps.
Does it rain 3 times in 10?
Calibration and Brier score for each source’s rain probability, next 48 hours. The pilots’ TAF is translated to probabilities by the ICAO rulebook, fixed before the start.
Same moments, same yardstick
A failed download stays in the chain as a gap and is never filled in. If DMI corrects a measurement later, the first value is kept and the correction logged. DMI’s trace code (−0.1 mm) counts as no rain.
What to keep in mind
DMI is both a participant and the source of the measurements. OpenWeatherMap is scored on its free forecast, not a paid product. The first 30 days are marked preliminary. Nothing here is a verdict on anyone’s forecasters.
Check it without trusting us
Three things anyone can verify today, before a single result is published.
The rules came first
Release v1.1 is the commit with the collector and the scoring rules, four hours before the first forecast. The code’s SHA-256 is in METHOD.lock.
Release v1.1 →History can’t be rewritten
Every day the newest hash of the chain is committed publicly. Change any earlier forecast and the chain no longer leads to a head that is already public.
The anchor file →Run it yourself
The collector is open source, Python standard library only, with offline tests. The same code scores the results that will be published here.
Read the code →Which one do you think wins?
DMI at home, the Norwegians, the free app, or the pilots’ forecaster? Seal your guess now with Lock your prediction: it’s sealed in your own browser, and you post the seal wherever you like. When the first scoreboard appears around 28 October, reveal it.
- 28 Sep 2026 · 07:57 UTCMethod and code published. Scoring rules, thresholds and baseline fixed in public.
- 28 Sep 2026 · 12:08 UTCFirst forecasts locked. All sources answered.
- 29 Sep 2026First public anchor. The chain head committed to GitHub, then every day.
- Now · day –Collecting. Forecasts four times a day, measurements every night, the chain checked every night.
- Around 28 Oct 2026First scoreboard, marked preliminary.
- November–December 2026Results after 60–90 days.
This is the same method we use to validate trading signals: lock it before the outcome, measure it against data we don’t control, and compare it with a lazy guess.