Accuracy
Our forecasts, graded against pilot reports
Most forecast products quote the accuracy of the model they build on. We grade our own forecasts instead: after a flight or route check completes, we automatically compare what we briefed against the official pilot reports (PIREPs) filed along that route while it was flying - and publish the running record here, including the misses.
This page is generated from our verification database and refreshes hourly. Nothing on it is hand-picked: every scored check since the loop went live is counted, and checks where no pilot report could be matched are reported as exactly that - not quietly dropped, and never counted as wins.
The record so far
The live numbers are temporarily unavailable - this page refreshes automatically, so they will reappear shortly. The methodology below is the standard every forecast is held to.
How a check is scored
For each completed flight or route check we reconstruct the flight path and collect every pilot report filed within 45 minutes of overflight, within about 50 km of the path, and within 3,000 ft of cruise altitude (reports filed without an altitude are matched without that vertical filter). Each report is compared with the strongest intensity band we forecast anywhere within that window - deliberately generous to the report, because pilot positions carry real error - and the flight gets one verdict:
- Confirmed - every matched report landed within one intensity band of the briefing. Forecasts are bands rather than points, and pilot-report intensity is subjective, so one band is the honest tolerance.
- Bumpier in places - at least one report came in two or more bands above what we briefed. These are the misses we most want to count.
- Smoother than briefed - the air ran at least two bands calmer than the briefing, with nothing two or more bands above it. We'd rather brief a bump that never arrives than the reverse - but it still counts as a miss, not a win.
- No data - no pilot report could be matched to the flight window: none were filed nearby, or those filed were too far off the path, at a different altitude, or filed before our forecast was computed. Common on quieter corridors; reported honestly rather than scored.
Two kinds of checks feed the pool. Paid briefings are graded on the final pre-flight forecast we emailed - the like-for-like test of the product. Daily route checks grade a forecast computed about 30 minutes before departure on popular US routes, which measures short-lead accuracy - the easier end of the forecast problem, and we label it as such rather than blending it silently.
One more rule keeps the grading fair: pilot reports filed before the forecast was computed are excluded from scoring. Our live forecasts incorporate recent pilot reports, so without that cut-off a report could inform a forecast and then be counted as evidence the forecast was right. This rule was added in August 2026; forecasts computed before then are graded without it (or with an approximate cut-off), so the older part of the corpus can, in a narrow window around departure, have matched a report the forecast had already seen - a small optimistic bias that washes out as new checks accrue.
The strict scores
Meteorologists score each matched report against the forecast with no tolerance bands, excluding light-intensity reports (pilots agree on those only about two-thirds of the time). Matching still uses the 50 km / 45-minute window above - the standard tolerance for pilot position error - but the intensity comparison is exact, which is deliberately tougher than the flight-level verdicts, so the numbers read lower. Publishing both is the point.
Scores publish automatically once scored point-pairs accrue.
For context on what the field achieves: NOAA's published verification for its operational guidance reports probability of detection around 70-85% for moderate-or-greater turbulence with a false-alarm ratio of 25-35%, and roughly 0.85 AUC is the practical ceiling nobody - human or model - meaningfully beats. Those are the upstream model's numbers, not ours; this page exists so you can see ours.
Read more
The full picture of where the forecast is strong and where it is genuinely hard is in our honest accuracy breakdown. For how this compares with other turbulence tools, see Turbli vs Turbcast. Or just check your flight - the forecast this page grades is free, with no signup.
Updated hourly from the verification database.
