Verified accuracy

We check our own homework.

Every forecast ClearSkys makes is recorded, then checked against what the sky actually did. No hand-picked examples, no marketing claims — the numbers below are computed from every verified night and update automatically as new nights are checked.

61.4%
of next-night forecasts within 10 points of the verified actual
98.5%
of the nights we called good turned out good (5,558 go calls, one day ahead)
26,886
city-nights verified since 2026-06-21

For next-night forecasts, 7,080 out of 7,742 verified nights landed within one quality band of what the sky actually did.

And when the score does miss, it nearly always errs on the side of caution. It predicts cloud that later clears, rather than promising a clear night that clouds over. You are far more likely to be pleasantly surprised than disappointed.

Accuracy by lead time

Weather models sharpen as the night approaches, and the score inherits that. Use the 7-day view to spot promising nights, and trust the 1-day score for the go/no-go call.

Forecast madeNights verifiedAvg score error BiasWithin ±10 ptsWithin ±15 ptsAvg cloud error
Same day 6,280 ±10.4 pts +8.4 60.4% 73.7% ±17.7%
1 day ahead 7,742 ±10.5 pts +7.4 61.4% 74.7% ±19.3%
3 days ahead 7,479 ±12.0 pts +7.8 57.5% 70.2% ±23.2%
7 days ahead 5,385 ±14.7 pts +7.5 51.3% 63.0% ±29.4%
VERIFIED NIGHTS 2026-06-21 → 2026-08-20 · UPDATED DAILY AS ERA5 DATA BECOMES AVAILABLE

Bias is the average signed difference (verified minus predicted) — negative means the forecast ran optimistic, positive means nights turned out better than promised. The same-day row is a baseline, not a forecast: it measures the built-in disagreement between the forecast weather source and the ERA5 verification record. Genuine forecast skill is how little the other rows exceed it.

Part of the positive bias is bookkeeping, not forecast error: predicted scores are docked when the forecast expects haze, but the ERA5 record has no visibility data, so verified scores are never docked the same way. On the 5,296 most recent next-night forecasts where this is measurable, removing that one-way penalty leaves an underlying bias of +3.2.

How often does each verdict hold up?

Every forecast gets one of four verdicts: Poor, Mixed, Good or Excellent. Here is how the next night actually turned out for each one, checked against real weather records.

Poor 691 nights
Turned out · Poor 40% · Mixed 35% · Good 14% · Excellent 11%

A Poor call can only be beaten. 40 in 100 stayed Poor, 35 improved to Mixed, and 25 in 100 cleared all the way to Good or better.

Mixed 1,140 nights
Turned out · Poor 6% · Mixed 27% · Good 31% · Excellent 36%

27 in 100 stayed Mixed, 67 in 100 improved to Good or Excellent, and only 6 in 100 dropped to Poor.

Good 1,272 nights
Turned out · Poor 1% · Mixed 8% · Good 21% · Excellent 71%

92 in 100 turned out Good or better — a Good call was more likely to be beaten (71 in 100 verified Excellent) than to fall short (9 in 100).

Excellent 4,639 nights
Turned out · Poor 0% · Mixed 1% · Good 3% · Excellent 95%

95 in 100 verified exactly Excellent, and 98 in 100 Good or better.

The misses run one way. When a Poor, Mixed or Good verdict is wrong, the night nearly always turned out better than called, not worse — even the Poor row fails upward, toward clear skies. Excellent, the only verdict that can miss downward, is also the one that almost never misses. The cost of a ClearSkys miss is a missed opportunity, not a wasted trip.

The go/no-go question

Point error is abstract; what matters is whether the score would have made the right call. So we also check the decisions the day-before score implied:

98.5%
of nights scored 70+ the day before were verified 60+ (5,558 go calls)
57.4%
of nights scored below 40 the day before were verified below 50 (502 no-go calls)

THRESHOLDS ARE FIXED AND STATED, NOT TUNED: GO = PREDICTED ≥70 VERIFIED ≥60 · NO-GO = PREDICTED <40 VERIFIED <50

The two numbers are not symmetric, and that is the cautious skew again: a go call that misses means a promised night let you down, and that almost never happens. A no-go call that misses means the sky beat the forecast — the night you skipped turned out fine. We would rather miss in that direction.

How verification works

01
Record the prediction. Every day, the score for each tracked location is computed and stored for the same-day night and the nights 1, 3, and 7 days ahead — exactly as the app would have shown it, using the same scoring engine.
02
Let the night happen. Nothing is touched until the night has passed and the ERA5 reanalysis archive — a globally complete, satellite-and-observation-derived weather record — has caught up, about six days later.
03
Rescore on reality. The actual hourly conditions for that night are fed through the identical scoring engine. Same code, same weights — the only difference is forecast inputs versus what actually happened.
04
Publish the difference. The gap between predicted and actual score is stored per night, per location, per lead time, and aggregated into the table above. Bad nights count the same as good ones.

Honest caveats

ERA5 is reanalysis, not a telescope. Ground truth here is the best available global weather record, not a human standing under the sky. It is spatially complete and independent of the forecast models we score with, but it is still a model of what happened.

Moonlight is held constant. The moon is identical on both sides of each comparison (same night, same place), so it cancels out. These numbers measure the weather component of the score — which is the part a forecast can get wrong.

Rain is the weakest comparison. Forecasts express rain as a probability; the archive records what fell. Cloud cover — by far the largest factor in the score — compares directly.

The haze penalty is one-way. Predicted scores are docked when the forecast expects poor transparency, but ERA5 records no visibility — so verified scores are never docked the same way. Part of the positive bias above is therefore our own deliberate caution being marked against us, not forecast error. Where measurable, the underlying bias with that penalty removed is shown under the lead-time table.

Directional cloud on satellite passes

Everything above measures the night score, which describes the sky as a whole. Satellite passes now carry a separate, narrower estimate: cloud along the specific bearing and altitude the pass crosses, rather than across the whole sky. It appears on a pass only when it disagrees with the ordinary forecast.

This estimate is not yet verified

It is not in the numbers above. Every figure on the rest of this page is checked against ERA5. The directional estimate is not, because it has not been running long enough to accumulate the cases that would test it.

The comparison is awkward by nature. Reanalysis is gridded at your location the same way the forecast is, so it records what the sky did overhead rather than along a bearing. The nights that matter are the ones where the directional and whole-sky readings disagree, and those are uncommon by design, so it will take time either way.

A grid square is not a sight line. A cell reporting 40% low cloud does not say whether your particular line of sight falls in the covered 40%. This sharpens a probability; it does not produce a certainty.

2,052
Passes assessed
17.4%
Differed from whole-sky
1.4%
Grid too coarse

Counts only. These say how often the estimate differs from the whole-sky forecast, and how often the model grid was fine enough to support it. Neither says whether it was right. When there is enough data to answer that, the result goes here whichever way it falls.

How the directional estimate works, and what it cannot tell you

Questions

What does “within 10 points” actually mean?

The score runs 0–100 and 10 points is roughly one quality band — the difference between a good night and a great one. A forecast within 10 points of the verified actual would have led you to the same decision about setting up.

Why publish this at all?

Because every forecast site claims to be accurate and almost none show their working. ClearSkys exists to answer “is tonight worth it?” — and that answer is only useful if you know how much to trust it at each lead time.

How is the score itself calculated?

Cloud cover across three altitude layers, moonlight and its overlap with darkness, transparency, wind, and more — the full breakdown is in How the Stargazing Score Works.

See tonight's score for your location →