We check our own homework.
Every forecast ClearSkys makes is recorded, then checked against what the sky actually did. No hand-picked examples, no marketing claims — the numbers below are computed from every verified night and update automatically as new nights are checked.
For next-night forecasts, 7,080 out of 7,742 verified nights landed within one quality band of what the sky actually did.
And when the score does miss, it nearly always errs on the side of caution. It predicts cloud that later clears, rather than promising a clear night that clouds over. You are far more likely to be pleasantly surprised than disappointed.
Accuracy by lead time
Weather models sharpen as the night approaches, and the score inherits that. Use the 7-day view to spot promising nights, and trust the 1-day score for the go/no-go call.
| Forecast made | Nights verified | Avg score error | Bias | Within ±10 pts | Within ±15 pts | Avg cloud error |
|---|---|---|---|---|---|---|
| Same day | 6,280 | ±10.4 pts | +8.4 | 60.4% | 73.7% | ±17.7% |
| 1 day ahead | 7,742 | ±10.5 pts | +7.4 | 61.4% | 74.7% | ±19.3% |
| 3 days ahead | 7,479 | ±12.0 pts | +7.8 | 57.5% | 70.2% | ±23.2% |
| 7 days ahead | 5,385 | ±14.7 pts | +7.5 | 51.3% | 63.0% | ±29.4% |
Bias is the average signed difference (verified minus predicted) — negative means the forecast ran optimistic, positive means nights turned out better than promised. The same-day row is a baseline, not a forecast: it measures the built-in disagreement between the forecast weather source and the ERA5 verification record. Genuine forecast skill is how little the other rows exceed it.
Part of the positive bias is bookkeeping, not forecast error: predicted scores are docked when the forecast expects haze, but the ERA5 record has no visibility data, so verified scores are never docked the same way. On the 5,296 most recent next-night forecasts where this is measurable, removing that one-way penalty leaves an underlying bias of +3.2.
How often does each verdict hold up?
Every forecast gets one of four verdicts: Poor, Mixed, Good or Excellent. Here is how the next night actually turned out for each one, checked against real weather records.
A Poor call can only be beaten. 40 in 100 stayed Poor, 35 improved to Mixed, and 25 in 100 cleared all the way to Good or better.
27 in 100 stayed Mixed, 67 in 100 improved to Good or Excellent, and only 6 in 100 dropped to Poor.
92 in 100 turned out Good or better — a Good call was more likely to be beaten (71 in 100 verified Excellent) than to fall short (9 in 100).
95 in 100 verified exactly Excellent, and 98 in 100 Good or better.
The misses run one way. When a Poor, Mixed or Good verdict is wrong, the night nearly always turned out better than called, not worse — even the Poor row fails upward, toward clear skies. Excellent, the only verdict that can miss downward, is also the one that almost never misses. The cost of a ClearSkys miss is a missed opportunity, not a wasted trip.
The go/no-go question
Point error is abstract; what matters is whether the score would have made the right call. So we also check the decisions the day-before score implied:
THRESHOLDS ARE FIXED AND STATED, NOT TUNED: GO = PREDICTED ≥70 VERIFIED ≥60 · NO-GO = PREDICTED <40 VERIFIED <50
The two numbers are not symmetric, and that is the cautious skew again: a go call that misses means a promised night let you down, and that almost never happens. A no-go call that misses means the sky beat the forecast — the night you skipped turned out fine. We would rather miss in that direction.
How verification works
Honest caveats
ERA5 is reanalysis, not a telescope. Ground truth here is the best available global weather record, not a human standing under the sky. It is spatially complete and independent of the forecast models we score with, but it is still a model of what happened.
Moonlight is held constant. The moon is identical on both sides of each comparison (same night, same place), so it cancels out. These numbers measure the weather component of the score — which is the part a forecast can get wrong.
Rain is the weakest comparison. Forecasts express rain as a probability; the archive records what fell. Cloud cover — by far the largest factor in the score — compares directly.
The haze penalty is one-way. Predicted scores are docked when the forecast expects poor transparency, but ERA5 records no visibility — so verified scores are never docked the same way. Part of the positive bias above is therefore our own deliberate caution being marked against us, not forecast error. Where measurable, the underlying bias with that penalty removed is shown under the lead-time table.
Directional cloud on satellite passes
Everything above measures the night score, which describes the sky as a whole. Satellite passes now carry a separate, narrower estimate: cloud along the specific bearing and altitude the pass crosses, rather than across the whole sky. It appears on a pass only when it disagrees with the ordinary forecast.
This estimate is not yet verified
It is not in the numbers above. Every figure on the rest of this page is checked against ERA5. The directional estimate is not, because it has not been running long enough to accumulate the cases that would test it.
The comparison is awkward by nature. Reanalysis is gridded at your location the same way the forecast is, so it records what the sky did overhead rather than along a bearing. The nights that matter are the ones where the directional and whole-sky readings disagree, and those are uncommon by design, so it will take time either way.
A grid square is not a sight line. A cell reporting 40% low cloud does not say whether your particular line of sight falls in the covered 40%. This sharpens a probability; it does not produce a certainty.
Counts only. These say how often the estimate differs from the whole-sky forecast, and how often the model grid was fine enough to support it. Neither says whether it was right. When there is enough data to answer that, the result goes here whichever way it falls.
How the directional estimate works, and what it cannot tell you
Questions
What does “within 10 points” actually mean?
The score runs 0–100 and 10 points is roughly one quality band — the difference between a good night and a great one. A forecast within 10 points of the verified actual would have led you to the same decision about setting up.
Why publish this at all?
Because every forecast site claims to be accurate and almost none show their working. ClearSkys exists to answer “is tonight worth it?” — and that answer is only useful if you know how much to trust it at each lead time.
How is the score itself calculated?
Cloud cover across three altitude layers, moonlight and its overlap with darkness, transparency, wind, and more — the full breakdown is in How the Stargazing Score Works.