We check our own homework.
Every forecast ClearSkys makes is recorded, then checked against what the sky actually did. No hand-picked examples, no marketing claims — the numbers below are computed from every verified night and update automatically as new nights are checked.
For next-night forecasts, 14,025 out of 15,213 verified nights landed within one quality band of what the sky actually did.
And when the score does miss, it nearly always errs on the side of caution. It predicts cloud that later clears, rather than promising a clear night that clouds over. You are far more likely to be pleasantly surprised than disappointed.
Accuracy by lead time
Weather models sharpen as the night approaches, and the score inherits that. Use the 7-day view to spot promising nights, and trust the 1-day score for the go/no-go call.
| Forecast made | Nights verified | Avg score error | Bias | Within ±10 pts | Within ±15 pts | Avg cloud error |
|---|---|---|---|---|---|---|
| Same day | 13,751 | ±10.2 pts | +7.8 | 61.0% | 75.1% | ±16.4% |
| 1 day ahead | 15,213 | ±10.4 pts | +7.0 | 61.0% | 74.8% | ±18.3% |
| 3 days ahead | 14,950 | ±12.4 pts | +7.8 | 55.3% | 68.9% | ±22.3% |
| 7 days ahead | 12,856 | ±16.0 pts | +7.7 | 47.3% | 59.9% | ±29.6% |
Bias is the average signed difference (verified minus predicted) — negative means the forecast ran optimistic, positive means nights turned out better than promised. The same-day row is a baseline, not a forecast: it measures the built-in disagreement between the forecast weather source and the ERA5 verification record. Genuine forecast skill is how little the other rows exceed it.
Part of the positive bias is bookkeeping, not forecast error: predicted scores are docked when the forecast expects haze, but the ERA5 record has no visibility data, so verified scores are never docked the same way. On the 12,767 most recent next-night forecasts where this is measurable, removing that one-way penalty leaves an underlying bias of +2.9.
How often does each verdict hold up?
Every forecast gets one of four verdicts: Poor, Mixed, Good or Excellent. Here is how the next night actually turned out for each one, checked against real weather records.
A Poor call can only be beaten. 47 in 100 stayed Poor, 34 improved to Mixed, and 19 in 100 cleared all the way to Good or better.
28 in 100 stayed Mixed, 64 in 100 improved to Good or Excellent, and only 7 in 100 dropped to Poor.
90 in 100 turned out Good or better — a Good call was more likely to be beaten (65 in 100 verified Excellent) than to fall short (10 in 100).
94 in 100 verified exactly Excellent, and 99 in 100 Good or better.
The misses run one way. When a Poor, Mixed or Good verdict is wrong, the night nearly always turned out better than called, not worse — even the Poor row fails upward, toward clear skies. Excellent, the only verdict that can miss downward, is also the one that almost never misses. The cost of a ClearSkys miss is a missed opportunity, not a wasted trip.
The go/no-go question
Point error is abstract; what matters is whether the score would have made the right call. So we also check the decisions the day-before score implied:
THRESHOLDS ARE FIXED AND STATED, NOT TUNED: GO = PREDICTED ≥70 VERIFIED ≥60 · NO-GO = PREDICTED <40 VERIFIED <50
The two numbers are not symmetric, and that is the cautious skew again: a go call that misses means a promised night let you down, and that almost never happens. A no-go call that misses means the sky beat the forecast — the night you skipped turned out fine. We would rather miss in that direction.
How verification works
Honest caveats
ERA5 is reanalysis, not a telescope. Ground truth here is the best available global weather record, not a human standing under the sky. It is spatially complete and independent of the forecast models we score with, but it is still a model of what happened.
Moonlight is held constant. The moon is identical on both sides of each comparison (same night, same place), so it cancels out. These numbers measure the weather component of the score — which is the part a forecast can get wrong.
Rain is the weakest comparison. Forecasts express rain as a probability; the archive records what fell. Cloud cover — by far the largest factor in the score — compares directly.
The haze penalty is one-way. Predicted scores are docked when the forecast expects poor transparency, but ERA5 records no visibility — so verified scores are never docked the same way. Part of the positive bias above is therefore our own deliberate caution being marked against us, not forecast error. Where measurable, the underlying bias with that penalty removed is shown under the lead-time table.
Directional cloud on satellite passes
Everything above measures the night score, which describes the sky as a whole. Satellite passes now carry a separate, narrower estimate: cloud along the specific bearing and altitude the pass crosses, rather than across the whole sky. It appears on a pass only when it disagrees with the ordinary forecast.
This estimate is not yet verified
It is not in the numbers above. Every figure on the rest of this page is checked against ERA5. The directional estimate is not, because it has not been running long enough to accumulate the cases that would test it.
The comparison is awkward by nature. Reanalysis is gridded at your location the same way the forecast is, so it records what the sky did overhead rather than along a bearing. The nights that matter are the ones where the directional and whole-sky readings disagree, and those are uncommon by design, so it will take time either way.
A grid square is not a sight line. A cell reporting 40% low cloud does not say whether your particular line of sight falls in the covered 40%. This sharpens a probability; it does not produce a certainty.
Counts only. These say how often the estimate differs from the whole-sky forecast, and how often the model grid was fine enough to support it. Neither says whether it was right. When there is enough data to answer that, the result goes here whichever way it falls.
How the directional estimate works, and what it cannot tell you
Sky darkness
Forecast pages show how dark the sky is at a location, as a band such as “Rural” or “City sky” with an approximate Bortle range and an SQM value (sky brightness in magnitudes per square arcsecond; higher is darker). It is a property of the place, not of the night, so it is shown alongside the forecast score and never changes it.
We estimate it ourselves, empirically rather than by simulating the atmosphere. Starting from NASA’s satellite measurements of light emitted at night, we add up the light from every lit area within 250 km, weighted by distance, and calibrate the result against thousands of sky-brightness readings taken on the ground by Globe at Night volunteers. The model is rebuilt each year from the latest satellite composite. The full story, including how it was calibrated and why we built our own, is in the guide How dark is your sky?
How accurate it is
Typically within about 0.4 SQM (median 0.3) of ground measurements it was not fitted to. That places most locations in the right band, but it cannot separate Bortle 1, 2 and 3, so the darkest places share one band.
City centres are probably brighter than shown, by around one magnitude. They still fall in the “City sky” band. Satellites under-count some modern LED lighting, and readings taken in town centres often include glare from nearby lamps.
Local conditions vary. A hill, a tree line or a single nearby light can make your own spot noticeably darker or brighter than the surrounding square kilometre.
Night lights: NASA Black Marble VNP46A4 (VIIRS/NPP Lunar BRDF-Adjusted Nighttime Lights Yearly), doi:10.5067/VIIRS/VNP46A4.002. Ground calibration: Globe at Night, licensed CC BY 4.0.
Questions
What does “within 10 points” actually mean?
The score runs 0–100 and 10 points is roughly one quality band — the difference between a good night and a great one. A forecast within 10 points of the verified actual would have led you to the same decision about setting up.
Why publish this at all?
Because every forecast site claims to be accurate and almost none show their working. ClearSkys exists to answer “is tonight worth it?” — and that answer is only useful if you know how much to trust it at each lead time.
How is the score itself calculated?
Cloud cover across three altitude layers, moonlight and its overlap with darkness, transparency, wind, and more — the full breakdown is in How the Stargazing Score Works.