Brier skill vs. the "it always blows" baseline — the honest, threshold-free measure, since ~78% of summer days are windy anyway.
| Version | Accuracy* | P(good) | R(good) | Brier ↓ | kt MAE |
|---|---|---|---|---|---|
| v1 — baseline features | 81% | 85% | 90% | 0.153 | — |
| v2 — label fix, gradient stations, buoy | 84% | 89% | 91% | 0.121 | — |
| v3 — cloud cover, kt regression | 85% | 90% | 91% | 0.117 | 1.8 kt |
| v4 — strength is the headline (same model) | 85% | 90% | 91% | 0.117 | 1.8 kt |
| v5 — sustained-wind label + build-up clouds | 82% | 85% | 86% | — | 1.8 kt |
| v6 — coastal-eddy features + advisory | 82% | 85% | 85% | — | 1.8 kt |
*at operating threshold 0.4. v1 measured against its own (buggy) labels and 75% base rate; v2/v3 against the corrected labels (78% base). Brier skill = 1 − Brier/Brieralways-yes. v5 is scored against the stricter sustained-wind labels (a good day must hold ≥12 kt for ~1h, not just touch it), so its numbers sit on a harder target than v3/v4 — the point of v5 is fewer trust-busting misses, not a higher score. v6 adds the eddy signal: aggregate accuracy barely moves (eddy days are rare), but it correctly downgrades the eddy-bust class and flags the rest as low-confidence.
Ten years of METAR, labeled from Moffett Field (peak ≥ 12 kt NW, 14–18h). Logistic regression on season, valley–coast ΔT, 850 hPa sounding, marine layer, persistence, SFO onset.
What it taught us: the ΔT sweet spot (~19 °F — hotter is not better, weight +0.62), hot air aloft kills the breeze (t850 −0.35), and the morning surface pressure gradient SFO−SJC is useless before the afternoon.
78% vs 75% base — barely better than "always yes". April was actively broken (59% vs 69% base).
Dissecting 119 misclassified test days found three things:
1. A label bug — 89 windy days (5%) blew from 351–356°, just past the 350° sector edge, and were scored "no wind". The model had been right all along. Sector widened to 280–360°.
2. April is a different animal — its missed days are post-frontal NW gradient wind (cold aloft, no ΔT, 10–26 kt at 850 hPa), not thermals.
3. New stations where the errors pointed: Sacramento, Arcata, Winnemucca (large-scale pressure gradients) and NDBC buoy 46012 (onshore flow at the source). The SFO−Winnemucca gradient — the offshore-event detector — instantly became the #2 feature (+0.57).
84% vs 78% base, Brier 0.153 → 0.121, April fixed (72% vs 70%), September 55% → 76%.
The old sun features only looked at the morning. The thermal lives off afternoon valley heating — so: forecast afternoon cloud cover (low/mid/high), afternoon insolation, and stratus clear-out hour.
Mid-level afternoon cloud went straight to #4 overall (−0.48) — monsoon moisture and cirrus shields are real thermal killers.
Plus a regression head on the same features: expected peak knots, MAE 1.8 kt, 82% within ±3 kt — because for winging, 12 kt means the big wing and 18 kt means the small one.
The binary GO/NO is nearly saturated — in high summer 90%+ of days blow, so "it'll be windy" is hard to beat. The real question for a winger is how windy: 12 kt (big wing) vs 18 kt (small wing) vs a holey 14. So the expected-peak-knots regression moved from a footnote to the headline; GO/NO shrank to a chip.
Does the knots forecast actually rank strength? On the held-out test set it does — cleanly and monotonically:
| model says | days | actual mean | actually ≥15 kt |
|---|---|---|---|
| <11 kt | 80 | 11.0 kt | 9% |
| 11–13 kt | 119 | 12.6 kt | 19% |
| 13–15 kt | 187 | 14.9 kt | 56% |
| 15–18 kt | 144 | 16.4 kt | 83% |
Correlation predicted-vs-actual r = 0.67, MAE 1.8 kt over a 6–21 kt spread. When it says "15–18", 83% of those days truly hit ≥ 15 kt; when it says "under 11", only 9% do. Shown now as a 1–5 strength scale (too light · marginal · good · strong · send it) with a likely range.
Tested but shelved: the coastal eddy. A southerly morning/midday coast (eddy) was floated as a bust cause. Against 10 years it turned out to nudge good days toward marginal (holey 10–16 kt), not toward busts — a quality effect, not an on/off switch. Filed for the knots model, not the binary. A buoy southerly-component feature (we currently only keep the westerly component, blind to S-vs-N) is the cheap follow-up.
A Monday called GO (0.77) turned out cloudy and light — and the old track record still scored it a hit, because a single 13 kt gust cleared the bar. Two failures, one investigation:
1 · A gust is not a session. "Good" was max ≥ 12 kt —
one spike sufficed. Now a good day must hold usable wind:
≥ 30% of the 14–18 window at/over 12 kt (~1 hour), fraction-based so
it survives the uneven METAR cadence. Both the training labels and the live
track record use it; re-scoring flipped four borderline spike-days
good→no, including that Monday.
2 · Clouds killed it during the build-up, not the window. Mid-level
cloud sat at 99–100% over the valley from 10–13h — starving the heating
that fires the thermal — then cleared by 16:00. The old cloud/radiation
features averaged 12–16h, so they both missed the killer hours
and flattered the day as it cleared. Retimed to the
10–13h build-up window (cc_mid_build,
cc_high_build, swrad_build):
| old (12–16h) | new (10–13h) | |
|---|---|---|
| Monday mid cloud | 46% | 90% |
| model P(good) | 0.77 → GO ❌ | 0.26 → NO ✓ |
Aggregate test accuracy is unchanged (~82%) — most days are
cloudy or clear across both windows alike. The gain is precisely on the
rare "cloudy-morning-that-clears-late" busts, the ones that cost trust.
The retimed cc_mid_build inherited the strong −0.29 weight;
simply adding a build-up feature next to the old one did nothing
(collinear — the 12–16 mean had already claimed the signal).
A hot-valley day was called GO and blew a holey ~11 kt. The tell was out at sea: buoy 46012 showed a southerly ocean wind (150°) and warm water (16°C, +2.6° above normal). When the coast wind swings south, coastal upwelling shuts off, the ocean surface warms, and the marine air that feeds the sea breeze arrives warm instead of cold — so the breeze goes soft no matter how hot the valley is.
The model was blind to it: buoy_w_am is only the westerly
component, so a south wind reads like calm, and sea-surface temp was unused.
Added buoy_s_pos (the rectified southerly excursion) and
dT_e16_sst (valley→ocean driving contrast). Validated over
10 years and out-of-sample:
| buoy in the morning | days (test) | model P(go) | actually good |
|---|---|---|---|
| calm / northerly (upwelling) | 44 | 0.48 | 68% |
| southerly (eddy) | 8 | 0.36 | 12% |
The honest limit. On a normal day the eddy correctly pushes the call toward NO. But on an extreme hot-valley + eddy day the model still over-commits — it adds "hot valley" and "eddy" instead of letting the eddy cap the valley's benefit (a real interaction: hot-valley days drop from 60% to 40% good under an eddy). Rather than overfit the weight to a couple of bust days, eddy days now raise a low-confidence advisory banner on the page — the honest move: show the call and the reason to distrust it.
The pipeline was generalised from Shoreline-only to a spot registry: everything that differs between spots (ground-truth station, wind sector, session window, coordinates) is config, so the label→build→train→predict→ verify machinery stays a single shared codebase. A new spot is a config entry, not a code fork — the same fix now lands on all of them.
Candlestick Point was added as spot #2: truth station SFO, WNW gap-flow sector 250–320°, and a later 14–20h window (the SF sea breeze holds into the evening, ~12 kt at 20:00, unlike the south-bay thermal that dies by 18:00). Its physics differs sensibly — the raw inland→coast temperature gradient dominates (gap flow scales with the pressure gradient), and it weights the eddy harder than Shoreline (−0.48 vs −0.21), being a purer coastal spot.
Redwood City followed as spot #3 (truth station San Carlos / KSQL, mid-bay WNW sector 245–335°, same 14–18h south-bay thermal window as Shoreline). Its shorter usable record — the AWOS only reports densely from 2019 — made the 37-feature model overfit at the default regularization (expected-kt skill went negative), so it runs a stronger per-spot L2 ridge; a plausibility filter also drops the AWOS's occasional 47–165 kt sensor spikes before they poison the labels. Switch spots at the top of the forecast page.
| Month | Base | v1 | v2 | v3 |
|---|---|---|---|---|
| April | 70% | 59% ⚠ | 72% | 72% |
| May | 88% | — | 91% | 92% |
| Jun–Aug | 91–96% | — | ~90% | ~90% |
| September | 55% | 71% | 76% | 76% |
| October | 56% | 60% | 63% | 65% |
High summer is nearly unbeatable by construction — the value is in the shoulder months, where generic forecasts flounder.
· Morning surface gradient SFO−SJC — the thermal gradient doesn't exist
before noon.
· The buoy (yet) — 2025 data hole (buoy adrift) and overlap with the
pressure gradients.
· "Delta competition" (Sacramento stealing the flow) — cute hypothesis,
+0.04 weight. Reality declined.
· nw850 as an explicit feature — the April fix came from the label repair
and seasonality, not from the flag.