📈 Model History

Which features moved the needle — Shoreline Lake thermal predictor. All numbers: held-out test 2024–2026 (537 days). ← back to forecast · stats for nerds

🚀 Skill over the versions

Brier skill vs. the "it always blows" baseline — the honest, threshold-free measure, since ~78% of summer days are windy anyway.

always-yes
0%
v1 baseline
39%
v2 label fix + gradients
45%
v3 + cloud cover
47%
VersionAccuracy*P(good)R(good)Brier ↓kt MAE
v1 — baseline features81%85%90%0.153
v2 — label fix, gradient stations, buoy84%89%91%0.121
v3 — cloud cover, kt regression85%90%91%0.1171.8 kt
v4 — strength is the headline (same model)85%90%91%0.1171.8 kt
v5 — sustained-wind label + build-up clouds82%85%86%1.8 kt
v6 — coastal-eddy features + advisory82%85%85%1.8 kt

*at operating threshold 0.4. v1 measured against its own (buggy) labels and 75% base rate; v2/v3 against the corrected labels (78% base). Brier skill = 1 − Brier/Brieralways-yes. v5 is scored against the stricter sustained-wind labels (a good day must hold ≥12 kt for ~1h, not just touch it), so its numbers sit on a harder target than v3/v4 — the point of v5 is fewer trust-busting misses, not a higher score. v6 adds the eddy signal: aggregate accuracy barely moves (eddy days are rare), but it correctly downgrades the eddy-bust class and flags the rest as low-confidence.

v1🌱 Baseline — the thermal engine

Ten years of METAR, labeled from Moffett Field (peak ≥ 12 kt NW, 14–18h). Logistic regression on season, valley–coast ΔT, 850 hPa sounding, marine layer, persistence, SFO onset.

What it taught us: the ΔT sweet spot (~19 °F — hotter is not better, weight +0.62), hot air aloft kills the breeze (t850 −0.35), and the morning surface pressure gradient SFO−SJC is useless before the afternoon.

78% vs 75% base — barely better than "always yes". April was actively broken (59% vs 69% base).

v2🐛 Error analysis pays the rent

Dissecting 119 misclassified test days found three things:

1. A label bug — 89 windy days (5%) blew from 351–356°, just past the 350° sector edge, and were scored "no wind". The model had been right all along. Sector widened to 280–360°.

2. April is a different animal — its missed days are post-frontal NW gradient wind (cold aloft, no ΔT, 10–26 kt at 850 hPa), not thermals.

3. New stations where the errors pointed: Sacramento, Arcata, Winnemucca (large-scale pressure gradients) and NDBC buoy 46012 (onshore flow at the source). The SFO−Winnemucca gradient — the offshore-event detector — instantly became the #2 feature (+0.57).

84% vs 78% base, Brier 0.153 → 0.121, April fixed (72% vs 70%), September 55% → 76%.

v3☁️ Cloud cover + expected knots

The old sun features only looked at the morning. The thermal lives off afternoon valley heating — so: forecast afternoon cloud cover (low/mid/high), afternoon insolation, and stratus clear-out hour.

Mid-level afternoon cloud went straight to #4 overall (−0.48) — monsoon moisture and cirrus shields are real thermal killers.

Plus a regression head on the same features: expected peak knots, MAE 1.8 kt, 82% within ±3 kt — because for winging, 12 kt means the big wing and 18 kt means the small one.

v4💨 Strength becomes the headline

The binary GO/NO is nearly saturated — in high summer 90%+ of days blow, so "it'll be windy" is hard to beat. The real question for a winger is how windy: 12 kt (big wing) vs 18 kt (small wing) vs a holey 14. So the expected-peak-knots regression moved from a footnote to the headline; GO/NO shrank to a chip.

Does the knots forecast actually rank strength? On the held-out test set it does — cleanly and monotonically:

model saysdaysactual meanactually ≥15 kt
<11 kt8011.0 kt9%
11–13 kt11912.6 kt19%
13–15 kt18714.9 kt56%
15–18 kt14416.4 kt83%

Correlation predicted-vs-actual r = 0.67, MAE 1.8 kt over a 6–21 kt spread. When it says "15–18", 83% of those days truly hit ≥ 15 kt; when it says "under 11", only 9% do. Shown now as a 1–5 strength scale (too light · marginal · good · strong · send it) with a likely range.

Tested but shelved: the coastal eddy. A southerly morning/midday coast (eddy) was floated as a bust cause. Against 10 years it turned out to nudge good days toward marginal (holey 10–16 kt), not toward busts — a quality effect, not an on/off switch. Filed for the knots model, not the binary. A buoy southerly-component feature (we currently only keep the westerly component, blind to S-vs-N) is the cheap follow-up.

v5🕙 The Monday bust: timing beats averaging

A Monday called GO (0.77) turned out cloudy and light — and the old track record still scored it a hit, because a single 13 kt gust cleared the bar. Two failures, one investigation:

1 · A gust is not a session. "Good" was max ≥ 12 kt — one spike sufficed. Now a good day must hold usable wind: ≥ 30% of the 14–18 window at/over 12 kt (~1 hour), fraction-based so it survives the uneven METAR cadence. Both the training labels and the live track record use it; re-scoring flipped four borderline spike-days good→no, including that Monday.

2 · Clouds killed it during the build-up, not the window. Mid-level cloud sat at 99–100% over the valley from 10–13h — starving the heating that fires the thermal — then cleared by 16:00. The old cloud/radiation features averaged 12–16h, so they both missed the killer hours and flattered the day as it cleared. Retimed to the 10–13h build-up window (cc_mid_build, cc_high_build, swrad_build):

old (12–16h)new (10–13h)
Monday mid cloud46%90%
model P(good)0.77 → GO ❌0.26 → NO ✓

Aggregate test accuracy is unchanged (~82%) — most days are cloudy or clear across both windows alike. The gain is precisely on the rare "cloudy-morning-that-clears-late" busts, the ones that cost trust. The retimed cc_mid_build inherited the strong −0.29 weight; simply adding a build-up feature next to the old one did nothing (collinear — the 12–16 mean had already claimed the signal).

v6🌀 The coastal eddy: when the ocean wind turns warm

A hot-valley day was called GO and blew a holey ~11 kt. The tell was out at sea: buoy 46012 showed a southerly ocean wind (150°) and warm water (16°C, +2.6° above normal). When the coast wind swings south, coastal upwelling shuts off, the ocean surface warms, and the marine air that feeds the sea breeze arrives warm instead of cold — so the breeze goes soft no matter how hot the valley is.

The model was blind to it: buoy_w_am is only the westerly component, so a south wind reads like calm, and sea-surface temp was unused. Added buoy_s_pos (the rectified southerly excursion) and dT_e16_sst (valley→ocean driving contrast). Validated over 10 years and out-of-sample:

buoy in the morningdays (test)model P(go)actually good
calm / northerly (upwelling)440.4868%
southerly (eddy)80.3612%

The honest limit. On a normal day the eddy correctly pushes the call toward NO. But on an extreme hot-valley + eddy day the model still over-commits — it adds "hot valley" and "eddy" instead of letting the eddy cap the valley's benefit (a real interaction: hot-valley days drop from 60% to 40% good under an eddy). Rather than overfit the weight to a couple of bust days, eddy days now raise a low-confidence advisory banner on the page — the honest move: show the call and the reason to distrust it.

🗺️ Three spots, one model

The pipeline was generalised from Shoreline-only to a spot registry: everything that differs between spots (ground-truth station, wind sector, session window, coordinates) is config, so the label→build→train→predict→ verify machinery stays a single shared codebase. A new spot is a config entry, not a code fork — the same fix now lands on all of them.

Candlestick Point was added as spot #2: truth station SFO, WNW gap-flow sector 250–320°, and a later 14–20h window (the SF sea breeze holds into the evening, ~12 kt at 20:00, unlike the south-bay thermal that dies by 18:00). Its physics differs sensibly — the raw inland→coast temperature gradient dominates (gap flow scales with the pressure gradient), and it weights the eddy harder than Shoreline (−0.48 vs −0.21), being a purer coastal spot.

Redwood City followed as spot #3 (truth station San Carlos / KSQL, mid-bay WNW sector 245–335°, same 14–18h south-bay thermal window as Shoreline). Its shorter usable record — the AWOS only reports densely from 2019 — made the 37-feature model overfit at the default regularization (expected-kt skill went negative), so it runs a stronger per-spot L2 ridge; a plausibility filter also drops the AWOS's occasional 47–165 kt sensor spikes before they poison the labels. Switch spots at the top of the forecast page.

⚖️ Current feature weights (top 12)

dT_opt2
+1.04
doy_cos
−0.80
buoy_w_am
−0.46
pgrad_sfo_wmc
+0.38
pgrad_sfo_sac
+0.36
t11_delta
+0.35
t850
−0.35
buoy_wspd_am
+0.34
haf_ceil
+0.31
cc_mid_build
−0.29
yday_good
+0.29
buoy_s_pos
−0.21
pushes toward GO pushes toward NO · standardized logistic weights · Shoreline model

🗓️ Where the skill actually lives

MonthBasev1v2v3
April70%59% ⚠72%72%
May88%91%92%
Jun–Aug91–96%~90%~90%
September55%71%76%76%
October56%60%63%65%

High summer is nearly unbeatable by construction — the value is in the shoulder months, where generic forecasts flounder.

⚰️ What did not work (kept for honesty)

· Morning surface gradient SFO−SJC — the thermal gradient doesn't exist before noon.
· The buoy (yet) — 2025 data hole (buoy adrift) and overlap with the pressure gradients.
· "Delta competition" (Sacramento stealing the flow) — cute hypothesis, +0.04 weight. Reality declined.
· nw850 as an explicit feature — the April fix came from the label repair and seasonality, not from the flag.