Why weather_alerts Deduplication Is Harder Than It Looks: Event IDs, Polygon Overlap, and What to Actually Store

Weather alerts look straightforward from the API side: you get an array of active alerts for a location, each with a headline, severity, and some timing fields. Easy enough to display. The problem shows up the moment you try to build a reliable notification system on top of them — one that doesn’t spam users with duplicates, doesn’t miss a new alert because it looks similar to an old one, and doesn’t silently drop an upgrade from Watch to Warning.

The fundamental issue is that alert identity is messier than it looks. NWS uses event IDs called UGC codes and VTEC strings — a VTEC string like /O.NEW.KBMX.TO.W.0042.260812T2130Z-260813T0000Z/ encodes the issuing office, event type, action code, and time window in a single string. WeatherAPI’s alerts.alert[] response doesn’t expose raw VTEC strings. You get a headline, msgtype, effective, expires, and an areas string. That means you’re reconstructing identity from fields that aren’t guaranteed to be stable across the full lifecycle of a single NWS event.

Why headline + expires Isn’t a Reliable Key

The obvious approach is to hash headline + effective + expires and use that as a dedup key. It works until an alert gets extended. When NWS extends an expiry time, the effective timestamp stays the same but expires changes — so your hash changes, and you fire a second notification for what the user experiences as the same event. Worse, when an alert upgrades from Watch to Warning, both the headline and the expires typically change, which is exactly the case where you do want to notify.

The hash-based approach fails in two opposite directions: it fires a duplicate on extensions, and it can miss an upgrade if you deduplicate too aggressively.

A More Stable Identity Model

What works better is a composite key that’s sensitive to the right things. Based on what fields are reliably available in the response:

  • Event type — Tornado Warning, Winter Storm Watch, etc. This is stable within an alert’s lifecycle.
  • Issuing office — the areas field combined with the issuing office in the headline gives you enough to identify the source NWS WFO (Weather Forecast Office).
  • Effective time — the original issuance timestamp. This doesn’t change on extension; NWS issues a new VTEC action code and a new effective time for genuine upgrades.

Using event_type + office_slug + effective_epoch as a composite key is more stable than including expires. Extensions bump expires but leave effective alone. Genuine upgrades (CON to UPG in VTEC language) come with a new effective time, so they get a new key and correctly trigger a new notification.

The limitation: you’re deriving a proxy for VTEC identity from fields that don’t expose VTEC directly. It’s not perfect. If two different NWS offices issue simultaneous Tornado Warnings with the same event type and the same effective minute — rare but possible — you need location to disambiguate. Including a normalized version of the areas field handles that edge case.

The Polygon Overlap Problem

Alert polygons are another source of spurious duplicates that’s easy to miss. NWS alert polygons are county-based or zone-based rather than drawn to a specific point. If a user is near a county boundary, they may simultaneously be inside two overlapping alert zones covering the same underlying event. Your API request for their coordinates returns both alerts in the array.

You need to deduplicate within the response before you even get to the cross-request dedup logic. Two alerts with the same event type, the same effective time, and a time-window overlap above roughly 80% are almost certainly the same physical event reported across two administrative zones. Collapsing those before persisting anything to your store avoids a whole class of false duplicates.

One practical approach: normalize the event type string, strip trailing zone and county suffixes from the headline, and compare effective and expires windows. If they overlap past your threshold, keep the alert with the broader coverage area — a longer areas string is a rough proxy — and discard the other.

Handling Cancellations Correctly

An alert disappearing from the array doesn’t necessarily mean it expired. It could mean NWS cancelled it early (VTEC action CAN) or that it was replaced by a more severe successor (UPG). In both cases you’d want to notify the user that the alert is no longer active, but the reason matters: a cancellation means the hazard is over, an upgrade means it got worse.

Without VTEC action codes in the response, the only signal you have is whether a previously-stored alert is absent from the current response before its expires time. If it’s gone early, flag it as possibly cancelled. You can’t distinguish CAN from UPG automatically from the API response alone — but you can check whether a new alert with the same event type appeared in the same response. If the old one vanished and a new one with a higher severity arrived in the same poll, that’s an upgrade pattern. Not airtight, but workable.

What to Actually Persist

Your alert store needs at minimum:

  • Your composite dedup key
  • The full alert payload (don’t just store the key — you’ll want the headline and severity later)
  • A first_seen_at timestamp from your own system clock
  • A last_seen_at timestamp, updated on every poll where the alert is still present
  • A notified boolean per user/location combination
  • An expires_at from the API response, for cleanup

The last_seen_at pattern earns its keep. An alert that stopped appearing in the response but hasn’t passed its expires_at time is worth a closer look — it may have been cancelled early or upgraded. An alert that’s still in the response and hasn’t been notified yet is straightforward. An alert whose last_seen_at is older than your poll interval but whose expires_at is still in the future is a signal to trigger a possible-cancellation notification.

Polling Cadence Matters More Than You’d Think

NWS issues and updates alerts on irregular schedules, but Tornado Warnings in particular can be issued with effective times only minutes out. If you’re polling every 15 minutes, you may never see a Tornado Warning that’s issued and expires within the same polling window. For severe convective events, a 3-5 minute poll interval on the alerts.json or the realtime endpoint for affected locations is the floor for having any hope of catching short-fuse warnings before they expire.

This isn’t a complaint about any particular API — it’s a fundamental mismatch between polling latency and NWS warning lead times for fast-moving events. If your application is safety-critical in this sense, polling alone is probably not the right architecture regardless of the data source.

A Caveat on Alert Coverage

WeatherAPI’s alert data is sourced primarily for the US, UK, and a subset of other regions — it’s not globally uniform. Outside the US, alert density and timeliness vary considerably depending on each national met service’s data-sharing practices. Don’t build a notification system that implies global coverage if the underlying data doesn’t support it for the regions you actually care about. Pull test data for a real location in your target region before committing to this approach in production.

If you’re starting from scratch, log every raw alert payload for a week across a handful of representative locations before you write a line of dedup logic. The patterns in when alerts appear, repeat, and drop out of the response will tell you more about what your composite key needs to handle than any upfront design exercise will.

Scroll to Top