How to Use epoch Fields Correctly When Stitching History and Forecast Data Together

If you’re building anything that combines historical observations with a forward forecast — a 7-day view with actuals up to now and model output after — you will eventually get burned by how epoch timestamps work across WeatherAPI’s different endpoints. Not because the API is wrong. Because the assumptions baked into how developers join the two datasets are usually wrong.

What You’re Actually Getting Back

The history.json endpoint and the forecast.json endpoint both return hourly data with a time_epoch field per hour. But here’s the detail that bites people: time_epoch represents the local clock time of that hour in the location’s timezone, converted to Unix epoch. It is not a pure UTC epoch.

That sounds fine until you’re in a location that observes DST. The hour that doesn’t exist during the spring-forward transition, or the hour that appears twice during fall-back, causes real misalignment when you’re sorting or deduplicating a stitched array. Sort all records by time_epoch and deduplicate by that value alone, and the fall-back duplicate hour produces two rows with identical epoch values pointing at genuinely different observation windows. You can’t tell them apart from time_epoch alone.

The Spring-Forward Trap

The missing hour is arguably worse for stitching. Take a US Central Time location on the March DST transition day. Local clocks jump from 01:59 to 03:00. Your history response will have no 02:00 row — that hour didn’t exist locally. Your forecast response, depending on when you made the request relative to the model run, may or may not handle this correctly either.

If your frontend expects 24 hourly records per calendar day and renders them by index position, you get a gap or a silent off-by-one that shifts every subsequent hour label. Energy consumption dashboards and HVAC scheduling tools run into this. The fix isn’t complicated, but you have to account for it explicitly: validate that the count of hourly records per local date is actually 24 before you render by index, and fall back to rendering by time string instead.

The Right Way to Stitch

The join key should not be time_epoch alone. Use the time string field — formatted as YYYY-MM-DD HH:mm in local time — as your primary key for deduplication and ordering within a single location. This sidesteps the DST epoch collision entirely because the string representation is unambiguous for a given location context.

Concretely, something like this in Python:

from datetime import datetime

def stitch_hours(history_hours, forecast_hours):
    seen = {}
    for h in history_hours:
        seen[h['time']] = h  # 'time' string as key
    for h in forecast_hours:
        if h['time'] not in seen:
            seen[h['time']] = h
    return sorted(seen.values(), key=lambda x: x['time_epoch'])

History takes precedence over forecast for any overlapping window — that’s the right default because the history endpoint reflects what actually happened (or the best observational estimate of it), while the forecast for the same period is model output that’s already verifiable against reality.

Sorting the final result by time_epoch for display ordering is fine once deduplication is done. At that point you’re ordering within a single timezone context, not joining across them.

The Overlap Window Problem

WeatherAPI’s forecast.json endpoint with days=3 typically starts at today’s date at 00:00 local time. If you’re making this request at 14:00, hours 00 through 13 of today exist in both the forecast response and in history — and they’re not necessarily identical values. The forecast for those hours was issued at some earlier model run and represents modeled output; history for those hours reflects observed or nowcast data.

History is more accurate for anything more than a couple of hours back. The subtlety: for the current hour and possibly the previous one, the history endpoint may still be returning a nowcast estimate rather than a finalized observation, depending on how quickly METAR and synoptic station data flows through. There’s a lag between an observation being taken at a station and it propagating through ingestion into the historical record — we see this in our own pipeline. For most use cases it doesn’t matter. For anything where you’re trying to do precise validation against a sensor, it does.

Cross-Timezone Stitching Is a Different Problem

If you’re stitching records across multiple locations — a comparison dashboard for five cities in different timezones, say — don’t use the time string as a join key across locations. That joins London’s 14:00 to New York’s 14:00, which are completely different moments in absolute time.

For multi-location work, flip back to time_epoch as your global join key and accept the per-location DST limitation (usually acceptable when you’re comparing weather, not building a legal audit trail). Or convert everything to UTC explicitly using the tz_id and localtime_epoch fields in the location block to derive the UTC offset, then normalize before joining. That’s the more correct approach. It’s also more code.

One Edge Case Worth Knowing

Locations that use half-hour or 45-minute UTC offsets — India (UTC+5:30), Nepal (UTC+5:45), parts of Australia (UTC+9:30) — produce epoch values that don’t fall on hour boundaries when converted to UTC. If you’re doing any server-side bucketing by UTC hour (rounding time_epoch to the nearest 3600), you’ll misplace these records. Bucket by the time string’s hour component in local time instead.

This matters more than it sounds if your app has any agricultural or logistics angle in South Asia. A request for Kolkata hourly data bucketed by UTC hour will have its records shifted 30 minutes relative to where they should sit in a local-time chart.

Practical Checklist Before You Ship

  • Deduplicate history/forecast overlaps using the time string field, not time_epoch.
  • Validate hourly record count per local date — 23 or 25 records on DST transition days is correct, not a bug.
  • Don’t sort by time string lexicographically across timezone boundaries.
  • If you’re normalizing to UTC for cross-location joins, use tz_id from the location block and a proper timezone library (pytz, Luxon, NodaTime) — don’t derive UTC offset from time_epoch arithmetic alone.
  • Prefer history over forecast data for any hour that’s already passed, even if the overlap window means you have both.

Before you finalize your data layer, decide explicitly what your app needs to do with times that don’t exist or appear twice. Most apps can ignore DST edge cases entirely. Apps touching scheduling, billing intervals, or sensor correlation usually can’t — and you’d rather make that call now than trace it back from a support ticket filed in November.

Scroll to Top