GRIB2 files are not friendly to read sequentially. HRRR model output for a single cycle can run into the gigabytes when you include the full variable set across all pressure levels, and if your ingestion pipeline is naive about how it reads those files, you’ll spend most of your time doing I/O on data you were going to throw away anyway.
We write our GRIB2 ingestion pipeline in C#, which means we’re not using the Python cfgrib/xarray stack that most of the meteorological tooling community defaults to. Fewer pre-built helpers, but it also means we’ve had to understand what the format is actually doing at a lower level rather than letting a library absorb the complexity. Some of those lessons apply regardless of language.
What makes GRIB2 retrieval expensive
A GRIB2 file is a flat sequence of independent messages. Each message encodes a single variable, at a single level, for a single forecast hour. HRRR f01 through f18 (3km grid, CONUS domain) might contain hundreds of messages per cycle — surface temperature, wind components at multiple pressure levels, CAPE, composite reflectivity, and so on. When you want temperature and dewpoint at 2m for 40,000 locations, you need exactly two of those messages. Reading the whole file to find them is wasteful.
NOAA publishes .idx index files alongside each GRIB2 file on their HTTPS servers. These are plain-text offsets: each line maps a message number to a byte offset in the parent GRIB2 file, and includes the variable name and level. That means you can do HTTP range requests — fetch only the byte ranges you actually need, rather than downloading the entire file. For HRRR surface output that can exceed 400MB per cycle, the difference between fetching two messages versus the whole file is the difference between sub-second retrieval and tens of seconds of download time.
The pattern: parse the .idx to build an offset map, identify the messages you need, request only those byte ranges. Most cloud storage and NOAA’s own HTTPS endpoints support range requests natively. This is how we avoid pulling down gigabytes of GRIB2 data we’ll never decode.
The merge-threshold problem
Even with range requests, the size of each request matters. GRIB2 messages aren’t uniform — a full 3km HRRR domain grid for a single variable at a single level is roughly 1059×1799 grid points. At 32 bits per value before compression, that’s over 7 million floats. After JPEG2000 or grid-point packing, message sizes vary considerably depending on the field — reflectivity compresses differently than temperature — but you’re typically looking at hundreds of kilobytes per message.
If you’re fetching multiple messages and they happen to be contiguous or near-contiguous in the file, it usually pays to merge adjacent byte ranges into a single request rather than making N separate calls with small gaps between them. We use a simple heuristic: if two needed message byte ranges are within some threshold of each other, merge them into one request and discard the bytes in the gap locally. This reduces round-trip overhead at the cost of pulling a small amount of unwanted data. The right threshold depends on your network latency — over a high-latency connection to NOAA’s servers, merging aggressively makes sense; over a local cache or nearby mirror, it matters less.
Getting that threshold wrong in the expensive direction means you issue a lot of small HTTP requests, each with its own connection overhead. Getting it wrong in the cheap direction means you’re occasionally pulling a few hundred kilobytes of irrelevant messages and discarding them. The second failure mode is almost always less bad than the first.
Field selection: upstream of all the performance work
None of the above matters if you’re not deliberate about which fields you actually ingest. HRRR has a lot of variables. Most weather API use cases need a small subset: 2m temperature, 2m dewpoint, 10m U and V wind components, surface pressure, APCP (accumulated precipitation), cloud cover, and maybe CAPE and lifted index for severe weather signals. That’s maybe 10–15 messages out of hundreds in a typical output file.
We process HRRR, NAM, and GFS — each with different grid resolutions and variable naming conventions. HRRR’s 3km grid is genuinely useful for surface conditions in complex terrain where NAM’s 12km grid smooths over real topographic variation. But HRRR covers CONUS only and runs hourly out to 18 hours (48 hours for the 00z and 12z cycles), while GFS runs out to 384 hours at 0.25-degree resolution globally. The practical decision we landed on: HRRR for short-range CONUS where resolution matters, GFS for global and extended range, NAM as a middle layer. Stitching these into a coherent hourly grid for a given location requires knowing which model takes precedence and at which forecast hour the handoff happens — and that logic lives in the ingestion pipeline, not downstream.
Grid interpolation isn’t free
Once you’ve retrieved and decoded the GRIB2 messages you need, you still have to map grid-point values to actual lat/lon coordinates. HRRR uses a Lambert conformal conic projection. GFS uses a regular latitude/longitude grid. NAM uses a different Lambert conformal specification. None of them align natively.
For a given query location, you need to find the surrounding grid points in the native projection, then interpolate. Bilinear interpolation across the four nearest points is standard and usually sufficient. But doing this per-request, per-location, for tens of thousands of locations in real time is too slow — so we pre-compute and cache the grid index mappings for every location we serve. The mapping from (lat, lon) to (grid_i, grid_j) in each model’s native projection is fixed until the grid changes, so computing it once and storing it is an obvious win.
What’s less obvious: you also have to handle grid point selection carefully near coastlines. A bilinear interpolation that mixes an ocean point with adjacent land points will produce nonsense for surface air temperature. We mask ocean-flagged grid points and fall back to the nearest valid land point rather than letting the interpolation quietly average across a land/sea boundary.
A concrete failure mode worth knowing about
If you’re building your own GRIB2 ingestion rather than using a managed layer, one failure mode we hit early was not validating that the GRIB2 file had actually completed uploading before fetching it. NOAA uploads files progressively, and the .idx file can appear before the GRIB2 is fully written. If you parse the index and immediately start range requests, you can get truncated messages — which may decode silently to plausible-looking garbage rather than obvious errors. The fix is straightforward: check that the file size matches what you actually received, or add a retry window after the nominal cycle availability time before treating a partial file as valid.
It’s an easy bug to miss because the values aren’t obviously wrong — they’re in the right ballpark, just subtly off. The kind of thing that surfaces in user reports as “temperature looks a bit high this morning” rather than a clear error.
What this means if you’re consuming an API rather than building one
If you’re using WeatherAPI rather than running your own ingestion, most of this is invisible to you — which is the point. But it does explain why forecast values for a mountainous CONUS location tend to be more accurate than what you’d get from a naive nearest-grid-point lookup against raw GFS output: model selection, elevation correction, and grid interpolation are already applied.
Where it becomes relevant to API consumers is data latency. HRRR hourly cycles nominally complete around 45–55 minutes after the top of the hour. GFS takes considerably longer — the 00z cycle isn’t fully available until several hours after midnight UTC. If your application depends on the freshest possible HRRR data, there’s a window each hour where the previous cycle is still what’s being served. That’s not a bug; it’s model processing time. Knowing this helps you reason about forecast age rather than assuming a 3am API call is using a 3am model run.
If you are building your own pipeline: get field selection right first, then optimize retrieval with range requests and index parsing, then tackle interpolation and edge cases. Doing it in the other order means optimizing something you’ll want to throw away once the field list changes — and it will change.
