Analytica Data Science SolutionsContact
ESC

to move to open

← All work

Case study

Predicting how a hurricane cascades through US critical infrastructure

A hurricane doesn't damage one thing. It damages a port, which idles a refinery, which starves a region — and the asset records that would let anyone see that coming lived in multiple federal and commercial data sources that shared no identifier.

Client
Cybersecurity and Infrastructure Security Agency (CISA), US Department of Homeland Security
Sector
Federal government · Homeland security · Critical infrastructure
Role
Subcontractor
Period
2018 – present
Result
2018 – continuing, as a subcontractor
the entity-matching engine built for it
fastEMthe entity-matching engine built for it
infrastructure sectors linked
5+infrastructure sectors linked

The problem

CISA is responsible for understanding what a disaster will do to the country's critical infrastructure before it does it. The hard part wasn't the modeling: the asset records needed to answer the question were spread across multiple federal and commercial data sources — each built separately, for different purposes, with no identifier in common.

Until those records are reconciled, no forecast is possible. A model can't predict what a storm does to a facility when the same facility appears four times under four names.

What we built

The first half of the work is ingesting, cleaning and normalizing heterogeneous data — the IP Gateway, All Hazards Analysis, NOAA, US Census, Harvard Business School economic data, and sector-specific asset holdings — and then resolving it into one coherent picture.

fastEM
An entity-matching engine built for this, combining probabilistic linkage (fastLink), spatial reasoning (Haversine distance with K-D tree indexing) and deterministic rules to link and de-duplicate infrastructure records that share no key. Geography does work here that string comparison cannot: two records naming the same site differently are still at the same coordinates.
Critical Asset List and Significant Asset Set
Rule sets applied over the resolved records, and continuously refined, across banking, communications, energy, transportation and healthcare — delivered in standard formats (CSV, RDS) the client's own analysts work with directly.
The forecast layer
Models that project the impact of a natural disaster onto the assets identified, and onto the economic clusters that depend on them.
The Infrastructure of Concerns Tool
A secure cloud-hosted GIS platform: interactive maps, filtering by sector and region, reporting, and predictive-impact visualization — plus the architecture documentation, function-level docs for fastEM and workflow guides needed to maintain it.
Named data sources with no common identifier, resolved into one asset pictureThe source families named in the engagement record and shown in this figure — the IP Gateway, All Hazards Analysis, NOAA, the US Census, Harvard Business School economic data and per-sector asset holdings — none of which share an identifier. They feed an entity-matching engine that combines probabilistic linkage, spatial reasoning and deterministic rules. The output is one reconciled asset picture spanning banking, communications, energy, transportation and healthcare, which is what every downstream forecast is computed against.SOURCE FAMILIESNo identifier in common.ONE MATCHING ENGINEfastEMONE ASSET PICTUREEverything downstream reads this.IP Gatewayfederal asset inventoryAll Hazards Analysishazard modeling dataNOAAweather and storm tracksUS Censuspopulation and economyHarvard Business Schooleconomic cluster dataSector holdingsper-sector asset recordsProbabilisticfastLinksame record, or notSpatialHaversine distanceK-D tree indexDeterministicrules that must hold, not mayBankingCommunicationsEnergyTransportationHealthcareEvery downstream forecast is computed against this reconciled picture.
If two records for one substation stay separate, the forecast counts it twice and overstates resilience; if two substations get merged, it understates the exposure. Neither error announces itself in the output, which is why the effort went into the matching engine rather than the forecasting model.

Why the record linkage comes first

Every model downstream inherits the quality of that reconciliation. If two records for one substation stay separate, the forecast counts it twice and overstates resilience. If two different substations get merged, it understates the exposure. Neither error announces itself in the output, so most of the engineering effort went into the matching engine.

The same capability — record linkage across systems that were never designed to be joined — recurs in our data-integration work whenever source systems do not align.

Result

A cloud platform that lets decision makers see, before a storm makes landfall, which critical assets and which economic clusters are in its path. The client reports improved readiness and preparedness, better resource allocation through prioritization and crew positioning, and reduced preparedness costs.

Discuss a similar problem

If one of these engagements resembles a problem you're facing, we can walk through how it was built and what it would take in your environment.

Start a conversation