Asset Location Data

The physical-footprint layer under every risk score.

Asset Location Data is the methodology behind alphaX's physical-footprint intelligence — the coordinate-level layer that every nature, climate and geopolitical risk score at alphaX is sampled against. It can't be bought off the shelf as a location file.

Aerial view of forest canopy representing alphaX's natural-world data
12,400
Companies tracked
99%+
Of global equity
3.2M
ISIN-attributed locations
8M+
Facilities geolocated
Why it matters

Risk doesn't hit a company. It hits a place.

Every position an institution holds has a physical footprint — mines, plantations, refineries, terminals, warehouses. A drought does not hit a company, it hits a river basin, and with it the eleven facilities inside that basin rather than the ninety outside it. Nearly every risk analytic on the market still scores at company or sector level, because pinpointing where a company actually operates has been too expensive to do at scale.

12,400
Companies tracked
MSCI ACWI IMI plus client-requested extensions
99%+
Of global equity
By market capitalisation covered
3.2M
ISIN-attributed locations
Meta-data rich, high-fidelity records
8M+
Facilities geolocated
Across the full discovery pipeline
Map showing a river-basin catchment boundary with a facility location pinned inside it
Example: one river-basin catchment (outlined) and a facility inside it. The same drought, flood or water-allocation decision reaches every asset inside that boundary — regardless of company or sector.
Reference data

Every basin has more than one tenant.

An asset location on its own says nothing about how contested the ground underneath it is. The same discovery pipeline also geolocates everything else sharing that basin, watershed or grid connection — competitors, farms, water utilities, ports — so a scarcity signal can be attributed to the right assets instead of averaged across a sector.

8M+
Reference locations mapped
70+
Reference asset categories
World map heatmap showing density of reference asset locations
Reference data

70+ categories, from mining to medical centres

Reference assets are typed and geolocated with the same discovery pipeline used for target companies — spanning agriculture, energy, logistics, government and more. Grounded against over 8,000,000 reference locations, so a basin-level view can account for who else is drawing on the same resources.

Grid of asset category photos including ports, warehouses, vineyards, mining and utilities
The asset base

Not just a pin on a map

Asset Location Data is the layer everything else at alphaX sits on. Every risk layer is only as precise as the coordinate it is sampled at, and every financial attribution is only as sound as the entity the asset is tied to. Beyond name, address, latitude and longitude, each record carries the fields that let an analyst judge how much to trust it, and how much to care.

Precision, published not implied

coord_precision is one of rooftop, street, neighbourhood, city or country — so an analyst can exclude anything coarser than street level before running a basin-level intersection.

Live status, not a snapshot

Operational, under construction, planned, closed, mothballed, temporarily closed, closing, or under exploration — read alongside the date that status was determined.

Materiality, not just a count

Capacity and units capture operational throughput where it is disclosed. Without it, a site count cannot be weighted for how much it actually matters.

Full provenance on every row

Source, attribution and research date travel with the record, alongside the source's own asset description next to our standardised classification — so any mapping decision can be re-examined.

Discovery

One staged pipeline, not one tool

There is no single method that finds corporate facilities across every sector and country, so we didn't build one. Each stage below is an independent component under a central orchestrator, and every stage reads and writes the same strictly-typed asset schema — so sources as different as a mining registry and a company's own site map converge into one structure.

01

Profiling

Build the company's corporate family, subsidiaries and ownership structure first — every later stage reads this.

02

Dataset gathering

Pull candidate assets from structured public registries covering specific sectors, then resolve each row to the company.

03

Web discovery

An agentic system with search tooling nominates candidate URLs, fetches them, and extracts structured asset rows targeting registry gaps.

04

Enrichment

Geocode addresses that arrive without coordinates, record the source and precision of every coordinate, then classify the asset type.

05

Validation

Clean, prioritise, deduplicate, and run a coverage check against the company's expected footprint.

06

Correction

A final verification pass over the coordinates themselves, closing the loop back to the schema.

Discovery

Fellegi–Sunter, not fuzzy matching

Registry rows are matched back to companies using Fellegi–Sunter probabilistic record linkage — the standard statistical framework for deciding whether two records describe the same entity. Non-matches are dropped rather than guessed at.

A statistical framework, not a heuristic

The same probabilistic linkage method used across official statistics agencies decides whether a registry row and a company are the same entity — with a defined confidence, not a string-similarity guess.

Confirmed matches do double duty

They become known assets, and they tell the web-discovery stage which sites are already covered — so it searches for what is missing instead of re-finding what it already has.

Non-matches are dropped rather than guessed at — a missing asset is safer than a wrong one.
Quality control

Precision tells you an asset is right. It says nothing about what's missing.

Asset location data always contains a degree of error — the question is whether the pipeline is built to find it. Ours applies three checks, different in kind, so they fail independently.

Geospatial plausibility

Ocean boundary layers flag land-based asset types geocoded into the sea, with genuine offshore types such as rigs and wind farms excluded. Urban boundary layers flag large-scale mining sites placed in the middle of a city.

Statistical deduplication

A duplicate is the same company, the same asset type, within a defined proximity. Candidates are confirmed individually, then collapsed to the more authoritative record. Every merge keeps an audit trail.

Coverage QA

Rather than only asking whether each asset found is correct, an automated check asks whether the company's expected footprint has been captured at all. A high-severity gap re-triggers discovery, up to three times.

Judging assets one at a time tells you about precision and nothing about recall — the coverage loop is what's built to catch it.
Classification

50 asset types, resolved with the company's own numbers

Raw descriptions arrive as things like "Fab 4", "Assembly Plant" or "Coal Unit 1". We map them to a custom taxonomy of 50 asset type codes in 14 top-level groups, grounded in NAICS 2022 but adapted for physical and nature materiality.

1

Deterministic lookup

A persistent mapping table resolves any description seen before — no model call needed.

2

Model assignment

Unseen descriptions go to a language model, which assigns a best-fit code and a confidence flag. High-confidence results join the mapping table.

3

Revenue disambiguation

Ambiguous cases — "Manufacturing Plant" could be food, chemical or machinery — are resolved per company using that company's own NAICS revenue breakdown. Only the remainder reach a model again, with full company context.

The dependency argument in miniature — a location file alone cannot reproduce this without the same revenue-split data underneath.
Refresh

Monthly publication, tiered by how fast the world changes

A new asset database publishes every month. Underneath that, investigation cycles are tiered by how fast the physical world actually changes.

Stable sectors
12–18 months

Offices and plants that stay put are re-investigated on a slower cycle — the physical reality barely moves year to year.

Fast-changing sectors
Materially faster

Data centres are the obvious case: build-out is fast enough that an annual cycle would publish a map that's already wrong.

Prioritised regions
Investment-led

Countries or regions seeing heavy infrastructure investment or regulatory change get prioritised the same way — the cadence follows the risk.

The difference

Why this doesn't look like the rest of the market

The legacy approach
alphaX
Large offshore teams reading filings by hand
A multi-source AI pipeline — registries, web crawling, geospatial and satellite data, run on commercial models and locally hosted GPUs
Partial coverage, refreshed annually
Monthly publication, with a continuous GPU-based verification pipeline running around the clock
A snapshot that decays until the next manual pass
Corrections feed back into the next run, so accuracy compounds instead of decaying
Asset types assigned in isolation from the company
Ambiguous types resolved against the company's own revenue-split data — a link a location file alone can't reproduce
The bottom line

The hinge the rest of the platform stands on

Nature, climate and geopolitical risk are all sampled at the coordinates Asset Location Data provides. Building it required the same financial-data foundation that most competitors sell as separate products, if they sell them at all.