August 8, 2026

Humanitarian Data Exchange (HDX): Intelligence Source Guide

0

HDX is the UN’s open catalogue of humanitarian data: administrative boundaries, population figures, displacement counts, food prices, conflict events and needs assessments, contributed by hundreds of organisations. It is a catalogue, not a dataset, and the difference determines how you use it.

humanitarian-data-exchange-hdx-intelligence-source-guide

HDX is the UN's open catalogue of humanitarian data: administrative boundaries, population figures, displacement counts, food prices, conflict events and needs assessments, contributed by hundreds of organisations. It is a catalogue, not a dataset, and the difference determines how you use it.

At a glance

Source Humanitarian Data Exchange (HDX)
Category Conflict, Crime & Human Security › Conflict & Event Databases
Homepage https://data.humdata.org/
Machine interface https://data.humdata.org/api/3/action/package_search
Format JSON
Access Open — no account required
Disciplines Geospatial Intelligence, Open Source Intelligence
Mission domains Conflict & Humanitarian

Open humanitarian datasets catalog. — as catalogued in the platform’s own source registry.

The Humanitarian Data Exchange is operated by the United Nations Office for the Coordination of Humanitarian Affairs through its Centre for Humanitarian Data, and it is built on CKAN, the open-source data portal software. That architectural fact is the most useful thing to know about it, because it means the whole catalogue is queryable through the standard CKAN action API – dataset search with facets, dataset detail, organisation listing, resource metadata – and anything you have written against another CKAN portal works here. What sits in the catalogue is contributed by hundreds of organisations: UN agencies, international and national NGOs, government statistical offices, academic projects and private data providers. A dataset entry carries metadata – title, description, contributing organisation, location tags, time period, expected update frequency, licence, methodology notes – and one or more resources, which are the actual files, most often CSV or Excel, frequently shapefiles or GeoJSON for boundaries, occasionally an API reference. A significant subset of the tabular data carries HXL hashtags, a lightweight standard that annotates columns with machine-readable semantic tags so that files from different organisations can be combined without hand-mapping every column. The catalogue also maintains the Common Operational Datasets, the reference administrative boundaries and population baselines that humanitarian coordination in a country is supposed to be built on.

The job HDX does that nothing else does is make the humanitarian information ecosystem addressable from one place. Without it, finding the current administrative boundary file for a country in crisis, the agreed population baseline, the displacement tracking figures and the food price series means knowing which agency publishes each of those, in which country, in which year, and whether the version you found is the one everybody else is using. HDX collapses that into a search. For conflict and crisis analysis the value is specifically in the reference layers rather than the headline numbers: the boundaries and population baselines are the denominators under every rate you will ever compute, and using a different boundary set from the organisation you are comparing against produces figures that are incomparable in ways nobody notices. It is also the practical route into datasets that are individually well known but scattered – conflict event data, displacement tracking, food security classifications, market prices – packaged with metadata that tells you when they were updated and under what licence. For GEOINT and OSINT work on conflict, humanitarian access and population movement, this is the layer that turns geographic and demographic context from a research project into a lookup.

Who publishes it, and why that matters

OCHA runs HDX as a public good with no charging model, funded through the humanitarian system's own budgets, and the Centre for Humanitarian Data exists specifically to improve the use and management of data across that system. That gives it durability and a genuine mandate, and it also shapes what is in it. HDX is a platform rather than a publisher: with narrow exceptions it does not create the data, does not verify it, and does not take editorial responsibility for its accuracy. The contributing organisation owns the content and the quality. What OCHA does do is set the rules – a terms of service, a data responsibility framework that governs what may be uploaded, and an active review posture on datasets containing personal or demographically identifiable information, which are restricted or refused. The incentive structure is worth understanding: organisations contribute to be visible and to meet donor and coordination expectations, which means the catalogue is comprehensive for well-resourced international actors and thin for national and local ones. It also means datasets get uploaded at the moment an organisation wants credit for them and updated far less reliably afterwards, which is the single most important thing to know about using this catalogue.

Provenance is the first question to ask of any dataset and the one most often skipped. Who collects it, what their incentive is, whether they publish a methodology, and whether they correct the record when they get something wrong all bear directly on how much weight a finding drawn from it can carry.

What a record actually contains

The fields you will be working with, what each one means, and whether it is something you can pivot on. Read the meanings carefully — more analysis is wrecked by misreading a field than by failing to find one, and a field that looks like an observation is often an inference.

Field Type What it means Pivot value
name string The dataset's URL slug and stable identifier within the catalogue. This is what you key on, because titles change and organisational renaming is common in this sector. Dataset detail retrieval, resource enumeration, revision history.
title string Human-readable dataset name as the contributor wrote it. Naming conventions vary enormously between organisations, so title text is unreliable for automated classification and good only for display. None reliable; use tags, groups and organisation instead.
organization string The contributing organisation. This is the single most important quality signal in the record, because it tells you the methodology, the incentives and the likely update discipline behind the data. Other datasets from the same contributor, the organisation's own publications and methodology documentation.
groups array Location tags, normally countries, that place the dataset geographically. A dataset may carry several, and a regional dataset tagged with a dozen countries is not the same thing as a country dataset. Country dashboards, cross-referencing with other datasets for the same location, regional aggregation.
tags array Free and controlled vocabulary tags describing the subject. Coverage is uneven because tagging is the contributor's responsibility, so tag-based discovery finds well-curated datasets and misses poorly described ones entirely. Thematic discovery; always combine with full-text search because tags alone under-recall badly.
dataset_date timestamp The time period the data refers to, which is distinct from when it was uploaded and distinct from when it was last modified. Conflating these three is the most common error in using this catalogue. Temporal filtering, currency assessment, alignment with an event timeline.
last_modified timestamp When the dataset record or its resources last changed. A recent modification can mean new data or merely a corrected description, and the record does not always distinguish the two. Freshness monitoring, change detection between collection runs.
data_update_frequency enum The contributor's declared update cadence – daily, weekly, monthly, annually, never, as needed. It is a statement of intent, and the gap between it and actual behaviour is the catalogue's most useful quality signal. Staleness detection; compare declared frequency against observed modification history.
license_id string The licence for that dataset, set by the contributor. Licences vary across the catalogue from public domain dedications through attribution licences to bespoke restrictive terms, and there is no catalogue-wide licence. Reuse and redistribution assessment; must be checked per dataset and per resource, never assumed.
resources array The actual files, each with a format, a URL, a size and its own modification date. A dataset with a current metadata record can contain resources that have not been touched in years. Direct file retrieval, format-based filtering, per-resource freshness assessment.
format enum Resource file format – CSV, XLSX, shapefile, GeoJSON, PDF and others. The presence of a PDF where you expected tabular data usually means the underlying numbers are not actually machine-readable. Ingest routing; discrimination between genuinely open data and a document posted as data.
methodology string Free-text methodology note where the contributor provided one. Its presence and quality is a strong proxy for the seriousness of the dataset, and its absence should lower your confidence rather than be ignored. Quality assessment; the starting point for any question about how a figure was produced.
caveats string Contributor-declared limitations. Where present these are usually honest and specific, and they are routinely ignored by consumers who take the numbers and discard the metadata. Uncertainty documentation; should travel with any figure extracted from the dataset.
hxl_tags array Machine-readable column semantics inside HXL-tagged tabular resources, appearing as a hashtag row beneath the header. This is what allows automated combination of files from different organisations without manual column mapping. Automated schema alignment, cross-dataset joins on standardised concepts such as location, sector, population and date.

Coverage — and what is not in it

Coverage tracks the humanitarian system's attention, which is a specific and knowable bias rather than a random one. Countries with a declared emergency, an active coordination structure and an international presence are covered densely: administrative boundaries at multiple levels, population baselines, displacement figures, needs assessments, sectoral response data and market monitoring. Countries with chronic but undeclared crises, or with conflicts where international access is denied, are covered thinly and often by remote-monitoring proxies rather than by field data. Wealthy countries are largely absent because the humanitarian system does not operate there. Thematically the strength is in the operational categories the coordination system runs on – boundaries, population, displacement, food security, health facilities, education, water and sanitation, protection – and the weakness is in anything the system does not manage, including most economic, security and political data. Temporally the catalogue spans roughly the past decade with growing depth, and its update rhythm is entirely contributor-dependent: some datasets update daily by automated pipeline, many are annual, and a substantial number were uploaded once and abandoned. The honest summary is that HDX shows you the world as the international humanitarian system sees it, which is a partial view that is nonetheless the best organised one available.

Known blind spots

Absence of evidence here is not evidence of absence. These are the conditions under which Humanitarian Data Exchange (HDX) will not show you something that is nevertheless real:

  • HDX creates almost nothing. It is a catalogue, so every quality, methodology and bias question belongs to the contributing organisation, and a dataset's presence here confers no validation whatsoever.
  • Staleness is pervasive and semi-visible. Many datasets carry a declared update frequency that has not been honoured for years, and the metadata record can look current while the underlying file has not changed since it was uploaded.
  • Coverage follows humanitarian declarations and funding. Crises that are chronic, politically inconvenient or inaccessible to international agencies are systematically underrepresented, and their absence is a statement about access rather than about need.
  • Data on affected populations is deliberately restricted. The data responsibility framework limits what may be published about identifiable individuals and groups, so the most granular protection, trafficking and vulnerability data is not and should not be here.
  • Administrative boundaries are politically contested in exactly the places where they matter most, and the versions published carry UN framing and disclaimers that some parties to a conflict do not accept.
  • Multiple entries frequently describe the same underlying data, republished by different organisations at different times, so counting datasets or naively merging them double-counts and produces spurious agreement.
  • Licences are set per dataset and vary from open dedications to restrictive bespoke terms, so there is no catalogue-level permission and any assumption of blanket reusability is wrong.
  • Format quality varies from clean HXL-tagged CSV to scanned PDFs uploaded as data, and the catalogue metadata does not reliably distinguish between machine-readable data and a document that happens to contain numbers.
  • Aggregation and geographic units differ between contributors even within one country, so figures from two datasets covering the same place are frequently not comparable without reconciling to a common boundary set.

Write the blind spot into the product. A statement that something “was not observed in Humanitarian Data Exchange (HDX)” is defensible; a statement that it “did not happen” is not, and the difference is what survives cross-examination.

Access, licensing and what you may do with it

Access model: Open — no account required

Everything is open and unauthenticated for reading. The catalogue is a CKAN instance, so the standard action API is available: search datasets with facets and filters, retrieve a dataset record with its resources, list organisations and groups. Resources are ordinary files at ordinary URLs and can be downloaded directly once you have the record. There is also a harmonised interface offering standardised subsets of core data across countries, which reduces the per-dataset cleaning burden considerably for the categories it covers; check the current documentation on the site rather than assuming its scope. Practical guidance: search returns paginated results and you must page properly rather than assuming the first page is the answer, because relevance ranking in CKAN is not sophisticated enough to be trusted with a single-page read. Retrieve the full dataset record rather than working from search results alone, because the fields that determine whether a dataset is usable – licence, methodology, caveats, per-resource dates – are in the detail record. Contributing data, rather than consuming it, requires an account and is subject to review under the data responsibility framework.

Licence

There is no HDX licence. Each dataset carries the licence its contributor chose, and the range is wide: public domain dedications, attribution licences, attribution licences with intergovernmental organisation variants, and bespoke terms including some that permit viewing but not redistribution. This is the single most commonly mishandled aspect of the platform. An organisation that harvests broadly and republishes without checking each licence will be redistributing restricted material, and in a sector where the contributors are UN agencies and NGOs with reputational sensitivities, that is a real problem rather than a theoretical one. The correct engineering approach is to treat the licence field as mandatory on ingest, refuse to store or redistribute any resource whose licence you have not recorded, and carry the licence through to any export or publication. Where a dataset's licence is marked as other or unspecified, treat it as all rights reserved until the contributor says otherwise. Separately, the platform's own terms of service govern your use of the catalogue and its API, and the data responsibility framework governs what may be uploaded – relevant if you ever contribute.

Rate limits and fair use

No hard published quota, and the usual CKAN behaviour applies: search queries are relatively expensive and file downloads are cheap. The etiquette is straightforward. Use search facets and filters to narrow at the server rather than downloading everything and filtering locally. Page properly and with a sensible page size. Cache dataset records, which change infrequently, and check modification dates before re-downloading resources rather than re-fetching files blindly. If you are building a mirror or a broad harvest, run it on a schedule measured in days rather than minutes, spread it out, and set a descriptive user agent naming your organisation with a contact address. This is infrastructure that humanitarian responders use operationally during emergencies, and a harvester that degrades it during a crisis is causing real harm for no analytical gain. Where a harmonised or bulk interface exists for the data you want, use it in preference to iterating the catalogue.

Licensing changes, and it changes without warning. A dataset that was free for research this year may not be free for commercial or evidential use next year. Confirm the current terms before you build a dependency on it, and record the terms you relied on alongside the data — the licence in force at the time of collection is part of the provenance.

Collecting it

How Humanitarian Data Exchange (HDX) is actually pulled, in the order you would set it up. Prefer the bulk or export interface over per-item lookups wherever one exists: it is kinder to the publisher, faster for you, and gives a reproducible snapshot rather than a series of point-in-time answers you cannot reconstruct later.

Method Format Cadence Notes
Faceted catalogue search JSON daily to weekly The CKAN package search action with location, organisation, tag and format facets. The right way to discover what exists for a country or theme without downloading anything.
Dataset detail retrieval JSON on discovery and on change Full record including resources, licence, methodology and caveats. Never work from search results alone; the fields that determine usability are only in the detail record.
Resource download CSV per declared update frequency The files themselves. Check the per-resource modification date before re-downloading, and store the file with a hash so you can detect silent content changes.
Reference layer baseline bulk quarterly Pull the administrative boundaries and population baselines for your countries of interest once and treat them as a controlled reference layer, because every rate you compute depends on them being consistent.
Harmonised interface JSON as documented Where the standardised cross-country interface covers your data category, prefer it – it removes most of the per-contributor cleaning work and gives consistent units and geography.
Change monitoring JSONL weekly Record dataset modification dates on every run and diff them. The pattern of what updates and what does not is itself an indicator of where the humanitarian system is actually active.

Ingesting it into the platform

Every step below is idempotent and cursor-based: interrupt one and it resumes from where it stopped rather than duplicating rows or losing progress. Collection is recorded per source, so a feed that quietly stops publishing shows up as a stale timestamp instead of silently thinning your coverage.

  1. Register HDX as a catalogue, not a feed — In sources.php, record that this is a metadata catalogue over many independent publishers, so that every downstream record carries the contributing organisation as its real provenance rather than being attributed to HDX.
  2. Harvest metadata before data — Run collect.php against the search API to build a local index of dataset records first. Deciding what to download is an analytical act, and doing it from metadata is far cheaper than downloading files and discovering they are PDFs.
  3. Make licence a mandatory ingest field — In ingest.php, refuse any resource whose licence is not recorded, and store the licence identifier with every derived record. Redistributing restricted humanitarian data because a field was null is an avoidable and serious mistake.
  4. Separate the three dates — Store the data reference period, the upload date and the last modification date as distinct fields. Collapsing them into one currency indicator is the most common way analysts end up citing five-year-old figures as current.
  5. Pin the reference geography — Load administrative boundaries and population baselines as a controlled layer and reconcile every other dataset to it. Where a contributor used different units, record the mismatch rather than silently reprojecting or reallocating.
  6. Exploit HXL tags where present — Parse the hashtag row on tagged resources to map columns to standard concepts automatically, and flag untagged resources for manual schema mapping so the cost of using them is visible rather than hidden.
  7. Carry caveats through to the record — Attach the contributor's methodology and caveat text to every derived figure via enrich.php, so an analyst reading a number in a case sees the limitation the publisher stated rather than a bare value.
  8. Monitor declared versus actual updates — Configure alerts.php to flag datasets whose declared update frequency has lapsed, so staleness surfaces as a managed signal and your country baselines do not quietly become historical.

Registered sources and their last-collected state are listed in sources.php, and the scheduled chain that keeps them current is in automation.php.

How it is wrong, and how to tell

Every dataset is wrong in characteristic ways. Knowing which ways is the difference between using a source and being used by one, and it is the part of source evaluation most often skipped because it is the part that takes work.

Quality is not a property of HDX; it is a property of each contributor, and the platform's value is that it tells you who the contributor is. That makes assessment tractable but unavoidably manual. At the strong end, datasets from established statistical and monitoring programmes – conflict event databases with published coding rules, displacement tracking with documented survey methods, food price monitoring with defined market sampling – are as good as anything in their field and carry methodology documentation you can read. At the weak end are one-off uploads with no methodology note, no caveats, ambiguous geographic units and a declared update frequency that was never honoured. The distinguishing signals are consistent and easy to check: does the record carry a methodology, does it carry caveats, has it actually been updated in line with its declared frequency, is the resource machine-readable, does the contributor publish the same data elsewhere with more documentation. A dataset that passes all five is generally reliable. A dataset that fails three is a lead, not a source. The structural quality issue that applies to nearly everything here is that humanitarian data is collected under access constraints – where teams could go, when they could go, who they could speak to – and those constraints are correlated with severity, so the data systematically understates conditions in the worst-affected places.

Characteristic false positives

  • The dataset date, the upload date and the last modification date are three different things, and using the wrong one produces confident claims about current conditions based on data from years ago.
  • Presence in the catalogue is read as validation. HDX does not verify contributed data, so a badly constructed dataset sits alongside a rigorous one with identical visual weight in search results.
  • The same underlying data appears under several entries from different contributors, so merging without deduplication produces double counting and a false impression of independent corroboration.
  • Geographic units differ between contributors even within one country, so figures aggregated across datasets are frequently summed over incompatible boundaries without anyone noticing.
  • Population baselines are estimates with their own methodologies and vintages, and dividing an incident count by a mismatched baseline produces rates that look precise and are not comparable to anyone else's.
  • Access-constrained collection is read as complete coverage. A displacement or needs figure describes what was measurable where teams could operate, and its absence in a district usually means no access rather than no need.
  • Restricted or sensitive data is sometimes present in a form the contributor did not intend to be reidentifiable, and combining several innocuous datasets can reidentify populations that each dataset alone protected.
  • Files change in place without a version change, so a resource re-downloaded at the same URL can contain different content from the copy your analysis was built on unless you hashed it.

None of these make the source unusable. They make it a source that requires corroboration before an assertion built on it goes into a product, which is true of every source and admitted by few.

Ageing

Ageing varies by category more than by dataset, and knowing the category tells you most of what you need. Administrative boundaries age slowly and change discontinuously – a boundary reform or a new administrative level makes every dataset keyed to the old units incompatible overnight, and this happens more often than outsiders expect. Population baselines age on a census cycle where a census exists and on a projection cycle where it does not, and in protracted displacement situations they become badly wrong within a couple of years. Displacement figures, needs assessments and market prices age in weeks to months and are close to worthless outside their reference period, though they retain value as a historical series. Conflict event data ages not at all as history but is continuously revised as coding is corrected, so the version you downloaded is not the version available now. The catalogue metadata gives you the tools to see all of this if you use them: the declared update frequency states the contributor's intent and the modification history states what actually happened, and the gap between the two is the most honest staleness indicator available. A stale record here looks entirely healthy – a current-looking dataset page whose only resource has not changed in three years.

What this source feeds

A source is only worth what it lets you conclude. These are the disciplines that collect through it, the mission domains it serves and the data points it yields — every one is a tag, so you can follow any thread from here into the rest of the library.

Collected by these intelligence disciplines

Serves these mission domains

Yields these data points

How each sector uses Humanitarian Data Exchange (HDX)

The same dataset is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The records are shared — the constraints, thresholds and outputs are not.

🎖 Military and defence

For civil-military coordination, stabilisation and any operation with a population dimension, this is the reference layer: agreed administrative boundaries, population baselines, settlement locations, health and education facility inventories, and displacement figures. It is also the practical route to understanding where humanitarian actors are operating, which is a deconfliction requirement rather than an intelligence one and should be treated with the seriousness that distinction implies. Two constraints are non-negotiable. Humanitarian data is contributed under an expectation that it serves humanitarian purposes, and using it to support targeting or any activity that endangers affected populations or responders breaches that expectation and puts real people at risk. And humanitarian actors are frequently reluctant to be associated with military users, so consuming the data is unproblematic while approaching the contributors may not be.

🕵 National intelligence

The analytical contribution is context and denominators rather than events. Population baselines, boundaries and settlement data are what turn a raw incident count into a rate, and rates are what make comparisons across places and times meaningful. Displacement data is a genuine indicator stream in its own right – population movement is one of the earliest and most reliable observable responses to deteriorating security, often preceding reporting – and food security classifications track the humanitarian consequences of conflict and economic disruption with a documented methodology. For OSINT and GEOINT work the catalogue is also a discovery tool that finds datasets you did not know existed for a country of interest. The discipline point is that this data is collected by organisations operating on the basis of neutrality and access negotiated with parties to a conflict, and analytical use that could be perceived as compromising that neutrality has consequences for the people who collect it.

👮 Law enforcement

The direct law enforcement applications are in transnational cases with a displacement or humanitarian dimension: understanding population movement corridors in trafficking investigations at a strategic level, establishing the administrative geography that jurisdictional questions turn on, and contextualising conditions in origin and transit locations. The essential constraint is that individual-level data about affected populations is deliberately excluded from this platform under the data responsibility framework, and that exclusion is protective rather than incidental. Nothing here should be used to locate, identify or approach individuals, and any investigative work touching trafficking or exploitation should run through the established referral mechanisms and national hotlines rather than through data mining. Treat the catalogue as strategic context, not as a source of leads about people.

🔍 Private investigation and corporate security

Corporate security, due diligence and supply chain risk work uses this for country and subnational context that commercial risk products summarise but do not source: which districts are affected by displacement, where food insecurity is classified as severe, what the actual administrative geography of an operating area is, which facilities exist. It is free, it is sourced, and it lets you check the claims in a commercial risk report against the underlying data. The limitations to hold onto are that the catalogue covers humanitarian rather than commercial concerns, that it is thin outside declared emergencies, and that currency varies enormously between datasets. Never present a humanitarian dataset as a security assessment; it describes need and response, not threat.

📰 Journalism and OSINT media

For reporting on crises this is a primary source that supports specific, checkable claims and that can be linked to directly in a published piece. Its particular value is in supplying the denominators that make numbers meaningful – a displacement figure means something different against a population baseline than in isolation – and in providing boundary files for maps that match what the humanitarian system actually uses. The reporting practices that matter are to attribute to the contributing organisation rather than to HDX, to state the reference period rather than the download date, to read and reflect the caveats the contributor supplied, and to be alert that administrative boundaries in contested areas carry political framing that will be contested by someone. Where a dataset has not been updated in line with its declared frequency, say so rather than presenting it as current.

🌍 NGO, humanitarian and human rights

For humanitarian and human rights organisations this is infrastructure rather than a source, and the main advice is about contributing as well as consuming. Use the Common Operational Datasets as your reference geography so that your figures are comparable with everyone else's, which is the entire point of their existence. Check the data responsibility framework before uploading anything derived from work with affected populations, because the reidentification risk from combining datasets is real and the framework exists because it has happened. Where you consume, prefer datasets with methodology and caveat documentation and be sceptical of figures with neither. And recognise the access bias in your own use: the districts with no data are usually the ones with no access, and reporting on them as though data absence indicates low need inverts the truth.

🎓 University and research

The catalogue supports development economics, conflict studies, public health, migration research and geography, and its CKAN foundation makes systematic harvesting straightforward. Three methodological points should appear in any paper using it. First, cite the contributing organisation and the specific dataset version and reference period, not the platform, because the platform is a repository and the science belongs to the contributor. Second, address the selection mechanism explicitly – inclusion tracks humanitarian declarations, funding and access, so a panel built from HDX is a sample with a non-random and correlated missingness structure. Third, address boundary and baseline consistency, because analyses that pool across countries or years using whatever boundary file was available produce results that are not reproducible. Archive the exact resource files used, since resources change in place at stable URLs.

Playbook: working Humanitarian Data Exchange (HDX) end to end

A repeatable sequence from first pull to finished product. Each phase states what you are trying to establish, not merely what to click — the objective is a defensible chain of reasoning, not a completed checklist.

Phase 1 — Decide what layer you actually need

Separate reference layers – boundaries, population, settlements, facilities – from measurement layers such as displacement, needs and prices. Reference layers should be pinned once and used consistently across the whole analysis; measurement layers are time-bounded observations. Mixing the two conceptually is what produces incomparable figures later.

Phase 2 — Search the catalogue before searching the web

Use faceted search by location, organisation, tag and format to establish what exists for your country and theme. Page through results properly rather than reading the first page, because CKAN relevance ranking will not surface the best dataset reliably and the one you need is often not on page one.

Phase 3 — Read the full record, not the search result

Retrieve each candidate dataset's detail record and read the contributor, methodology, caveats, licence, declared update frequency and per-resource dates. This is where you decide whether the dataset is usable, and it takes minutes; downloading first and evaluating later wastes hours.

Phase 4 — Establish the three dates and the real currency

Determine the reference period the data describes, when it was uploaded and when the resource actually last changed. Compare the declared update frequency against the observed modification history. A dataset declaring monthly updates whose file has not moved in two years is a historical record and should be labelled as one.

Phase 5 — Pin the geography before anything else

Adopt one administrative boundary set and one population baseline for the country and record their vintage. Every other dataset must be reconciled to these, and where a contributor used different units the reconciliation must be documented rather than performed silently. Skipping this step invalidates every rate and comparison downstream.

Phase 6 — Assess the contributor, not the dataset

Look up the contributing organisation's other datasets and its own published methodology. Established monitoring programmes with public coding rules are a different class of evidence from one-off uploads, and the contributor is a more reliable quality signal than anything in the dataset record itself.

Phase 7 — Deduplicate before combining

The same underlying data frequently appears under multiple entries from different contributors. Identify the original producer for each figure and drop the republications, because merging them creates double counts and manufactures the appearance of independent corroboration where there is none.

Phase 8 — Use HXL tags where they exist and cost the alternative honestly

Tagged resources can be aligned to a common schema automatically. Untagged resources need manual column mapping, which is real work that must be budgeted rather than assumed away. Deciding which untagged datasets are worth the mapping cost is an analytical decision, not a data engineering one.

Phase 9 — Model the access bias explicitly

Ask, for every geographic gap in your data, whether it reflects low need or no access. In conflict settings it is nearly always the latter, and the districts with the worst conditions are the ones with the least data. Write this into the analysis as a stated direction of bias rather than treating missing values as zeros.

Phase 10 — Corroborate against independent measurement

Test humanitarian figures against sources with different collection mechanisms – conflict event databases, remote sensing, market and price data, media reporting – and treat convergence as meaningful and divergence as a question. Humanitarian reporting has known incentive pressures around funding appeals, and independent measurement is what disciplines it.

Phase 11 — Record licence and provenance with every figure

Carry the contributing organisation, dataset identifier, resource hash, reference period and licence through into the case record and any export. Humanitarian data is contributed under varied terms and the reputational cost of redistributing restricted material in this sector is high and falls on real relationships.

Phase 12 — Freeze what you used

Archive the exact resource files your analysis rests on, with hashes and retrieval times, because resources change in place at stable URLs and a rerun three months later will silently produce different results. Nothing about the catalogue's design prevents this and only your own discipline does.

The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.

What to pair it with

No single source carries a finding. These are the datasets that corroborate, extend or contradict this one — and a source that contradicts is worth more than one that agrees, because it is the only thing that will tell you when you are wrong.

Source Relationship What it adds
ReliefWeb extends The narrative and documentary counterpart from the same parent organisation – situation reports and analysis that explain what the numbers in HDX mean and who collected them.
ACLED corroborates Coded conflict event data with published methodology, providing an independent measurement of violence against which humanitarian need figures can be tested.
UCDP corroborates Academic conflict data with conservative inclusion criteria and a long time series, useful precisely because its coding rules differ from event-based commercial datasets.
IOM Displacement Tracking Matrix extends The primary displacement measurement programme whose outputs populate a large share of the movement data in the catalogue, with its own methodology documentation.
IPC extends The food security classification system that turns market and nutrition data into a comparable severity scale, and the reference for interpreting food security datasets.
FEWS NET corroborates Independent famine early warning analysis with its own methodology and a different institutional sponsor, valuable as a cross-check on food security classifications.
GDACS extends Automated disaster alerting and impact estimation that provides a rapid first indication before humanitarian datasets exist for a new event.
geoBoundaries corroborates An independent, openly licensed administrative boundary database, useful when the humanitarian boundary set is contested or when a permissive licence is required.
Humanitarian OpenStreetMap Team extends Community-mapped infrastructure, buildings and roads in crisis areas, filling the gaps between official facility inventories and reality on the ground.

Legal, ethical and operational constraints

Three distinct legal and ethical layers apply. The first is licensing, which is per dataset and genuinely varied, and which must be recorded and honoured rather than assumed – this is the constraint that most often bites organisations that harvest broadly. The second is data protection: although the platform's data responsibility framework restricts publication of personal and demographically identifiable data, humanitarian datasets can still be sensitive in aggregate, and in most jurisdictions the combination of several datasets that reidentifies a small population is your legal problem regardless of each dataset's individual compliance. Assess reidentification risk before combining, particularly for datasets covering minority groups, displaced populations or protection cases. The third layer is the humanitarian principles under which the data was collected. This material exists because organisations negotiated access on the basis of neutrality, impartiality and independence, and using it in ways that could be perceived as serving military, intelligence or law enforcement targeting endangers both affected populations and the responders who collect it. In most jurisdictions nothing makes that unlawful; it is nonetheless a line that should be drawn deliberately, documented, and enforced within your organisation.

Operational security

Catalogue searches are unauthenticated but logged, and a query pattern focused on a specific country, district or theme is a legible statement of interest to anyone with access to those logs or to your network traffic. In a humanitarian context that exposure has a particular character: the parties to a conflict, and the states hosting it, have a demonstrable interest in who is examining data about their territory. The mitigations are ordinary – harvest broadly on a routine schedule rather than querying narrowly when interest spikes, mirror what you need locally so that repeated analysis touches no network, and avoid query cadences that correlate with operational events. The more serious exposure is downstream: publishing analysis that names specific locations, facilities or contributing organisations can put field staff at risk, and in some contexts identifying which organisation supplied a dataset is enough to endanger it. Consider whether the provenance you are correctly recording internally should also appear in an external product.

Two rules that hold regardless of jurisdiction. Collection that is lawful is not automatically proportionate, and a dataset assembled for one purpose does not carry consent for another. Where the records concern identifiable people, the question is not only whether you may hold the data but whether holding it serves the purpose you are accountable for.

Is it earning its place?

Sources accumulate. Feeds get added during an incident and are never reviewed again, and a decade later the pipeline is carrying dead weight that nobody dares remove. These are the measures that show whether Humanitarian Data Exchange (HDX) is contributing anything, and they are worth baselining now so the answer is available later.

  • Proportion of ingested resources with a recorded licence, which must be complete and is the single most consequential compliance measure for a catalogue of this kind.
  • Gap between declared update frequency and observed modification interval, per dataset, aggregated into a staleness score for your country baselines.
  • Share of datasets in use that carry contributor-supplied methodology and caveats, tracked as a quality profile of your own source selection rather than of the catalogue.
  • Number of analyses in which the reference geography was pinned to a single documented boundary and baseline set, as a process measure of whether comparability is being maintained.
  • Rate at which harvested datasets turn out to be non-machine-readable documents, which measures how much of the apparent catalogue volume is actually usable data.
  • Count of findings where a humanitarian dataset supplied a denominator or context that materially changed an assessment derived from other sources.
  • Number of resource content changes detected by hash without a corresponding metadata update, which quantifies how much silent revision your pipeline is absorbing.

Beware of volume. Indicator counts rise easily and say almost nothing. Unique contribution — findings this source produced that no other source in your stack would have — is the measure that matters, and it is usually far lower than anyone expects.

Tradecraft notes

The distinctions that separate a competent analyst from a fast one:

  • Attribute to the contributor, never to the catalogue. HDX is a repository; the methodology, the incentives and the responsibility for every number belong to the organisation that uploaded it, and a citation to the platform hides the only fact that matters.
  • Three dates, always. Reference period, upload date and last modification are distinct, and the discipline of recording all three is what prevents an analysis from silently describing conditions from several years ago.
  • Pin the geography first and never change it mid-analysis. Boundary and population baselines are the denominators under everything, and mixing vintages produces figures that are individually correct and collectively meaningless.
  • Declared update frequency versus actual modification history is the best staleness signal in the catalogue, and it is free. A dataset that promises monthly updates and has not moved in two years is telling you something important about the programme behind it.
  • Missing data in a conflict zone means no access far more often than it means no need, and treating geographic gaps as zeros inverts the direction of your error in exactly the districts that matter most.
  • Deduplicate to the original producer before combining anything. Republication across organisations manufactures the appearance of independent corroboration, which is the most dangerous artefact this catalogue produces.
  • Read the caveats and carry them forward. Contributors in this sector are unusually honest about limitations, and consumers are unusually good at discarding that honesty at the point of extraction.
  • Hash every resource you download. Files change in place at stable URLs with no version signal, and an analysis you cannot reproduce because the underlying file moved is not an analysis.
  • Know where the line is on use. This data is contributed under humanitarian principles by organisations whose access depends on being seen as neutral, and there are applications that are technically permitted and professionally indefensible.

Questions analysts actually ask

Does HDX verify the data it publishes?

No, with narrow exceptions. It is a catalogue, and the contributing organisation is responsible for accuracy and methodology. OCHA sets terms and enforces a data responsibility framework governing what may be uploaded, particularly regarding personal and identifiable data, but presence in the catalogue is not validation of a dataset's quality.

What is the API and do I need a key?

It is a CKAN instance, so the standard CKAN action API applies – dataset search with facets, dataset detail, organisation and group listings – and no key is required for reading. Anything you have written against another CKAN portal works here. There is also a harmonised interface offering standardised subsets across countries for some data categories; check its current scope on the site.

Can I redistribute what I download?

It depends entirely on the dataset. Licences are set per dataset and range from public domain dedications through attribution licences to bespoke restrictive terms. There is no catalogue-wide permission. Record the licence on ingest, treat unspecified licences as all rights reserved, and carry the licence through to any export.

How do I tell whether a dataset is current?

Compare three things: the reference period the data describes, the last modification date of the actual resource, and the contributor's declared update frequency. A record can look current while its only file has not changed in years. The gap between declared and observed update behaviour is the most useful currency signal available.

What are HXL tags and why do they matter?

They are machine-readable hashtags placed in a row beneath the column headers of tabular data, annotating each column with a standard concept such as location, sector, date or population. They let files from different organisations be combined automatically without hand-mapping every column, which is the difference between an afternoon and a week when working across many datasets.

Why are administrative boundaries controversial?

Because in conflict and disputed-territory settings the boundary is the argument. The published sets carry UN framing and disclaimers that parties to a conflict do not necessarily accept, and using them in a public product is a choice with political implications. Where this matters, state which boundary set you used and why, and be aware that independent alternatives exist with different framing.

Can I find data on individual affected people?

No, and you should not try. The data responsibility framework specifically restricts publication of personal and demographically identifiable data about affected populations, and that restriction exists because reidentification has caused real harm. Casework about individuals belongs with the appropriate protection agency or national referral mechanism, not with an open data catalogue.

Why does a country I care about have almost nothing?

Usually because the humanitarian system is not operating there at scale – no declared emergency, no coordination structure, or no access. Absence in the catalogue reflects the international system's presence and attention rather than the level of need, and in the hardest places those two are inversely related.

Is it appropriate for military or intelligence use?

Consuming published open data is not restricted, but the professional line matters. This data exists because humanitarian organisations negotiated access on the basis of neutrality and independence, and uses that could be perceived as supporting targeting or that endanger responders and affected populations breach the expectation under which it was contributed. Draw the line explicitly and enforce it internally rather than leaving it to individual judgement.

Standards, formats and interoperability

What this source speaks natively, and what it has to be translated into before a partner can consume it. Work that arrives in a recognised format is easier to defend, easier to hand over and easier to automate against:

  • CKAN is the underlying platform, so the action API, dataset and resource model and faceted search behaviour follow CKAN conventions and are documented in the CKAN project's own documentation.
  • HXL, the Humanitarian Exchange Language, provides the hashtag vocabulary that makes tabular humanitarian data machine-combinable across contributing organisations.
  • The Common Operational Datasets define the reference administrative boundaries, population baselines and settlement data that humanitarian coordination in a country is built on.
  • P-codes are the place-code identifiers linking datasets to administrative units, and they are the practical join key between humanitarian datasets for the same country.
  • IPC and the Cadre Harmonisé are the classification systems that make food security severity comparable across countries and seasons.
  • The OCHA data responsibility guidelines set the framework governing what may be published about affected populations and how sensitivity is assessed.
  • GLIDE numbers provide a shared identifier for disaster events, linking datasets here to disaster records in other humanitarian information systems.
  • The platform exports derived entities, locations and events in STIX 2.1, MISP, CSV, JSON and JSONL, so humanitarian context travels into a case alongside technical and conflict observables.

References

Primary documentation and authoritative references for this source. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.

  1. Humanitarian Data Exchange — UN OCHA Centre for Humanitarian Data. The catalogue itself. Start with the country pages to see how coverage is actually organised before writing any harvesting code.
  2. Centre for Humanitarian Data — UN OCHA. The team behind HDX, and the home of the data responsibility guidance, data grids concept and the standards work that shapes what is in the catalogue.
  3. CKAN API documentation — CKAN Association. The authoritative reference for the action API HDX exposes. Read this rather than reverse-engineering the endpoints from network traffic.
  4. HXL Standard — Humanitarian Exchange Language. The hashtag vocabulary and tooling that make tabular humanitarian data combinable across organisations without manual column mapping.
  5. UN Office for the Coordination of Humanitarian Affairs — UN OCHA. The parent organisation, its coordination role and the humanitarian principles that govern how this data may appropriately be used.
  6. ReliefWeb — UN OCHA. The narrative counterpart. Situation reports and analysis that explain the context behind the datasets and identify which organisations are collecting what.
  7. IOM Displacement Tracking Matrix — International Organization for Migration. The methodology documentation behind a large share of the displacement data in the catalogue, and essential reading before citing any movement figure.
  8. Integrated Food Security Phase Classification — IPC Global Partners. The classification framework that makes food security severity comparable, and the reference for interpreting any phase classification found in the catalogue.
  9. ACLED — Armed Conflict Location and Event Data Project. Independent conflict event data with published coding rules, the standard cross-check against which humanitarian need figures should be tested.
  10. geoBoundaries — William and Mary geoLab. Independently maintained open administrative boundaries, useful when the humanitarian boundary set is contested or when licensing requires an alternative.
  11. Humanitarian OpenStreetMap Team — HOT. Community mapping of infrastructure in crisis areas, filling gaps between official facility inventories and observable reality.

Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.

Put it into practice

The Quantus Intel threat intelligence platform operationalises this source: it harvests catalogue metadata before data, makes licence and contributing organisation mandatory on every ingested record, keeps reference period separate from modification date, pins administrative boundaries and population baselines as a controlled layer, and flags datasets whose declared update frequency has quietly lapsed.. Browse the full source catalogue, or follow any tag above into the rest of the library.

Leave a Reply