UCDP Georeferenced Events: Intelligence Source Guide
UCDP GED records individual incidents of organised violence worldwide from 1989 onward, each geocoded to a point, dated to a day where possible, attributed to named actors and carrying low, best and high fatality estimates. It is the most conservative conflict event dataset in general use, and th…
UCDP GED records individual incidents of organised violence worldwide from 1989 onward, each geocoded to a point, dated to a day where possible, attributed to named actors and carrying low, best and high fatality estimates. It is the most conservative conflict event dataset in general use, and the conservatism is the point.
At a glance
| Source | UCDP Georeferenced Events |
|---|---|
| Category | Conflict, Crime & Human Security › Conflict & Event Databases |
| Homepage | https://ucdp.uu.se/ |
| Machine interface | https://ucdpapi.pcr.uu.se/api/gedevents/24.1 |
| Format | JSON |
| Access | Free registration — API key at no cost The platform catalogue records this as open. That is wrong, and the correction is explained under Access, licensing and what you may do with it below. |
| Disciplines | Geospatial Intelligence, Open Source Intelligence |
| Mission domains | Conflict & Humanitarian |
Uppsala organized-violence event dataset. — as catalogued in the platform’s own source registry.
The Uppsala Conflict Data Program Georeferenced Event Dataset is a hand-coded record of discrete incidents of organised violence. An event is defined narrowly: the use of armed force by an organised actor against another organised actor, or against civilians, resulting in at least one direct death, at a specific place on a specific date. Each row carries the parties on both sides with stable numeric identifiers, the conflict and dyad the event belongs to, a point coordinate with a precision code, a start and end date with a precision code, four disaggregated death counts (side A, side B, civilians, unknown), and three aggregate fatality estimates labelled low, best and high. Every event also carries the source articles it was coded from, concatenated into a single string, and a clarity code recording whether the incident was reported as a distinct event or disaggregated by coders from an aggregated report. The dataset is released in annual versions and served through a REST API at ucdpapi.pcr.uu.se in which every past version remains permanently retrievable, so a query embedded in replication material returns the same rows years later. As of version 26.1 the API reports 417,968 events. Alongside the annual release, UCDP publishes GED Candidate, a monthly and quarterly preliminary series coded to the same rules but not yet through the full annual quality-control cycle.
The analytical job GED does that nothing else does is to make an absolute, auditable claim about what counts. Most conflict data products are permissive: they ingest reports, they include protests, riots, arrests and strategic developments, and they let the user filter afterwards. GED is the opposite. It admits only incidents that meet a published definitional threshold, it requires that the incident be attributable to an actor registered in UCDP's own actor list, and it requires at least one death the coders judge to be a direct result of the violence. That exclusion discipline is what makes GED usable as a denominator. If you need to say that battle deaths in a country rose or fell between two years, that a dyad became more lethal after an intervention, or that a peace agreement was followed by a measurable change in organised violence, GED is the series that survives review, because the coding rules did not change under you and every event is traceable to the reporting it came from. It is the natural anchor for GEOINT work on conflict geography, since every row carries a PRIO-GRID cell identifier, and for OSINT work it functions as a calibrated baseline against which noisier, faster feeds can be judged. Use ACLED or GDELT when you need breadth and speed. Use GED when the number has to hold.
Who publishes it, and why that matters
UCDP sits in the Department of Peace and Conflict Research at Uppsala University, with the flagship UCDP/PRIO Armed Conflict Dataset produced jointly with the Peace Research Institute Oslo. The operating model is academic: research staff and trained coders read source material, apply a published codebook, and release annual versions with version histories alongside peer-reviewed articles documenting what changed. Plan around three consequences. Funding is grant-and-institution based rather than commercial, so there is no service contract and no committed response time, but the institutional base has been stable for decades and the dataset is embedded in enough published research that its continuity is effectively underwritten by the field. Coding is expensive and slow, which is why the annual release lags the year it covers and why the candidate series exists as a stopgap. And the API is maintained by a named individual in the department: as of the current documentation it requires an access token requested by email with a short description of your project, reviewed and answered within a few working days. Treat that as a research-infrastructure relationship rather than a vendor relationship. It is free, it is generous, and it depends on users not abusing it.
Provenance is the first question to ask of any dataset and the one most often skipped. Who collects it, what their incentive is, whether they publish a methodology, and whether they correct the record when they get something wrong all bear directly on how much weight a finding drawn from it can carry.
What a record actually contains
The fields you will be working with, what each one means, and whether it is something you can pivot on. Read the meanings carefully — more analysis is wrecked by misreading a field than by failing to find one, and a field that looks like an observation is often an inference.
| Field | Type | What it means | Pivot value |
|---|---|---|---|
id |
int | Event identifier within a version. Stable across versions for events that survive re-coding, but events can be merged, split or removed between versions, so it is not a permanent key. | Direct record retrieval; version-to-version diffing to see exactly what UCDP changed. |
relid |
string | Human-readable composite identifier encoding country, year and internal coding references. Useful for spotting which coding batch an event came from. | Grouping events coded together; tracing back to the coding office. |
type_of_violence |
enum | 1 state-based (a government is one party), 2 non-state (neither party is a government), 3 one-sided (an organised actor kills civilians). This one field determines which of UCDP's three separate definitional universes the event belongs to. | The corresponding yearly datasets: ucdpprioconflict, nonstate, onesided. |
conflict_new_id / dyad_new_id |
int | The conflict and the specific pair of belligerents. UCDP renumbered these at GED version 5.0, and the API filters use the new numbering rather than the historical one. | Dyad-year and conflict-year datasets; actor histories running across decades. |
side_a / side_b |
string | Named parties, each with a numeric identifier. Governments appear as Government of X, and that identifier persists across changes of regime unless UCDP judges the state itself to have changed. | UCDP actor list; Gleditsch and Ward country codes via gwnoa and gwnob; proscription and sanctions lists. |
latitude / longitude / geom_wkt |
string | Point coordinate of the event. It is the coordinate of the place named in the source reporting, not a surveyed location of the violence. | PRIO-GRID cell, administrative units, terrain and infrastructure layers, other geocoded event feeds. |
where_prec |
enum | How precisely the location is known, running from an exact known point through to only the country being known. The most under-used field in the dataset and the one that most often invalidates a map. | Quality gating for spatial analysis; deciding whether an event may legitimately be aggregated to a district. |
date_start / date_end / date_prec |
timestamp | The window within which the event occurred, plus a precision code from exact day to year only. An event coded to a month is stored with a start and end spanning that month. | Temporal joins with other event feeds; sequencing analysis, but only for the highest precision codes. |
best / low / high |
int | The fatality estimate UCDP considers most plausible, and the conservative and generous bounds. Best is not a mean of low and high and must never be treated as one. | Battle-related deaths dataset; casualty comparison against ACLED, medical and mortuary records. |
deaths_a / deaths_b / deaths_civilians / deaths_unknown |
int | Disaggregation of the best estimate by whose dead they were. Deaths_unknown absorbs everything the reporting did not attribute, and in many conflicts it is the largest of the four. | Civilian-harm analysis; comparison with protection-of-civilians reporting. |
event_clarity |
enum | Whether the incident was reported as discrete (1) or was disaggregated by coders out of an aggregated report covering a period or an area (2). Clarity 2 events are real, but their date and place are constructions. | Filtering for incident-level analysis; explaining apparent clustering. |
active_year |
boolean | Whether the conflict or dyad met UCDP's annual activity threshold that year. GED includes events from inactive years for dyads active in some other year, which is a frequent source of confusion. | Reconciling GED event counts against the yearly conflict datasets. |
source_article |
string | The reporting the event was coded from, concatenated with a triple-slash separator. This is the audit trail and it is what makes GED defensible. | Source-media analysis; checking whether a spike is real or is one wire agency filing twice. |
priogrid_gid |
int | The PRIO-GRID cell the coordinate falls in, precomputed. Saves a spatial join and guarantees your grid assignment matches everyone else's. | PRIO-GRID covariates: population, terrain, nightlights, ethnic settlement, infrastructure. |
Coverage — and what is not in it
Global from 1989 to the most recent completed annual release, with no geographic exclusions in principle. In practice coverage is the product of two filters applied in sequence. The first is definitional: only organised violence with at least one direct death, only actors that pass UCDP's organised-actor test, only violence belonging to a conflict or actor UCDP has registered. Communal violence between groups UCDP does not judge organised, criminal homicide, deaths from disease and displacement, and the enormous indirect mortality of war are all outside the universe by design. The second filter is source availability: coding is driven primarily by global newswire reporting supplemented with NGO, IGO and specialist material, so the density of the record tracks the density of the press. Update rhythm is annual for the definitive series, with GED Candidate released monthly and quarterly under the same coding rules but before final quality control. Versions are cumulative and each remains permanently addressable through the API, which is unusual and valuable: a citation to gedevents/24.1 returns the same rows indefinitely even after 26.1 supersedes it.
Known blind spots
Absence of evidence here is not evidence of absence. These are the conditions under which UCDP Georeferenced Events will not show you something that is nevertheless real:
- Violence that killed nobody is invisible. Shelling that misses, a failed ambush, a raid that displaces a village without a fatality, and the sustained coercion that precedes mass violence all fall outside the definition, so a conflict can escalate substantially without moving the event count.
- Indirect deaths are excluded entirely. Famine, collapsed health services, disease and exposure kill far more people in most modern conflicts than combat does, and none of it is here; treating GED totals as war mortality understates the human cost by a factor that varies enormously between conflicts.
- An actor UCDP has not registered as organised produces no events at all. Communal, vigilante and much cartel and militia violence sits at the boundary of the organised-actor test, so its presence or absence in your extract reflects a coding judgement rather than the ground.
- The twenty-five battle-death annual threshold governs which dyads exist in UCDP's universe, so low-intensity insurgencies that never cross it are absent even though people died, and a conflict entering the data for the first time may have been running for years.
- Where press access is restricted, closed or lethal, event density collapses. Parts of the Sahel, northern Myanmar, closed states and the interiors of several long-running African conflicts are systematically thinner than the violence warrants, and the thinning is worst exactly where repression works best.
- Coordinates are place-name geocodes, not observations. An event coded to a district capital because the source named only the district is a point on your map carrying no more spatial information than the district name did, and where_prec is the only thing that tells you so.
- Events coded from aggregated reporting are constructed. When a source says fifty people were killed across a province during a month, coders must distribute that into events, and the resulting rows have plausible but manufactured dates and locations.
- The annual release lags. For roughly the last eighteen months you are working with candidate data or with nothing, and candidate data is explicitly subject to revision including deletion.
- Attribution is to the dyad, not to the unit. GED tells you a government fought a named rebel group; it will not tell you which formation, which foreign advisers were present, or whether a proxy relationship existed.
Write the blind spot into the product. A statement that something “was not observed in UCDP Georeferenced Events” is defensible; a statement that it “did not happen” is not, and the difference is what survives cross-examination.
Access, licensing and what you may do with it
Access model: Free registration — an account or API key, at no cost
The dataset is free and available three ways. The download centre at ucdp.uu.se offers full versions as CSV, Excel, Stata, R and shapefile bundles with the codebook alongside, and for most analytical work this is the right route: you get the whole thing, you can diff it against last year, and you are not dependent on network availability. The REST API at ucdpapi.pcr.uu.se serves the same data with server-side filtering by country, bounding box, date range, violence type, dyad and actor, and it now requires an access token supplied in an x-ucdp-access-token header. Tokens are free and issued on request by email to the API maintainer with a short statement of intended use; the documentation says requests are answered within three to five working days. Because the token travels as a custom header, the API cannot be exercised from a browser address bar, which surprises people working from older write-ups. The third route is the UCDP web interface, a good orientation tool and a poor collection tool. Whichever route you use, download the codebook for the exact version you are working with and keep it beside the data.
Licence
UCDP makes the data freely available and the terms centre on attribution and correct citation of the specific dataset and version, with the associated Journal of Peace Research articles named as the citations of record. That is a light-touch regime by the standards of official statistics and it has supported an enormous downstream literature. It is not a blanket grant: confirm the current terms on the download page before building a commercial product on the data or redistributing it in bulk, because a research infrastructure's terms are set by its institution and can change. The stronger practical obligation is scientific rather than legal. If you publish a figure derived from GED, state the version, the violence types included, which fatality estimate you used and any precision filtering applied. All four change the answer and none of them is visible in the number itself.
Rate limits and fair use
The documented quota is five thousand requests per day per token, and errors count against it. That is generous for targeted work and inadequate for bulk extraction: paging through the whole of GED at a modest page size will exhaust it. The correct pattern is to take the bulk download for anything resembling a full extract and reserve the API for targeted, filtered, repeatable queries of the kind you would embed in replication material. Use large page sizes, iterate using the TotalPages value returned in the response rather than guessing, and note that requesting a page beyond the end of a result set can return a server error rather than an empty set. Cache aggressively: annual versions are immutable, so anything fetched from version 25.1 last month is still correct today and never needs re-fetching.
Licensing changes, and it changes without warning. A dataset that was free for research this year may not be free for commercial or evidential use next year. Confirm the current terms before you build a dependency on it, and record the terms you relied on alongside the data — the licence in force at the time of collection is part of the provenance.
Collecting it
How UCDP Georeferenced Events is actually pulled, in the order you would set it up. Prefer the bulk or export interface over per-item lookups wherever one exists: it is kinder to the publisher, faster for you, and gives a reproducible snapshot rather than a series of point-in-time answers you cannot reconstruct later.
| Method | Format | Cadence | Notes |
|---|---|---|---|
| Full version download | CSV | annual, on release | The default for analysis. Take the CSV or the statistical-package bundle, keep the codebook with it, and record the version string in your project metadata. |
| REST API, filtered | JSON | on demand; results are immutable per version | Best for reproducible extracts and for pulling one country or bounding box into a case. Requires the token header; server-side filters cover country, geography, dates, violence type, dyad and actor. |
| GED Candidate | JSON | monthly and quarterly | Preliminary coding for the current year. Use it for currency, mark it provisional everywhere it appears, and expect events to change or vanish at the annual release. |
| Yearly companion datasets | CSV | annual | ucdpprioconflict, dyadic, nonstate, onesided, battledeaths and the country-year organised violence dataset. Pull these alongside GED; they answer aggregate questions GED answers badly. |
| Shapefile release | bulk | annual | Prebuilt spatial layer for GIS work, which avoids the coordinate-parsing errors that plague hand-rolled conversions. |
Ingesting it into the platform
Every step below is idempotent and cursor-based: interrupt one and it resumes from where it stopped rather than duplicating rows or losing progress. Collection is recorded per source, so a feed that quietly stops publishing shows up as a stale timestamp instead of silently thinning your coverage.
- Register the version, not just the source — In sources.php, record UCDP GED with the exact version string as part of the source identity. Two versions are two different datasets, and conflating them silently is the most common ingestion error with this source.
- Pull with a filter that matches the case — Configure collect.php to retrieve by country identifier, bounding box or dyad rather than pulling everything, and let cron.php own the schedule so a new annual release arrives as a job rather than as a surprise.
- Preserve the precision codes as first-class attributes — During ingest.php, carry where_prec, date_prec and event_clarity through into the stored record and expose them in the interface. An event that has lost its precision codes has become a false claim about a place and a day.
- Split identity from geography — Store actor identifiers, dyads and conflicts as entities in their own right so actor histories can be assembled across decades, and keep the coordinate as an attribute of the event rather than as its identity.
- Attach the grid cell and administrative units — Use the supplied priogrid_gid and resolve administrative units through enrich.php, so events can be aggregated the way analysts actually ask for them without re-running a spatial join each time.
- Reconcile against the yearly datasets — Run correlate.php between GED event aggregates and the battle-related deaths and conflict-year series. Systematic divergence is expected and informative; unexplained divergence usually means a filter in your own pipeline that you have forgotten.
- Build the case timeline and the theatre view — Push filtered events into timeline.php and theater.php so the event series renders alongside other collection, then attach the extract to a case in cases.php with the version string and the exact query recorded.
- Diff on every new release — When a new version lands, run it against the stored prior version and surface added, removed and materially changed events. UCDP revises history, and a finding built on last year's coding may no longer be supported by this year's.
Registered sources and their last-collected state are listed in sources.php, and the scheduled chain that keeps them current is in automation.php.
How it is wrong, and how to tell
Every dataset is wrong in characteristic ways. Knowing which ways is the difference between using a source and being used by one, and it is the part of source evaluation most often skipped because it is the part that takes work.
GED is as good as hand-coded conflict event data gets, and its weaknesses are structural rather than sloppy. Coding is done by trained staff against a published codebook, with an annual quality-control cycle, documented version histories and a peer-reviewed literature explaining what changed and why. Inter-coder consistency is taken seriously in a way automated event extraction cannot match, and the source_article field means any individual event can be audited back to the reporting. What the rigour buys is precision at the cost of recall: the dataset is deliberately biased towards excluding events it cannot substantiate, so it undercounts systematically and predictably rather than overcounting erratically. Judge it accordingly. A GED event almost certainly happened. The absence of a GED event is weak evidence that nothing happened, and in a low-press environment it is no evidence at all. The fatality estimates deserve particular respect: the low and high bounds exist because the coders genuinely do not know, and analyses that discard them and use best alone throw away the most honest thing in the dataset.
Characteristic false positives
- Best is treated as a measurement. It is a coder's judgement between two bounds that may differ by an order of magnitude, and summing best across events with wide bounds produces a total with false precision that nobody downstream will question.
- Aggregated reporting becomes fake incidents. Events coded from a summary of a month or a province carry a plausible date and coordinate that were assigned rather than observed, and clustering analysis run without filtering on event_clarity will discover structure the coders created.
- Place-name geocoding puts violence in the wrong place. An event attributed to a town because the source named the district appears as a precise point, and only where_prec distinguishes it from an event someone actually located.
- Source duplication inflates spikes. A single incident reported by several agencies is meant to be coded once, but where reports differ enough in place, date or casualty figure, near-duplicates survive, and they concentrate in the heavily reported incidents that drive headlines.
- Actor labels imply more than they mean. A government identifier persists across regime change and an umbrella rebel label can cover factions that fought each other, so a dyad time series can look continuous across a real-world discontinuity.
- Events from inactive years are read as evidence of an ongoing conflict. GED includes events from dyads in years when they did not meet the annual threshold, and active_year is the only flag distinguishing them.
- Candidate data is cited as if final. The monthly series is structurally identical to the annual release and its events are routinely revised, merged or removed; a report built on candidate data without saying so will not reproduce.
- Country attribution follows the coordinate. Cross-border incidents, disputed territory and contested administrative boundaries produce country assignments reflecting UCDP's geographic conventions rather than any party's claim, which matters much more when the dispute is itself the subject.
None of these make the source unusable. They make it a source that requires corroboration before an assertion built on it goes into a product, which is true of every source and admitted by few.
Ageing
Individual events do not age; they are dated observations of the past and remain valid as such. Two other things age fast. The first is the current-period picture: anything drawn from the candidate series is provisional by construction and should be re-checked against the annual release, and a dashboard of this year's violence built on candidate data is showing a draft. The second is the coding itself. UCDP revises history at each release as new sources emerge and rules are refined, so a stale record here looks like an event that no longer exists in the current version, or one whose fatality estimate has moved materially, or a dyad that has been renumbered or merged. The defensive practice is to store the version string with every extract and diff releases rather than overwriting. An analysis that cannot say which version it used cannot be reproduced and, in any adversarial setting, cannot be defended.
What this source feeds
A source is only worth what it lets you conclude. These are the disciplines that collect through it, the mission domains it serves and the data points it yields — every one is a tag, so you can follow any thread from here into the rest of the library.
Collected by these intelligence disciplines
Serves these mission domains
Yields these data points
How each sector uses UCDP Georeferenced Events
The same dataset is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The records are shared — the constraints, thresholds and outputs are not.
🎖 Military and defence
For campaign assessment and operational design, GED is the reference series for organised violence intensity in an area of operations, and its principal military value is comparability: the same coding rules across thirty-five years and every theatre mean a change in the series is more likely to be a change in the world than a change in the instrument. Use it to establish pre-deployment baselines, to test whether an operation was followed by a measurable change in lethal activity, and to characterise the actor landscape through dyad structure rather than through a list of group names. Its limits are equally operational: it lags, it excludes non-lethal engagements entirely, and it carries no order-of-battle content, so it will not support targeting or tactical force protection and should not be asked to.
🕵 National intelligence
The dyad and actor identifiers are the underused asset. Because UCDP maintains stable identifiers for belligerents across decades, GED supports longitudinal actor analysis that name-matching cannot: tracing a group through renaming, splintering and merger, or establishing when a government first fought a given adversary. For estimative work, the low and high fatality bounds are a ready-made expression of uncertainty that maps cleanly onto probabilistic language in a finished product. The discipline point is that GED is a lagging, conservative indicator and therefore the wrong instrument for warning; it is the right instrument for validating whether a warning judgement made two years ago was correct, a task most shops perform badly and this dataset makes possible.
👮 Law enforcement
Relevant chiefly where organised violence intersects transnational crime, and where a court or an inquiry needs a defensible general picture of conditions in a place and period. The audit trail matters here more than anywhere: every event names the reporting it was coded from, so an assertion about the security situation in a district in a given month can be traced to source rather than asserted. It is not evidence of any individual act and must never be presented as such. Its practical uses are contextual: establishing that a region was in a state of armed conflict for the purposes of a legal test, corroborating a claimant's account of conditions, or bounding a period of instability relevant to an asset-tracing or sanctions matter.
🔍 Private investigation and corporate security
For corporate security, due diligence and country-risk work, GED is the cheapest credible answer to whether a location has a history of organised violence and, more usefully, what kind. The three violence types are the analytically important split: a district with state-based conflict, one with non-state communal fighting and one with one-sided violence against civilians present very different risks to a workforce and a supply chain, and clients routinely conflate them. Pair the event history with the precision codes so a map handed to a client does not imply precision the data lacks, and be explicit that an absence of recorded events in a poorly reported region is not a clean bill of health.
📰 Journalism and OSINT media
GED is the series to reach for when you need a number that will withstand a challenge, and the low and high bounds are a gift to honest reporting rather than an inconvenience. It supports the durable stories: that a conflict everyone stopped covering has become more lethal, that a declared ceasefire did or did not change anything measurable, that violence in a country is concentrated in districts nobody names. Two obligations come with it. State the version and the fatality estimate used, because a different choice yields a different headline. And be clear that the dataset counts direct deaths from organised violence only, so a comparison against any total war-mortality figure is a category error that will be pointed out.
🌍 NGO, humanitarian and human rights
For humanitarian analysis, protection monitoring and advocacy, the one-sided violence category is the closest thing to a systematic, comparably coded global record of organised violence against civilians, and the civilian death disaggregation supports protection reporting directly. The geographic precision codes let you decide honestly whether an event can be attributed to a district for programming purposes. The critical caution is the exclusion of indirect mortality: for populations whose deaths come from displacement, hunger and collapsed health systems, GED will show a modest and possibly declining series while the humanitarian situation deteriorates. Present it alongside displacement and mortality data, never as a proxy for suffering, and never let an absence of events be read as an absence of need.
🎓 University and research
GED is the standard dependent variable in quantitative conflict research and the reason a large literature is comparable at all. The permanent versioning of API endpoints is an unusually strong reproducibility feature and should be used: embed the versioned query in the replication package rather than shipping a CSV. The methodological obligations are well known and still routinely ignored. Model the reporting process rather than assuming events are observed at random; use where_prec and date_prec as filters or as measurement-error terms rather than discarding them; do not use best as a point estimate without acknowledging the bounds; and be explicit about whether the absence of a dyad reflects peace or a threshold. The codebook for your exact version is the primary methodological document and it is short enough to read in full.
Playbook: working UCDP Georeferenced Events end to end
A repeatable sequence from first pull to finished product. Each phase states what you are trying to establish, not merely what to click — the objective is a defensible chain of reasoning, not a completed checklist.
Phase 1 — Read the codebook for your version before writing a line of code
The codebook is the specification and it is not long. Definitions of the three violence types, the precision codes and the fatality estimates determine every subsequent decision, and most misuse of GED traces to an analyst who inferred the semantics from column names. Keep the codebook in the project directory alongside the data.
Phase 2 — Decide which violence universes you are in
State-based, non-state and one-sided violence are three separate definitional systems that happen to share a table. Merging them into one count is defensible only if you say so and can justify it; more often you want them separated, because they answer different questions and carry different reporting biases. Make the decision explicitly and record it.
Phase 3 — Establish the reporting environment before interpreting counts
For each country and period in scope, form a view on press access and coverage. A drop in events in a country where journalists were expelled is an artefact; the same drop where reporting was unchanged is a finding. Without this step every trend statement is uninterpretable, and it is the step analysts skip most often.
Phase 4 — Filter on precision before you map anything
Decide the minimum where_prec you will accept for spatial claims and the minimum date_prec for temporal claims, and apply them visibly. If the analysis needs district-level attribution, low-precision events must be excluded or handled separately rather than plotted as points and hoped for.
Phase 5 — Separate observed incidents from constructed ones
Use event_clarity to split incidents reported as discrete from those coders disaggregated out of summary reporting. Run any clustering, sequencing or pattern analysis on the discrete set first, then test whether adding the constructed events changes the conclusion. If it does, the conclusion is about the coding process.
Phase 6 — Work with the bounds, not the point estimate
Carry low and high through the whole analysis. Where the bounds are tight, best is meaningful; where they span a factor of five it is a placeholder. Reporting a total alongside the sum of lows and the sum of highs is more honest and, in practice, more persuasive than a single figure with a footnote.
Phase 7 — Reconcile GED against the yearly datasets
Aggregate your extract to conflict-years and compare with the battle-related deaths and conflict-year series. They come from the same coding effort and are expected to differ in defined ways. An unexpected gap almost always means your extract includes inactive years, mixes violence types, or has silently dropped events.
Phase 8 — Build actor histories from identifiers, not names
Use the side and dyad identifiers to assemble an actor's full history across renamings and factional changes, then check the dyad structure to see who that actor has actually fought. Name-based matching against other datasets will fragment one group across a dozen spellings and merge distinct groups that share a word.
Phase 9 — Cross-read against a faster, broader feed
Run the same country and period through ACLED or a media-event feed and examine the difference set rather than the intersection. Events present elsewhere and absent here are usually below the fatality threshold, involve an unregistered actor, or belong to a different category of violence, and identifying which teaches you what your extract is silently excluding.
Phase 10 — Audit a sample back to source
Take a dozen events that matter to your conclusion and read the reporting listed in source_article. This is the step separating a defensible finding from a plausible one, and it regularly reveals that the incident you are building on is a single wire report of a claim by one party.
Phase 11 — Version-lock and diff
Pin the version in your extract, store it with the results, and when a new release lands, diff rather than overwrite. Report material revisions to whoever consumed the earlier analysis. UCDP revising history is a sign of a healthy dataset, but it silently invalidates work that did not record which history it used.
Phase 12 — Write the exclusions into the product
Whatever you publish should state, in the body and not a footnote, that the series counts direct deaths from organised violence meeting a defined threshold, excludes indirect mortality and non-lethal violence, and is bounded by press access. Without that, a reader will treat your figure as the human cost of the war, and the error will be attributed to you rather than to the dataset.
The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.
What to pair it with
No single source carries a finding. These are the datasets that corroborate, extend or contradict this one — and a source that contradicts is worth more than one that agrees, because it is the only thing that will tell you when you are wrong.
| Source | Relationship | What it adds |
|---|---|---|
| ACLED | extends | Broader inclusion rules, faster cadence, and coverage of protests, riots and non-lethal events. The natural counterpart: where the two disagree, the disagreement usually identifies a definitional boundary rather than an error. |
| UCDP/PRIO Armed Conflict Dataset | prerequisite | The conflict-year backbone defining which conflicts exist and when they were active. Read it to understand why a dyad appears in GED at all. |
| PRIO-GRID | extends | Standardised global grid with population, terrain, infrastructure, nightlight and ethnic settlement covariates, joined directly by the priogrid_gid already present in every GED row. |
| GDELT | contradicts | Machine-coded global media events at enormous volume and much lower precision. Useful as an early signal and as a demonstration of what automated coding does to a conflict series; never as a substitute for GED counts. |
| UNHCR population statistics | corroborates | Displacement responds to violence with a lag and frequently reveals conflict intensity in periods when event reporting is suppressed. |
| Internal Displacement Monitoring Centre | corroborates | Conflict-driven internal displacement figures at country and increasingly subnational level, which often move when the event series does not. |
| EM-DAT | extends | Disaster and complex-emergency mortality. Helps place conflict deaths in context alongside the other things killing people in the same country and year. |
| UCDP Encyclopedia | prerequisite | Narrative descriptions of conflicts, actors and dyads written by the same programme. The fastest way to find out what a numeric dyad identifier actually refers to. |
Legal, ethical and operational constraints
There are no meaningful statutory restrictions on using published, aggregated conflict event data of this kind, and the source is designed for open research use. The constraints that do apply are ethical and reputational. First, the underlying reporting frequently concerns identifiable victims and communities; although GED itself carries no personal data, republishing precise coordinates of one-sided violence against a small community in an ongoing conflict can direct attention to survivors, so consider whether your output needs point-level publication or whether aggregation serves the same purpose. Second, characterising a party as a belligerent in an armed conflict is a legally loaded statement in some jurisdictions and can carry defamation exposure when applied to a named organisation that disputes the label; UCDP's actor coding is a research judgement, not a legal determination, and should be presented as such. Third, where output feeds targeting, sanctions designation or immigration adjudication, the dataset's exclusion rules become material to fairness, and using it as evidence that violence did not occur in a place is not defensible in any of those settings.
Operational security
Bulk downloads reveal nothing beyond an HTTP fetch of a public file, and this is one of the few sources in the catalogue where the low-exposure option is also the analytically better one. The API is different: a token is issued against a named person and a stated purpose, and queries associate that identity with specific countries, bounding boxes, dyads and date ranges. For most users that is unremarkable and the relationship is a benign academic one. If your interest in a particular actor or border region is itself sensitive, take the full download and filter locally rather than issuing narrow queries, and do not request a token in the name of an organisation whose interest you are trying not to advertise. The web interface carries ordinary web analytics. There is no facility for confidential querying here and none is claimed.
Two rules that hold regardless of jurisdiction. Collection that is lawful is not automatically proportionate, and a dataset assembled for one purpose does not carry consent for another. Where the records concern identifiable people, the question is not only whether you may hold the data but whether holding it serves the purpose you are accountable for.
Is it earning its place?
Sources accumulate. Feeds get added during an incident and are never reviewed again, and a decade later the pipeline is carrying dead weight that nobody dares remove. These are the measures that show whether UCDP Georeferenced Events is contributing anything, and they are worth baselining now so the answer is available later.
- Share of your extract with the highest two where_prec values, which tells you what proportion of your map is actually located rather than approximated.
- Share of events with day-level date precision, which bounds how much sequencing or day-level correlation the extract can support.
- Ratio of the summed high estimate to the summed low estimate across the extract, as a single number expressing how uncertain the casualty picture is.
- Events per country-year against an independent measure of reporting volume, tracked over time, to catch coverage changes masquerading as trends.
- Count of events added, removed and materially revised at each version release for the countries you cover, which directly measures how much your historical baseline moves.
- Proportion of GED events in your area of interest that also appear in a broader feed, and the composition of the difference set, as a running check on what the definitional filter removes.
- Number of finished products in which a GED figure was published without the version, violence types and estimate choice stated. This should be zero and rarely is.
Beware of volume. Indicator counts rise easily and say almost nothing. Unique contribution — findings this source produced that no other source in your stack would have — is the measure that matters, and it is usually far lower than anyone expects.
Tradecraft notes
The distinctions that separate a competent analyst from a fast one:
- The codebook is the dataset. Column names are shorthand for definitions carrying real analytical weight, and an analyst who has not read the definitions of the three violence types will produce counts that mean something other than what they think.
- Never sum best without carrying low and high. The bounds are the dataset's own statement of what it does not know, and discarding them converts honest uncertainty into false precision that survives every subsequent step.
- Absence is about reporting, not about peace. Before any statement that violence declined, establish that the capacity to observe violence did not decline first, and say which one you checked.
- Where_prec and event_clarity are analytical fields, not housekeeping. Dropping them at ingest is the single most common way a competent GED analysis becomes a misleading map.
- Use identifiers for actors and dyads and treat names as display labels. Any join to another dataset on an actor name will fragment and merge groups in ways you will not notice until someone checks.
- Reconcile with the yearly datasets before publishing an aggregate. They come from the same coding operation, and a divergence you cannot explain is a bug in your pipeline until proven otherwise.
- Candidate data is a draft. It is valuable and it is provisional, and any product using it must say so in the same sentence as the number, not in a methods note nobody reads.
- GED is a lagging indicator by design. If your requirement is warning, this is the wrong instrument, and using it as one produces the characteristic failure of noticing an escalation eighteen months after it began.
- Read the source articles for events that carry weight. The audit trail is the reason to prefer this dataset, and an analyst who never uses it has given up the advantage that justified the choice.
Questions analysts actually ask
How does GED differ from ACLED, and which should I use?
GED admits only organised violence with at least one direct death, coded against a fixed definitional threshold since 1989; ACLED admits a much wider range of political violence and unrest including non-lethal events, and updates far faster. Use GED when the number must survive review or when you need a stable multi-decade series, and ACLED when you need breadth, currency or non-lethal activity. Running both and examining the difference set is more informative than choosing.
Why does an event appear for a year when the conflict was supposedly inactive?
GED includes events from dyads that met UCDP's activity threshold in some year, including events from years when they did not. The active_year field distinguishes them. This is deliberate, and it is why GED event counts do not reconcile naively with the conflict-year datasets.
Can I use best as the number of people killed?
Only with the bounds attached. Best is a coder's judgement of the most plausible figure between low and high, and where reporting is poor those bounds can differ by a factor of several. Report the range, or state clearly that you are using the best estimate and what its bounds were.
How current is the data?
The definitive annual release lags the year it covers by several months to a year. GED Candidate fills the gap with monthly and quarterly preliminary coding under the same rules, but candidate events are revised and sometimes removed at the annual release. Do not build a real-time product on this source.
Do I need an API token, and how do I get one?
Yes, for the API. Tokens are free and issued on request by email to the maintainer named in the API documentation, with a short description of your project or intended use; the documentation states requests are answered within three to five working days. The token travels in an x-ucdp-access-token header, so browser address-bar testing will not work. The bulk downloads need no token at all.
Are civilian deaths from airstrikes, sieges and famine included?
Deaths from an airstrike that killed someone directly are included and appear in the civilian disaggregation. Deaths from a siege that killed by starvation, from displacement, or from collapsed health services are excluded entirely. This is the most consequential exclusion in the dataset and the one most often overlooked when GED totals are quoted as the cost of a war.
Why do coordinates look suspiciously round or land on a city centre?
Because the coordinate is a geocode of the place name in the source reporting, not an observation of where the violence happened. When a report gives only the district, coders assign the district's representative point. The where_prec field records exactly how much to trust it and should be checked before anything is plotted.
Does GED cover criminal and cartel violence?
Only where the parties satisfy UCDP's organised-actor test and the violence falls into one of the three defined categories. A great deal of lethal criminal violence, including much homicide driven by organised groups, sits outside the universe. If organised crime is your subject, GED is a partial and definitionally awkward source and you should say so rather than quietly using it as a proxy.
How should I cite it?
Cite the specific dataset and version string, and the Journal of Peace Research articles UCDP names as the citations of record for GED and the associated conflict datasets. If you used the API, the versioned URL is itself a durable citation, which is one of the more useful properties of this source and worth exploiting in replication material.
Standards, formats and interoperability
What this source speaks natively, and what it has to be translated into before a partner can consume it. Work that arrives in a recognised format is easier to defend, easier to hand over and easier to automate against:
- UCDP's own definitions of state-based, non-state and one-sided violence, and of armed conflict at the twenty-five battle-related deaths per year threshold, are the semantics behind every field in the dataset.
- Gleditsch and Ward country numbers are used for country identity rather than ISO codes, which matters for joins and for how historical state changes are handled.
- PRIO-GRID cell identifiers are precomputed in every event row, giving direct compatibility with the standard spatial covariate infrastructure used in quantitative conflict research.
- The API's permanent version addressing functions as a citation standard: a versioned endpoint returns identical data indefinitely, supporting reproducible research in a way few data services do.
- Coordinates are supplied as decimal degrees and additionally as WKT point geometry, so no projection assumptions are required.
- The precision coding scheme for place and date is documented measurement uncertainty and maps cleanly onto the confidence expressions used in structured analytic products.
- Derived indicators and entities exported from the platform travel in STIX 2.1, MISP, CSV, JSON and JSONL, so conflict events can be carried into the same case structures as any other observable.
References
Primary documentation and authoritative references for this source. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.
- Uppsala Conflict Data Program — Department of Peace and Conflict Research, Uppsala University. The programme's front door, with the interactive interface and links to every dataset and version. Start here to see what exists before deciding what to download.
- UCDP API documentation — Uppsala Conflict Data Program. Authoritative statement of endpoints, filters, paging, the token requirement and the daily quota. Read it rather than any secondhand endpoint list, including older ones written before the token existed.
- UCDP Dataset Download Center — Uppsala Conflict Data Program. Bulk downloads, codebooks and version histories for GED and every companion dataset. The right route for anything resembling a full extract.
- UCDP GED codebook — Uppsala Conflict Data Program. The specification: definitions of the violence types, the precision codes, the fatality estimates and the actor structure, in about thirty pages.
- UCDP definitions — Uppsala Conflict Data Program. The concise statement of what counts as organised violence, an armed conflict and an organised actor. The single most important page for anyone interpreting a GED count.
- UCDP Encyclopedia — Uppsala Conflict Data Program. Narrative context for conflicts, dyads and actors, and the way to find out what a numeric dyad identifier actually refers to.
- Department of Peace and Conflict Research — Uppsala University. Institutional home of the programme. Useful for understanding the funding and staffing model behind the coding effort and for the associated publications.
- Peace Research Institute Oslo — PRIO. Co-producer of the UCDP/PRIO Armed Conflict Dataset and publisher of much of the methodological literature on conflict event data.
- PRIO-GRID — PRIO. The grid system referenced by every GED row, with the covariate layers that turn an event point into a modelled observation.
- ACLED — Armed Conflict Location and Event Data Project. The main comparator. Reading its inclusion criteria alongside UCDP's is the fastest way to understand what each dataset chooses not to see.
Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.
Put it into practice
The Quantus Intel threat intelligence platform operationalises this source: it registers each annual version as a distinct source, carries the precision and clarity codes through to the analyst's screen, joins events to actors and grid cells rather than to place names, and diffs every new release against the last so a revised history does not quietly invalidate a finished assessment.. Browse the full source catalogue, or follow any tag above into the rest of the library.