August 27, 2026

Global Terrorism Database: Intelligence Source Guide

0

The GTD is the longest-running open event dataset on terrorist attacks, coded incident by incident from public reporting since 1970. For kidnap and hostage work it carries fields almost nothing else does: hostage counts, duration, ransom demanded, ransom paid, and how it ended.

global-terrorism-database-intelligence-source-guide

The GTD is the longest-running open event dataset on terrorist attacks, coded incident by incident from public reporting since 1970. For kidnap and hostage work it carries fields almost nothing else does: hostage counts, duration, ransom demanded, ransom paid, and how it ended.

At a glance

Source Global Terrorism Database
Category Conflict, Crime & Human Security › Organised Crime, Gangs & Piracy
Homepage https://www.start.umd.edu/gtd/
Machine interface https://www.start.umd.edu/gtd/
Format HTML
Access Open — no account required
Disciplines Open Source Intelligence, News Intelligence
Mission domains Kidnap, Hostage & Extortion, Extremism & Radicalization

Kidnapping/hostage incident dataset (START). — as catalogued in the platform’s own source registry.

The Global Terrorism Database is an incident-level event dataset maintained by START, the National Consortium for the Study of Terrorism and Responses to Terrorism, at the University of Maryland. Each row is one attack, identified by a composite event ID built from the date plus a sequence number, and coded against a fixed schema of well over a hundred variables. The coding is done by human analysts reading open-source reporting – wire services, national and local press, and later a machine-assisted pipeline that surfaces candidate articles for a coder to adjudicate. The variables cover the basics you would expect (date, country, first-order administrative division, city, latitude and longitude, attack type, weapon type, target type and subtype, perpetrator group name, killed, wounded) and a long tail you would not: whether the perpetrator group claimed the attack or was merely attributed, how confident the coder was in that attribution, whether the incident was part of a coordinated multi-attack, whether property damage occurred and at what order of magnitude, and an entire hostage and kidnapping block. Crucially, every record carries the coder's judgement on three inclusion criteria and a doubt flag, so the dataset ships with its own uncertainty attached rather than pretending to a clean count.

Its analytical job is longitudinal comparison. Almost every other terrorism data product is either a current-awareness feed, a regional specialism, or a derived index; the GTD is the only widely available series that lets you place a 1978 hostage-taking, a 1994 bombing and a 2018 armed assault in the same coding frame and ask whether something changed. For the kidnap and extortion mission specifically, it is the base rate. When a client asks whether kidnapping of foreign nationals in a given province is rising, or whether a group that has historically released hostages has started killing them, or what the historical relationship is between ransom demanded and ransom paid in a theatre, the GTD hostage block is the only open dataset that will answer at scale. It is also the substrate under a great deal of published quantitative terrorism research and under the widely cited annual terrorism index products, which means that if you are reading a study or an index score, you are frequently reading the GTD at second hand – and inheriting its coding decisions whether or not you know it.

Who publishes it, and why that matters

START is a university research centre, originally established as a US Department of Homeland Security Center of Excellence, and the GTD has been funded through a rotating series of government research grants and contracts rather than a stable subscription base. That funding model has consequences you must plan around. Collection has been interrupted before, the update cadence has been irregular, and the end date of the publicly released series has not always moved forward on an annual rhythm. Do not build an operational dependency on the next release arriving. The lineage is also not a single continuous effort: the earliest years derive from a commercial risk-consultancy collection effort running from the 1970s into the 1990s, later years were coded by two successive academic centres under different protocols, and START took collection fully in-house at the start of the 2010s with a substantially different, higher-recall methodology. The team has always been open about this in the codebook, which is to their credit and is also the single most important thing a new user fails to read. Treat START as a careful, transparent, chronically under-resourced academic operation, and treat any date-based trend that crosses a methodology boundary as suspect until you have checked which boundary it crosses.

Provenance is the first question to ask of any dataset and the one most often skipped. Who collects it, what their incentive is, whether they publish a methodology, and whether they correct the record when they get something wrong all bear directly on how much weight a finding drawn from it can carry.

What a record actually contains

The fields you will be working with, what each one means, and whether it is something you can pivot on. Read the meanings carefully — more analysis is wrecked by misreading a field than by failing to find one, and a field that looks like an observation is often an inference.

Field Type What it means Pivot value
eventid string A twelve-digit identifier: eight digits of date plus a four-digit sequence for that day. It is the stable join key and it encodes the coded date, so a record whose eventid prefix disagrees with its date fields has been edited. Join to derived datasets and published research replication files; cite it in reports so a reader can check your row.
iyear, imonth, iday int Year, month and day of the incident as three separate integers. Month or day may be zero where reporting only gave a month or a year, and a companion approximate-date text field carries the ambiguity in words. Timeline reconstruction; a zero day is a signal to widen any temporal correlation window rather than to discard the row.
country_txt, region_txt string Country and a coarse world region from the GTD's own regional scheme. Country is coded to the political geography at the time of the event, so Soviet-era and pre-partition records will not match a modern country list. Country dashboards and theatre views – but reconcile the historical country coding to a current ISO list first.
provstate, city string First-order administrative division and the named city or locality. Free text as reported, with real inconsistency in transliteration and in whether a district, a village or a nearby town was named. Administrative geocoding, gazetteer matching, and local press searches; expect to normalise before you aggregate.
latitude, longitude, specificity float Coordinates plus a specificity code stating what the coordinates actually refer to – the exact site, the nearest named city, a larger administrative unit, or a whole country. Specificity is the field that stops you drawing a falsely precise map. Mapping, distance-to-asset calculation and hot-spot analysis, but only at the resolution specificity permits.
crit1, crit2, crit3, doubtterr enum Three inclusion criteria – a political, economic, religious or social goal; intent to coerce or intimidate an audience wider than the immediate victims; occurrence outside legitimate warfare activity – plus a flag for whether the coder doubted the incident was terrorism at all. Filtering on these changes your counts materially. Definition-sensitivity testing; run any headline number under at least two criteria filters before you publish it.
attacktype1_txt enum The tactic: bombing or explosion, armed assault, assassination, hostage taking by kidnapping, hostage taking in a barricade situation, hijacking, facility or infrastructure attack, unarmed assault, or unknown. Up to three attack types can be coded per event. Tactic-over-time series, group tactical repertoire profiling, and the primary filter for kidnap work.
targtype1_txt, targsubtype1_txt, corp, target1 string Target category and subcategory, the name of the corporate or institutional entity targeted, and the specific target – a named person, a named vessel, a named building. The free-text target fields are the richest and least exploited part of the schema. Entity extraction to persons and organisations; sector risk profiling; matching a client asset list against historical targeting.
gname, guncertain1, claimed string Perpetrator group name as coded, a flag for whether attribution is uncertain, and a separate flag for whether the group claimed responsibility. Attribution and claim are different questions and the dataset deliberately keeps them apart. Actor profiles and cross-reference to designation lists; never treat the group name as confirmed attribution without reading both flags.
nkill, nwound, nkillter, nwoundte int Fatalities and injuries, with separate counts for perpetrators where reporting distinguished them. Blank means unknown, not zero, and conflating the two is the most common quantitative error made with this dataset. Severity weighting and lethality-per-attack series, and comparison against conflict-fatality datasets that count differently.
ishostkid, nhostkid, nhours, ndays int Whether hostages or kidnap victims were taken, how many, and the duration of the incident in hours or days. Sentinel values are used for unknown in some releases, which will silently poison a mean if you do not screen them. Kidnap duration distributions by group and geography; input to hostage-recovery and crisis-response planning assumptions.
ransom, ransomamt, ransompaid int Whether a ransom was demanded, the amount demanded in US dollars, and the amount reportedly paid. Both amounts are as reported in open sources, which means they are the numbers someone chose to say out loud. Extortion-economy analysis; corroborate against prosecution files and insurer data before quoting any figure.
hostkidoutcome_txt, nreleased enum How the hostage incident ended – released by the perpetrators, escaped, killed, rescued, attempted rescue, or a combination – and how many people were released. This is the outcome variable for almost all kidnap base-rate work. Survival and release-rate analysis by group, region and victim nationality; scenario weighting for response planning.
dbsource, scite1, scite2, scite3 string Which collection effort coded the row and up to three source citations for it. The citations are how you audit a record, and dbsource is how you detect that an apparent trend break is really a methodology break. Return to the underlying press report; segment any time series by dbsource to expose collection discontinuities.

Coverage — and what is not in it

Global in scope, from 1970 onward, at incident level. Every inhabited region appears, though density follows press coverage rather than violence: South Asia, the Middle East and North Africa, and sub-Saharan Africa dominate the later decades, while Western Europe and Latin America dominate the 1970s and 1980s. The single most important coverage fact is the 1993 gap – the records for that year were physically lost before digitisation and were never fully reconstructed at incident level, so any series spanning 1993 has a hole in it that is an artefact and not a lull. The second is that coverage depth changes with the collecting organisation: the volume of coded events rises sharply at the start of the 2010s when START moved to an in-house, higher-recall pipeline, and a naive year-on-year chart across that boundary will show a surge in terrorism that is mostly a surge in collection. The end of the series is not fixed. Releases have been periodic and funding-dependent, and the last year of data in the copy you hold may be several years behind the present. Check the release note on the site for the actual coverage window before you scope any piece of work, and state that window in anything you publish. This is a historical baseline source, not a current-awareness feed, and using it as the latter is the fastest way to be wrong in front of a client.

Known blind spots

Absence of evidence here is not evidence of absence. These are the conditions under which Global Terrorism Database will not show you something that is nevertheless real:

  • The series ends where funding ended. There is no rolling window and no guarantee of a next release, so the most recent eighteen to thirty-six months of any theatre may be missing entirely – exactly the period an operational customer cares most about.
  • 1993 is absent as a matter of physical record loss, not of low activity, and every chart drawn straight through it is misleading. Partial country-level reconstructions exist but they are not incident-level and cannot be joined on the event identifier.
  • State violence is out of scope by definition. Attacks by state military and police forces acting as such are not coded, so in theatres where the dominant threat to civilians is the government you will see a fraction of the violence and a distorted picture of who is dangerous.
  • Incidents inside conventional armed conflict are systematically thinned by the third inclusion criterion. In an active war the line between an insurgent attack and a battlefield engagement is a coder's judgement call, and the same event may be present, absent or doubt-flagged depending on how it was reported.
  • Press-poor geographies under-report at every level. Rural areas without a local newspaper, countries with censored media, and periods of communications blackout generate few or no records regardless of what happened on the ground.
  • Kidnapping for ransom is under-counted structurally, because the families and companies with the strongest incentive to keep an abduction out of the press are precisely those most able to resolve it quietly. The recorded ransom fields skew towards cases that went public, which usually means the ones that went badly.
  • Group attribution is only as good as the reporting behind it. A large share of records carry an unknown perpetrator, and where a group is named it is frequently named by a government spokesman with an interest in the naming rather than by the group itself.
  • Non-English reporting is unevenly captured. Coverage of local-language press has varied by collecting organisation and by period, which means comparability across regions is worse than comparability within one region over time.
  • Nothing here is forward-looking. The dataset records what happened, not capability, intent, current group strength, or the existence of an unexecuted plot, and it will never show you the attack that was disrupted.

Write the blind spot into the product. A statement that something “was not observed in Global Terrorism Database” is defensible; a statement that it “did not happen” is not, and the difference is what survives cross-examination.

Access, licensing and what you may do with it

Access model: Open — no account required

Access is through the START website. The dataset is distributed as a downloadable file for registered users who accept the terms of use, together with a codebook that is the authoritative definition of every variable and every code value. There is also a browser-based search interface for looking up individual incidents without downloading the whole file, which is the right tool for verifying one record inside a case rather than for analysis at scale. Registration is straightforward but is not anonymous, and the terms you accept at download time are the terms that bind you afterwards. Practical advice: download once, keep the file, its codebook and the release note together in your evidence store, and record the release version in the case file. Because the published series has been revised as well as extended – records get recoded, coordinates corrected, group names harmonised – a figure you quoted from an earlier release may not reproduce against a later one. If your report will be scrutinised, archive the exact file you analysed rather than a pointer to the download page.

Licence

The GTD is distributed under terms of use that you accept at download and that have been oriented towards non-commercial research, with attribution required and redistribution of the dataset restricted. Commercial exploitation and bulk republication of the data have historically required separate permission from START. The terms have been revised over the life of the project, so read the current statement on the site rather than relying on what a colleague told you or on what a paper published five years ago said. In practice the safe posture for a commercial intelligence provider is: derive statistics and analysis from the data, cite START and the specific release in every product, do not ship the underlying rows to a client, and seek written permission before anything that looks like resale. If you work in a law-enforcement or government context, check whether your organisation already holds an institutional use agreement before you download under personal terms that may be narrower.

Rate limits and fair use

There is no API to rate-limit. This is a bulk download plus a web search interface, and the etiquette is correspondingly simple: take the file once, do not scrape the search interface as a substitute for downloading it, and do not re-download on a schedule in the hope that the data changed. Releases are infrequent enough that a quarterly manual check of the release note is more than sufficient. If you need automated awareness of a new release, poll the download page at a low frequency with a descriptive user agent and an owning contact address, and back off immediately on any error.

Licensing changes, and it changes without warning. A dataset that was free for research this year may not be free for commercial or evidential use next year. Confirm the current terms before you build a dependency on it, and record the terms you relied on alongside the data — the licence in force at the time of collection is part of the provenance.

Collecting it

How Global Terrorism Database is actually pulled, in the order you would set it up. Prefer the bulk or export interface over per-item lookups wherever one exists: it is kinder to the publisher, faster for you, and gives a reproducible snapshot rather than a series of point-in-time answers you cannot reconstruct later.

Method Format Cadence Notes
Full dataset download bulk per release, historically irregular The primary route. One file, one codebook, one release note. Store all three together; the codebook is version-specific and a code value can change meaning between releases.
Web incident search HTML as needed For verifying or citing a single incident inside a case. Use it to confirm a record you already found in the bulk file, not to assemble a dataset one query at a time.
Codebook HTML per release Read it before the data, not after. Every analytical error catalogued in this guide is documented somewhere in it, including the inclusion criteria, the unknown-value conventions and the collection discontinuities.
Derived indices and replication files bulk annual or per publication Published research and annual terrorism indices are built on GTD rows. Useful for benchmarking your own aggregation logic against someone else's and for finding out why your numbers diverge.

Ingesting it into the platform

Every step below is idempotent and cursor-based: interrupt one and it resumes from where it stopped rather than duplicating rows or losing progress. Collection is recorded per source, so a feed that quietly stops publishing shows up as a stale timestamp instead of silently thinning your coverage.

  1. Register the release — Record the dataset in sources.php with the exact release identifier and coverage end date, so every downstream record inherits a provenance stamp that says which version of the GTD it came from.
  2. Load through import.php — Bulk-load the file rather than trickling it through the feed collector. This is a static historical corpus, not a stream, and treating it as a feed produces a meaningless last-collected timestamp on sources.php.
  3. Normalise unknowns at the door — Convert unknown markers and sentinel values in the count, duration and ransom fields to true nulls at ingest. If you do not do this once, someone will compute a mean over them at three in the morning and put it in a report.
  4. Split identity from attribution — Map the perpetrator name to an actor entity but carry the uncertainty and claim flags onto the relationship rather than the node. In actor-profile.php an attributed attack and a claimed attack must remain visually distinguishable.
  5. Geocode within specificity — Store coordinates alongside the specificity code and refuse to render a point at site resolution when specificity says the coordinates are a country centroid. Enforce this at ingest, because it will not be enforced by the analyst reading the map.
  6. Emit event and location data points — Each row yields a dp_event with typed attributes and a dp_location constrained by specificity. Kidnap records additionally yield duration, demand, payment and outcome attributes that feed the kidnap and extortion views.
  7. Correlate against live event feeds — Run correlate.php against current-awareness sources to establish where the GTD baseline ends and live reporting begins. The join is by place, date and actor, it will not be clean, and matches should be treated as candidates for analyst review rather than as merges.
  8. Publish as baseline, not as current — Surface it in timeline.php and country-risk.php with an explicit series end date rendered on the chart itself. Nothing in the platform's kidnap views should imply that the GTD knows anything about last month.

Registered sources and their last-collected state are listed in sources.php, and the scheduled chain that keeps them current is in automation.php.

How it is wrong, and how to tell

Every dataset is wrong in characteristic ways. Knowing which ways is the difference between using a source and being used by one, and it is the part of source evaluation most often skipped because it is the part that takes work.

Judged as a research dataset, it is one of the best-documented event collections in the social sciences: the coding rules are published, the uncertainty is coded rather than hidden, the source citations are retained, and the known defects are stated by the people who built it. Judged as an operational intelligence feed, it is weak in exactly the ways a research dataset is weak – latency measured in years, a fixed schema that cannot absorb a new tactic quickly, and a dependence on press reporting that inherits every distortion in the press. The correct basis for judging it is reproducibility. Take a set of incidents you already know from primary sources – a court file, an after-action report, a company's own kidnap history – and look them up. You will typically find that the events are present, the geography is approximately right, the casualty counts are conservative, the group attribution follows the dominant press narrative of the time, and the ransom figures are the ones that were reported rather than the ones that were paid. That profile is fine for base rates and dangerous for case-specific claims. Use it to say what is normal, and use primary material to say what happened.

Characteristic false positives

  • Blank read as zero. Unknown fatality, hostage and ransom values are empty or sentinel-coded, and software that treats them as zeroes will produce lethality and payment averages that are confidently and precisely wrong.
  • Collection discontinuity read as a real trend. The step change in event volume at the start of the 2010s, and the smaller shifts at each earlier change of collecting organisation, look exactly like escalation on a chart and are not.
  • Coordinates read at face value. A record with country-level specificity plots as a point in the middle of a country, and analysts routinely build density maps whose peak sits in an empty desert because that is where the centroid fell.
  • Attribution read as fact. The perpetrator name is what open sources said, filtered through a coder. In contested environments governments name convenient groups and rival groups claim each other's work, and the dataset records the claim rather than the truth.
  • Duplicate and split events. A coordinated multi-site attack may be one row or several depending on the coding rules in force, so both attack counts and casualty totals across a complex incident can double-count or under-count.
  • Definitional filtering left off. Headline totals differ substantially depending on whether you require all three inclusion criteria and whether you exclude doubt-flagged records, and two analysts quoting the GTD can disagree by a wide margin while both being right.
  • Ransom amounts treated as a market. The demanded and paid figures are open-source reports about an economy that operates in secret, dominated by cases that failed publicly, and denominated in dollars converted at an unstated rate on an unstated date.
  • The end of the series read as a decline. When the data stops, the chart goes to zero, and a surprising number of published graphics have shown terrorism collapsing in a year for which no data was ever collected.

None of these make the source unusable. They make it a source that requires corroboration before an assertion built on it goes into a product, which is true of every source and admitted by few.

Ageing

The individual records do not go stale in the ordinary sense – a 1983 hijacking is as true today as it was then – but three things do age. First, the series end date ages continuously and silently; a copy of the file that was current when you took it becomes a historical artefact within months, and there is nothing inside the file to warn a later reader. Second, the coding ages: group names get merged and harmonised across releases, coordinates get corrected, and records get added retroactively as new sources surface, so a count you produced from an earlier release will not reproduce exactly against a newer one. Third, interpretation ages: a group coded under one name in 1998 may be understood today as a faction, a franchise or a fiction, and the dataset preserves the contemporary label rather than the current understanding. A stale GTD record looks like a perfectly formed row with a group name nobody uses any more, a coordinate that is nearly right, and a citation to a wire report that is no longer online. The mitigation is to version-stamp everything at ingest and to re-baseline your derived statistics whenever a new release lands, rather than assuming last year's numbers still hold.

What this source feeds

A source is only worth what it lets you conclude. These are the disciplines that collect through it, the mission domains it serves and the data points it yields — every one is a tag, so you can follow any thread from here into the rest of the library.

Collected by these intelligence disciplines

Serves these mission domains

Yields these data points

How each sector uses Global Terrorism Database

The same dataset is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The records are shared — the constraints, thresholds and outputs are not.

🎖 Military and defence

For force protection and theatre entry planning, the GTD is where you build the historical threat picture that current intelligence is a delta against. Query by province and attack type to establish which tactics have historically been used against which target categories in your operating area, what the seasonal and event-linked patterns look like, and whether hostage-taking of foreign personnel has local precedent. The hostage block supports realistic planning assumptions for personnel recovery: typical durations, typical outcomes by actor, and whether a given group has historically released, ransomed or killed. Do not use it for current threat warning – the latency makes that impossible – and remember that it deliberately excludes state force actions, so it will not characterise the conventional threat at all.

🕵 National intelligence

For strategic assessment and warning, the value is in structure rather than in any single record. Longitudinal tactic and target profiles per group let you say whether an actor's repertoire has genuinely shifted or whether recent reporting is a collection artefact, and cross-group comparison within a theatre supports competing-hypothesis work on who is behind an unattributed attack. The claim and uncertainty flags make the dataset unusually honest about attribution, which makes it a useful discipline for an analytic culture that tends to harden attribution too fast. Its principal contribution to an assessment is the denominator: how unusual is this, historically, in this place, by this actor.

👮 Law enforcement

For investigators and prosecutors, the GTD is corroborative background rather than evidence. It will tell you whether the group your suspect is associated with has a documented history of a given tactic in a given place, which supports charging decisions and the framing of expert testimony, and the source citations give you a route back to contemporaneous press reporting that can be independently authenticated. It is not admissible as proof that a specific incident occurred, the coding is not a legal determination, and the perpetrator field must never be presented to a court as an attribution finding. For kidnap-for-ransom casework, the outcome and duration distributions are useful for negotiation planning and for explaining base rates to a family or a corporate crisis team.

🔍 Private investigation and corporate security

In corporate security and private investigation, this is the source behind most credible country and city risk statements about terrorism and kidnap exposure. Use it to answer the question a client actually asks – has anything like this happened here before, to people like us – with a defensible historical count rather than a vendor's colour-coded map. Filter by target type to isolate business, NGO or private-citizen targeting instead of quoting an all-target national total that is dominated by attacks on security forces. Be explicit with the client about the series end date; presenting a historical baseline as a current assessment is the single most common way this source is misused commercially.

📰 Journalism and OSINT media

For investigative and data journalism, the GTD is a well-documented, citable dataset that will survive a fact-check if you use it carefully and embarrass you if you do not. It supports stories about long-run patterns, about how a group's targeting has changed, and about the gap between public perception and recorded incidence. The obligations are specific: state the coverage window, state your inclusion-criteria filter, do not chart across 1993 without marking the gap, and do not describe the fall-off at the end of the series as a decline. Where a story turns on a single incident, use the source citations to return to the original reporting and verify it there rather than citing the database for the fact.

🌍 NGO, humanitarian and human rights

For humanitarian security management, the hostage and target blocks are directly operational. Filtering to NGO, journalist, private-citizen and religious-figure target types in a given country produces the historical basis for a duty-of-care assessment, a staff movement policy and an insurance conversation, and the duration and outcome fields inform what a realistic abduction scenario looks like locally. Handle it victim-centred: these rows are people, several of them identifiable from the free-text target fields, and republishing names attached to a violent incident can harm survivors and families long afterwards. Aggregate for planning, and keep individual records inside the security team.

🎓 University and research

This is the canonical teaching and research dataset for quantitative terrorism studies, and its documentation is a model of what event-data transparency should look like. Its research uses are broad – diffusion and contagion modelling, group lifecycle analysis, target substitution, the effect of counterterrorism policy on tactic choice – but every one of them lives or dies on how you handle the collection discontinuities and the definitional filters. Publish your filter choices alongside your results, run your models under at least two inclusion regimes as a robustness check, cite the specific release, and archive the file you used. Reviewers in this field know the 1993 gap and the methodology breaks, and a paper that ignores them does not survive.

Playbook: working Global Terrorism Database end to end

A repeatable sequence from first pull to finished product. Each phase states what you are trying to establish, not merely what to click — the objective is a defensible chain of reasoning, not a completed checklist.

Phase 1 — Read the codebook before the data

Open the codebook and read the inclusion criteria, the hostage block definitions and the list of collecting organisations by period. This costs an hour and prevents the four or five errors that dominate misuse of this dataset. Write down the release identifier and coverage end date at the top of your working notes, because you will need both in the final product.

Phase 2 — Fix your definition of the phenomenon

Decide, before you look at any number, whether your question is about terrorism as the GTD defines it or about political violence more broadly. Then set your filter: all three criteria or any of them, doubt-flagged records in or out. Record the choice and apply it consistently, because switching filters mid-analysis is how two sections of the same report end up disagreeing.

Phase 3 — Scope the geography honestly

Select your country and administrative divisions, then immediately check the specificity distribution of the coordinates in that selection. If most of your rows are coded to city or region rather than to an exact site, you have a subnational analysis and not a local one, and you should say so rather than drawing a map that implies otherwise.

Phase 4 — Establish the collection baseline

Chart annual event counts for your selection and overlay the collecting-organisation boundaries and the 1993 gap. This is not analysis, it is instrument calibration: you are learning where your series can and cannot support a trend claim. Any inflection sitting on a boundary is presumed an artefact until proven otherwise.

Phase 5 — Isolate the kidnap and hostage subset

Filter on the hostage flag and the two hostage-taking attack types rather than on keywords in the summary. Then examine how many of those rows have populated victim counts, durations and outcomes – typically far fewer than the total – because that populated subset, not the full filter, is the real denominator for anything you say about duration or outcome.

Phase 6 — Profile the actors present

Group by perpetrator name within your selection and split the result by the claim and uncertainty flags. You are looking for three things: which actors dominate, how much of the theatre is unattributed, and whether a named group's presence reflects claimed operations or press attribution. The unattributed share is itself a finding about the information environment.

Phase 7 — Characterise targeting, not just volume

Cross-tabulate target type and subtype against attack type for your actors of interest. This is where a group's operational character shows: whether it takes hostages opportunistically at checkpoints or plans abductions of specific categories of person, whether it attacks infrastructure or people, and whether it substituted tactics after a security change.

Phase 8 — Derive the outcome distribution

For the populated hostage subset, compute release, escape, rescue and fatal outcomes by actor and by victim nationality where coded, alongside duration quartiles. Present these as ranges with explicit sample sizes. A crisis-response client needs to know that an estimate rests on eleven cases, and they will make worse decisions if you round that away.

Phase 9 — Test the ransom fields before quoting them

Look at how many rows in your subset carry a demanded amount, how many carry a paid amount, and how many carry both. Usually it is a small and unrepresentative fraction. If you quote a figure, quote it as reported, name the incident, and state the sample. Never present an average ransom for a region as if it were a market price.

Phase 10 — Cross-check the recent end against live sources

Take the last two years of your GTD selection and compare them against a current-awareness event feed and local reporting. You are measuring the gap: how much has happened since the series ended, and whether the pattern in that gap resembles the baseline. This step converts a stale dataset into a usable one by making its staleness explicit and quantified.

Phase 11 — Verify the records your conclusion rests on

Identify the individual rows doing the analytical work and go back to their source citations. Confirm the incident, the date, the location and the attribution in the original reporting. If a citation is dead, find a replacement or drop the claim. A conclusion resting on three rows must be able to survive an editor pulling all three.

Phase 12 — Publish with the instrument attached

State the release, the coverage window, the inclusion filter, the geographic selection and the sample sizes. Mark the 1993 gap and the methodology boundaries on any chart that crosses them, and put the series end date on the face of the graphic rather than in a footnote. This is what separates a defensible baseline from a number a competitor can dismantle in one sentence.

The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.

What to pair it with

No single source carries a finding. These are the datasets that corroborate, extend or contradict this one — and a source that contradicts is worth more than one that agrees, because it is the only thing that will tell you when you are wrong.

Source Relationship What it adds
ACLED extends Near-real-time coded political violence and protest events with far shorter latency and explicit coverage of state actors. The natural partner for the years the GTD does not reach, though its event definition is broader and the two are not directly countable together.
UCDP Georeferenced Event Dataset corroborates Fatality-anchored organised violence events from Uppsala, framed by conflict dyad rather than by terrorism. Excellent for testing whether a GTD pattern survives a different definitional lens.
GDELT extends Machine-coded global news event stream. Vastly higher recall and vastly lower precision than the GTD; use it to detect that something happened recently, not to count it.
Global Terrorism Index extends An annual index built substantially on GTD rows. Useful as an independent implementation of aggregation logic you can benchmark against, and as the version of this data your client has probably already read.
START research portfolio extends The same centre maintains complementary collections on individual radicalisation and on the organisational characteristics of violent groups, answering the who and why questions an event schema cannot.
Europol TE-SAT corroborates Official European reporting on terrorist incidents, arrests and disrupted plots. Its coverage of failed and foiled plots fills a gap the GTD cannot, since a disrupted attack generates no incident.
UNODC data and analysis extends Homicide, organised crime and trafficking statistics that provide the criminal-violence context in which politically framed kidnapping frequently sits.
CTC Sentinel corroborates Analytical writing on specific groups and campaigns that is often the best available narrative explanation of a statistical cluster you have found in the data.

Legal, ethical and operational constraints

The immediate legal constraint is contractual rather than statutory: you accepted terms of use at download, and those terms have historically restricted commercial use and redistribution. Honour them, and check them again rather than assuming, because they have been revised. Beyond the licence, the data protection position deserves thought. The free-text target and summary fields contain the names of individuals – victims, hostages, named officials – and in most jurisdictions that makes portions of this dataset personal data, some of it in a special category because it concerns criminal offences and sometimes health or religion. Research and journalism exemptions exist in many regimes but they are conditional, and none of them survives republishing a named kidnap victim in a commercial risk product. Apply the handling you would to any victim-identifying material: minimise, aggregate for external products, keep individual rows inside the case, and document your basis for retention. In law-enforcement contexts, remember that a GTD row is a research coding of a press report and nothing more; it is not an official record, not evidence of an offence, and presenting it to a tribunal as an attribution finding is an error a competent defence will exploit.

Operational security

Downloading the GTD is a registered, attributable act: you supplied an identity and an institutional affiliation to a university research centre, and that registration links your organisation to an interest in terrorism data. For most users this is unremarkable and carries no risk. It matters in two situations. If you are working under cover or through a front entity, do not register with the operational identity and do not use infrastructure that ties the download to a covert operation. If you are a journalist or NGO staffer working in a state that treats terrorism research as suspect, consider that both a registration record and a downloaded file exist and can be found on a seized device. The queries you run afterwards are local and invisible, which is a practical advantage of a bulk dataset over an API: once you hold the file, your analytical interest in a particular group, province or victim category is disclosed to nobody. Take the file once, hold it locally, and do the sensitive work offline.

Two rules that hold regardless of jurisdiction. Collection that is lawful is not automatically proportionate, and a dataset assembled for one purpose does not carry consent for another. Where the records concern identifiable people, the question is not only whether you may hold the data but whether holding it serves the purpose you are accountable for.

Is it earning its place?

Sources accumulate. Feeds get added during an incident and are never reviewed again, and a decade later the pipeline is carrying dead weight that nobody dares remove. These are the measures that show whether Global Terrorism Database is contributing anything, and they are worth baselining now so the answer is available later.

  • Share of your kidnap-related questions that the GTD can answer with a populated field rather than a blank – if the outcome and duration fields are empty for your theatre, this source is not earning its place there.
  • Gap size: the number of months between the GTD series end and the present for each theatre you cover, tracked deliberately, because it determines how much live collection you must fund to stay current.
  • Reproduction rate when you verify sampled records against their source citations, and the proportion of those citations that still resolve.
  • Divergence between your GTD-derived counts and an independent event dataset over the overlapping years, expressed per country – a large but stable divergence usually indicates a definitional difference worth documenting once and reusing.
  • Proportion of records in your selection with an unknown or uncertain perpetrator, tracked over time as a measure of the information environment rather than of the data.
  • Number of published products in which the coverage window and inclusion filter were stated on the face of the graphic, as a straightforward quality-control count.
  • Analyst time spent normalising the file at each new release, which tells you whether your ingest absorbs schema drift or whether you pay for it manually every year.

Beware of volume. Indicator counts rise easily and say almost nothing. Unique contribution — findings this source produced that no other source in your stack would have — is the measure that matters, and it is usually far lower than anyone expects.

Tradecraft notes

The distinctions that separate a competent analyst from a fast one:

  • Blank is not zero and unknown is not none. Screen the sentinel values at ingest, and when you report a mean, report the count of populated rows next to it every single time.
  • A trend break that lands on a change of collecting organisation is a collection artefact until independently corroborated. Keep the boundaries drawn on your working charts so you cannot forget where they are.
  • Specificity governs geography. Never plot a point at a resolution the specificity code does not support, and never let a heat map aggregate exact sites and country centroids into the same surface.
  • Attribution and claim are separate variables because they are separate facts. An analyst who collapses them is asserting knowledge the dataset explicitly declines to assert.
  • The free-text target, corporate and summary fields are the least exploited and most valuable part of the schema. Entity extraction over them surfaces named people, companies and vessels that the coded fields never expose.
  • Filter to the populated subset before you characterise any distribution. The hostage outcome analysis you can defend rests on the rows where the outcome was coded, not on every row where a hostage was taken.
  • Use the dataset to establish the denominator and something else to establish the numerator. Its comparative advantage is base rates; the moment your question becomes what happened last quarter, you are holding the wrong instrument.
  • Archive the exact file, not the download link. Releases revise history, and a claim you cannot reproduce is a claim you will eventually have to retract.
  • When you brief a client, lead with the series end date. Everything else you say is conditional on it, and stating it first removes the single most damaging misunderstanding before it forms.

Questions analysts actually ask

Is the GTD current enough to support a live threat assessment?

No. It is a historical baseline with a fixed end date that has often lagged the present by years. Use it to establish what is normal for a place and an actor, and pair it with a current-awareness feed and local reporting for anything about the present. Presenting it as current is the most common professional misuse of this source.

Why do my totals not match the ones in a published paper?

Almost always because of a different inclusion filter, a different release, or a different treatment of doubt-flagged records. Check all three before assuming an error. The GTD supports several defensible counts of the same phenomenon, which is why stating your filter is not pedantry but a requirement.

What actually happened to 1993?

The records for that year were lost before the dataset was digitised and were never fully recovered at incident level. Some aggregate reconstructions exist but they cannot be joined to the incident data. Mark the gap on charts and exclude the year from year-on-year calculations rather than interpolating it.

Can I use the ransom fields to estimate what a kidnap costs in a region?

Not responsibly. The demanded and paid amounts are open-source reports about a deliberately opaque market, populated for a small and biased subset of cases, and skewed towards incidents that became public because they went wrong. Quote individual reported figures with their incident and sample size, and refuse to produce a regional average.

Does it include attacks by governments?

No. State military and police violence carried out as state action falls outside the definition. In theatres where the state is the principal source of violence against civilians, this omission is not a minor caveat – it changes what the dataset appears to say about who is dangerous.

Can I redistribute the rows to a client in a deliverable?

Assume not without checking. The terms of use have restricted redistribution and commercial exploitation, and they have changed over time. Derive analysis, cite the release, and seek written permission before anything resembling resale or bulk transfer.

How should I handle victim names in the free-text fields?

As personal data about people who were subjected to violence. Aggregate for external products, keep named rows inside the case file, apply your normal retention rules, and do not republish a named victim in a commercial risk report even though the row is technically public.

Is a rise in recorded attacks in a country a rise in terrorism?

Only if press coverage, collection methodology and your inclusion filter all held constant across the period. Frequently at least one did not. Test the alternative explanation before you assert the substantive one, and say in the product which you tested.

What is the fastest way to sanity-check my selection?

Pull ten incidents you already know from independent primary material and look them up. Within an hour you will learn whether the geography, the casualty counts and the attributions in your selection behave the way you assumed, and that hour will save you a retraction.

Standards, formats and interoperability

What this source speaks natively, and what it has to be translated into before a partner can consume it. Work that arrives in a recognised format is easier to defend, easier to hand over and easier to automate against:

  • Fixed-schema event coding with a published codebook, which makes it directly convertible to STIX 2.1 incident and attack-pattern objects with observed-data provenance pointing back at the source citation.
  • Its attack-type and weapon-type taxonomies predate and differ from ATT&CK-style frameworks; map them explicitly rather than assuming correspondence with taxonomies used elsewhere in your platform.
  • Geographic coding to country, first-order administrative division and point coordinates with an explicit precision qualifier – a cleaner precision model than most event datasets offer, and one that should be preserved through any GIS export.
  • Perpetrator names are free-text strings harmonised over time, not identifiers from any designation list; reconciliation to UN, EU, OFAC or national proscription lists is manual work you must do and document.
  • The three inclusion criteria constitute an explicit operational definition of terrorism that can be cited in a methodology section, and it is one of the few such definitions with a coded dataset attached.
  • Compatible with standard tabular analysis and with CSV, JSON and JSONL export once normalised; the coded and text variants of most variables travel together and both should survive export.
  • The hostage block maps cleanly onto kidnap-for-ransom case taxonomies used in crisis response, treating victim count, duration, demand, payment and outcome as separate variables rather than as narrative.

References

Primary documentation and authoritative references for this source. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.

  1. Global Terrorism Database — START, University of Maryland. The dataset home: download access, the search interface, the codebook and the release notes. Check the coverage end date here before scoping any work.
  2. START — University of Maryland. The parent research centre, its funding history and its other datasets on radicalisation and group characteristics. Useful for understanding why the GTD is shaped as it is.
  3. ACLED — Armed Conflict Location and Event Data Project. The main alternative event dataset, with far lower latency and explicit state-actor coverage. Read its methodology alongside the GTD codebook to see what each definition includes.
  4. UCDP — Uppsala University. Georeferenced organised violence events anchored on fatalities. The best third opinion when two terrorism datasets disagree about a theatre.
  5. GDELT Project — GDELT. Machine-coded global news events. Useful for closing the recency gap at the end of the GTD series, with the precision penalty that implies.
  6. Vision of Humanity — Institute for Economics and Peace. Home of the Global Terrorism Index, the most widely read derived product of this dataset and often the version a client has already seen.
  7. Europol — European Union Agency for Law Enforcement Cooperation. Publisher of the annual EU terrorism situation and trend report, which covers foiled and failed plots that no incident dataset can capture.
  8. UNODC — United Nations Office on Drugs and Crime. Statistical and analytical context on organised crime, homicide and kidnapping as criminal rather than political phenomena.
  9. Combating Terrorism Center — United States Military Academy. Group and campaign analysis that frequently supplies the narrative explanation for a statistical cluster you have identified.

Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.

Put it into practice

The Quantus Intel threat intelligence platform operationalises this source: it loads the GTD as a version-stamped historical baseline, carries specificity and attribution uncertainty through to the map and the actor profile, and shows on the face of every timeline where the coded series ends and live collection begins.. Browse the full source catalogue, or follow any tag above into the rest of the library.

Leave a Reply