August 16, 2026

Counter-Trafficking Data Collaborative: Intelligence Source Guide

0

CTDC publishes de-identified, case-level records of individual trafficking victims contributed by IOM, Polaris and partner organisations. It is the only open dataset that describes what happened to identified victims rather than estimating how many exist.

counter-trafficking-data-collaborative-intelligence-source-guide

CTDC publishes de-identified, case-level records of individual trafficking victims contributed by IOM, Polaris and partner organisations. It is the only open dataset that describes what happened to identified victims rather than estimating how many exist.

At a glance

Source Counter-Trafficking Data Collaborative
Category Conflict, Crime & Human Security › Human Trafficking & Child Protection
Homepage https://www.ctdatacollaborative.org/
Machine interface https://www.ctdatacollaborative.org/download-global-dataset
Format CSV
Access Open — no account required
Disciplines Human Intelligence, Open Source Intelligence
Mission domains Human Trafficking

Largest global victim-level dataset (IOM). — as catalogued in the platform’s own source registry.

The Counter-Trafficking Data Collaborative is a data hub run by the International Organization for Migration together with counter-trafficking partners including Polaris, which operates the United States national human trafficking hotline, and other case-holding organisations. Its core product is a global dataset in which each row is one identified victim of trafficking, described by a fixed set of categorical variables: year of registration, the contributing data source, gender, a banded age, country of citizenship, country of exploitation, the type of exploitation suffered, the sector of labour or the form of sexual exploitation where applicable, the means of control used by the traffickers, and the relationship between the victim and the recruiter. The data originates in case management systems – the records that assistance providers create when they identify and support a survivor – and is passed through statistical disclosure control before publication. That control is the defining technical feature: values are generalised and suppressed so that no combination of attributes in the published file corresponds to a small enough group of individuals to permit re-identification. CTDC also hosts derived and synthetic datasets, including work with UNODC on victim-perpetrator relationships, and publishes visualisations and a data dictionary alongside the downloads.

Every other open resource on human trafficking gives you an estimate, an index or a policy narrative. This one gives you cases. The analytical job it does is describe the shape of exploitation: which control mechanisms co-occur, which recruitment relationships are associated with which exploitation types, how age at entry relates to sector, which citizenship-to-exploitation-country pairs recur, and how these patterns differ between labour and sexual exploitation. Those are questions about mechanism, and mechanism is what determines whether an intervention works. A prevalence estimate tells a minister how big the problem is; a means-of-control distribution tells a caseworker what to screen for and a prosecutor which statutory element is most often provable. For HUMINT and OSINT work on trafficking networks, the dataset is best used as a base rate: it tells you what is typical for a corridor, so that an atypical case in front of you is legible as atypical. It is emphatically not a prevalence source and any use of it to say how many victims exist is a misuse the publishers explicitly warn against.

Who publishes it, and why that matters

IOM is the UN migration agency and the largest single provider of direct assistance to trafficking victims worldwide, which is why it holds the largest case corpus and why it anchors the collaborative. Polaris operates the US hotline and contributes a very different kind of case – hotline-derived, US-centric, self-reported. Other contributors have joined over time and the contributor list is not static. This matters more than it sounds. The dataset is not a sample of trafficking; it is a union of several organisations' caseloads, each with its own intake criteria, geography, funding-driven programme footprint and definitional practice. When a new contributor joins, the shape of the whole dataset changes, and that change is a change in who is contributing, not in the world. The incentive structure is genuinely good – the collaborative exists because the counter-trafficking field recognised that its evidence base was embarrassingly thin and that data hoarding by agencies was the reason – and the publishers are unusually explicit about the limitations. Funding is institutional and donor-based, which makes continuity reasonably likely but publication cadence irregular.

Provenance is the first question to ask of any dataset and the one most often skipped. Who collects it, what their incentive is, whether they publish a methodology, and whether they correct the record when they get something wrong all bear directly on how much weight a finding drawn from it can carry.

What a record actually contains

The fields you will be working with, what each one means, and whether it is something you can pivot on. Read the meanings carefully — more analysis is wrecked by misreading a field than by failing to find one, and a field that looks like an observation is often an inference.

Field Type What it means Pivot value
yearOfRegistration int The year the case was registered by the assisting organisation. This is not the year of trafficking, not the year of escape and not the year of the offence – it is an administrative date, and the gap between exploitation and registration can be years. Cohort analysis; alignment with programme funding cycles and with the contributor's operational timeline rather than with events.
Datasource enum Which contributing organisation supplied the record. The most important field in the file and the one most often ignored. Almost every apparent pattern in the data needs to be checked for whether it is really a difference between contributors. The contributor's own methodology, geographic footprint and intake criteria; stratify every analysis by this field.
gender enum Recorded gender of the victim, with a small set of values and substantial missingness. Reflects both who is trafficked and who is identified, and identification systems in most countries are far better at recognising female victims of sexual exploitation than male victims of labour exploitation. Cross-tabulate against exploitation type and sector, always with the identification-bias caveat attached.
ageBroad enum Age in bands rather than years, because exact age would be identifying. The bands are narrow at the young end and wide at the old end, which is a disclosure-control artefact and not a statement about where the analytically interesting boundaries are. Minority-status flags, education and school-age analysis, and sector-specific age profiles.
citizenship string Country of citizenship as an ISO code, frequently suppressed where the combination with other fields would be identifying. Suppression is not random – it hits small nationality groups hardest, which are often the most analytically interesting. Corridor construction with country of exploitation; migration route and visa regime analysis; diaspora and recruitment-network context.
CountryOfExploitation string Where the exploitation took place, again as an ISO code with suppression. Note that a case may involve exploitation in several countries and the field carries one, so multi-country trajectories are flattened. Corridor pairs, destination-country labour market and enforcement regime, and the national referral mechanism relevant to the case.
isForcedLabour / isSexualExploit / isForcedMarriage / isOtherExploit enum Binary flags for the broad exploitation category. They are not mutually exclusive – many victims experience more than one – and treating them as a single categorical variable will silently drop the overlapping cases that matter most. Legal element analysis, sector cross-tabulation, and comparison against national statutory definitions.
typeOfLabourConcatenated array The labour sectors involved, concatenated into one string of codes: agriculture, construction, domestic work, hospitality, manufacturing, mining, begging and others. Needs splitting before analysis, and the code set has changed between releases. Sector-level supply-chain risk, employer and recruitment-agency analysis, and alignment with forced-labour indicator frameworks.
typeOfSexConcatenated array Forms of sexual exploitation recorded, similarly concatenated. Handle with care in any published product: the categories are clinical, the underlying experience is not, and aggregate reporting is the appropriate level. Venue and platform typology in the destination country; screening indicators for frontline staff.
meansOfControl* (a family of flags) enum A set of binary indicators covering debt bondage, withholding of documents, withholding of wages, restriction of movement, threats, physical and sexual abuse, psychological abuse, false promises, psychoactive substances, threat of law enforcement, use of children against a parent, excessive working hours and denial of medical care. Collectively the richest part of the dataset. Map directly onto the ILO indicators of forced labour and onto the means element of the Palermo Protocol definition; use for screening tool design and prosecutorial element analysis.
recruiterRelation* (a family of flags) enum The relationship between the victim and the person who recruited them: intimate partner, family member, friend, other, unknown. The single most policy-relevant variable in the file because it destroys the stranger-abduction model of trafficking that shapes most public understanding. Prevention messaging design, community-level intervention targeting, and network analysis of recruitment structures.
majorityStatus / majorityStatusAtExploit / majorityEntry enum Whether the victim was an adult or a minor at registration, at the time of exploitation, and at entry into the trafficking situation. Three distinct questions that are routinely conflated, and the distinction is legally decisive because the means element is not required for child victims. Child-protection referral pathways; applicable legal framework; national child-protection authority in the country of exploitation.
traffickMonths int Duration of the trafficking situation in months where recorded. Heavily missing, and where present it rests on the survivor's recollection of a period during which time perception is frequently distorted by the conditions of exploitation. Sector-specific duration profiles; correlation with control mechanisms and with escape or exit route.
isAbduction / isOrganRemoval / isForcedMilitary / isSlaveryAndPractices enum Flags for less common exploitation forms. Counts are small, which means disclosure control suppresses them aggressively and any analysis of them from the public file will be badly underpowered. Specialist referral pathways; qualitative case literature rather than quantitative analysis of this file.

Coverage — and what is not in it

Coverage is global in the sense that citizenships and countries of exploitation span most of the world, and profoundly uneven in the sense that the distribution reflects where the contributing organisations operate. IOM's caseload is shaped by its programme footprint – strong in the post-Soviet space, South East Asia, the Middle East, parts of Africa and along the major migration corridors where IOM runs assistance programmes – and Polaris's contribution is overwhelmingly United States. Regions where neither operates at scale, or where victim identification systems barely function, are thinly represented or absent regardless of how much trafficking occurs there. The time window runs from the early 2000s onwards, with case volumes rising over time in a way that reflects the growth of case management systems and of the collaborative itself rather than a rise in trafficking. Entity types are individual victims described by categorical attributes only: no names, no identifiers, no free text, no dates finer than a year, and no perpetrator identity beyond the recruiter-relationship flags. The dataset is republished periodically rather than continuously, and the cadence has not been regular.

Known blind spots

Absence of evidence here is not evidence of absence. These are the conditions under which Counter-Trafficking Data Collaborative will not show you something that is nevertheless real:

  • Only identified victims appear. A victim becomes a record when they came into contact with an assistance provider that contributes data, which means the dataset is a picture of the identification and assistance system, not of trafficking. Everyone who was never identified, was identified by a non-contributing agency, or refused assistance is absent.
  • Male victims of labour exploitation are systematically under-identified in most national systems, so the dataset understates them regardless of underlying prevalence. The same is true of victims exploited in sectors that inspectors do not enter, notably domestic work and small-scale agriculture.
  • Disclosure control removes exactly the rare combinations that are analytically most interesting. A small nationality group in a specific destination and sector is suppressed precisely because it is distinctive, so the dataset is blind to niche corridors by design.
  • Multi-stage and multi-country trajectories are flattened into single-value fields. A victim recruited in one country, transited through a second and exploited in a third and then a fourth appears with one country of exploitation, which destroys route structure.
  • There is no perpetrator data beyond the recruiter relationship. Networks, facilitators, employers, recruitment agencies, transport and financial infrastructure are all outside the file, so it cannot support network analysis on its own.
  • Internal trafficking is under-represented relative to cross-border, because the contributing organisations' mandates and funding are oriented towards migration-related casework, while a substantial share of trafficking globally never crosses a border.
  • The year of registration is an administrative artefact, so any time series measures the contributors' operational activity – programme openings, funding cycles, a new hotline campaign – at least as strongly as it measures anything about the phenomenon.
  • Exploitation occurring inside conflict zones, detention, state-imposed forced labour and institutional settings is largely absent, because assistance providers do not have access to those populations and the victims do not reach referral systems.
  • Definitional variation between contributors is not fully harmonised. What one organisation records as trafficking another records as exploitation or smuggling with aggravating features, and the file cannot tell you which convention produced a given row.

Write the blind spot into the product. A statement that something “was not observed in Counter-Trafficking Data Collaborative” is defensible; a statement that it “did not happen” is not, and the difference is what survives cross-examination.

Access, licensing and what you may do with it

Access model: Open — no account required

The global dataset is downloadable from the CTDC site without authentication, delivered as a flat CSV with an accompanying codebook. Take the codebook seriously: the column set and code values have changed between releases, casing is inconsistent across columns, and several fields are concatenated multi-value strings that need splitting before use. Download the codebook that matches the file version you took, and keep both. There is no API, so treat collection as a periodic manual or scheduled fetch with version tracking rather than as a feed. Beyond the global file, the site publishes derived datasets and visualisations, and researchers wanting access to less aggregated data must approach the contributing organisations directly, which involves a research ethics process and, in most cases, an institutional agreement. Do not expect that route to be quick, and do not expect it to yield anything approaching case-level detail with identifiers – the contributing organisations hold survivor data under strict protection obligations and the disclosure-controlled public file exists precisely so that they do not have to release more.

Licence

The public dataset is released for research, policy and public-interest use with attribution to the collaborative and its contributors, subject to terms published on the download page which have been updated over time. Confirm the current terms at the point of download rather than relying on a description. Two conditions matter regardless of what the licence text says. First, no re-identification: any attempt to link records against other datasets in order to identify individuals violates the basis on which the data was released and, in most jurisdictions, data protection law. Second, responsible presentation: the contributing organisations' consent basis with survivors rests on the data being used to improve responses, and publishing it in ways that sensationalise or that could stigmatise a nationality or community breaches the spirit of that basis even where it does not breach the letter. If your intended use is commercial, ask before you build.

Rate limits and fair use

Not applicable in the usual sense – this is a file download, not a service. The etiquette is to fetch once per release rather than polling, cache the file and its codebook locally with the download date recorded, and check for a new version on a monthly or quarterly cadence. Nothing about this data changes fast enough to justify anything more frequent, and repeated automated downloading of a static research file is both pointless and rude.

Licensing changes, and it changes without warning. A dataset that was free for research this year may not be free for commercial or evidential use next year. Confirm the current terms before you build a dependency on it, and record the terms you relied on alongside the data — the licence in force at the time of collection is part of the provenance.

Collecting it

How Counter-Trafficking Data Collaborative is actually pulled, in the order you would set it up. Prefer the bulk or export interface over per-item lookups wherever one exists: it is kinder to the publisher, faster for you, and gives a reproducible snapshot rather than a series of point-in-time answers you cannot reconstruct later.

Method Format Cadence Notes
Global dataset download CSV Per release, check quarterly The primary route. Store the file with its download date and the matching codebook, and version it – analyses run against different releases are not comparable.
Codebook and data dictionary capture HTML With every download Non-optional. Column names, code values and the disclosure-control approach have all changed between releases, and the codebook is the only record of what a given file's values mean.
Derived and synthetic datasets CSV Ad hoc The collaborative publishes additional datasets including synthetic victim-perpetrator work with UNODC. These have different provenance and different limitations; read their documentation separately rather than assuming continuity with the global file.
Contributor-side reporting HTML Annual IOM, Polaris and other contributors publish their own analyses of their own caseloads. These carry context the harmonised file strips out and are the best way to understand what a contributor's rows actually represent.
Direct research request bulk Project-based For work needing more than the public file supports, approach the contributing organisation with an ethics-reviewed protocol. Expect a long process and a narrow grant, and expect to sign undertakings on re-identification and publication.

Ingesting it into the platform

Every step below is idempotent and cursor-based: interrupt one and it resumes from where it stopped rather than duplicating rows or losing progress. Collection is recorded per source, so a feed that quietly stops publishing shows up as a stale timestamp instead of silently thinning your coverage.

  1. Register as a versioned dataset, not a feed — Add the source in sources.php with its collection route recorded as a periodic file download and a version field for the release, so that datasets.php shows which release any analysis was run against and nothing downstream assumes currency.
  2. Load with the codebook as a schema contract — Import through import.php using an explicit column mapping derived from the release's codebook, failing loudly on unknown columns or unknown code values rather than coercing them, because silent coercion is how release changes become undetected analytical errors.
  3. Split concatenated multi-value fields — Expand the labour-type, sexual-exploitation-type, means-of-control and recruiter-relation fields into proper multi-valued attributes at load time. Analysis performed against the raw concatenated strings will miscount every co-occurrence.
  4. Preserve suppression as a distinct state — Encode suppressed and missing values as separate, explicit states rather than nulls. The difference between not recorded and removed for disclosure control is analytically decisive and disappears if both become blank.
  5. Stratify by contributor at load — Materialise the data source as a first-class dimension so every downstream aggregation is available split by contributor. Any cross-contributor total should be flagged in the interface as a union of caseloads rather than a population figure.
  6. Map control mechanisms to the standard frameworks — Use resolve-tags.php to map the means-of-control flags onto the ILO indicators of forced labour and onto the act-means-purpose structure of the Palermo Protocol, so that the data joins to legal and supply-chain frameworks rather than sitting in its own vocabulary.
  7. Build corridor aggregates without individual rows — Generate citizenship-to-exploitation-country aggregates for human-trafficking.php and country.php with a minimum cell size enforced in the platform itself, so that platform-side aggregation never reconstructs a small group the publisher deliberately protected.
  8. Attach the caveat to the output, not the documentation — Configure reports.php so that any product drawing on this source carries the identified-victims caveat inline. The most likely failure mode of this dataset in an intelligence platform is a clean-looking chart presented as prevalence, and the fix has to be structural rather than a note in a handbook.

Registered sources and their last-collected state are listed in sources.php, and the scheduled chain that keeps them current is in automation.php.

How it is wrong, and how to tell

Every dataset is wrong in characteristic ways. Knowing which ways is the difference between using a source and being used by one, and it is the part of source evaluation most often skipped because it is the part that takes work.

Within its own frame, this is careful, well-documented work of a standard rarely seen in the counter-trafficking field. The variables are defined, the codebook is published, the disclosure control is described, and the publishers state the selection-bias limitation prominently rather than burying it. The individual records rest on case management by professionals who interviewed the survivor, which is a far better provenance than the survey extrapolation underlying most trafficking statistics. The weakness is entirely structural and entirely unavoidable: the selection mechanism is victim identification, which is one of the most biased processes in the whole of criminal justice, differing by country, by gender, by exploitation type, by immigration status and by whether a victim is treated as an offender on first contact. That bias is not noise you can model away with the fields available, because the file contains no information about the identification process itself. So judge it as excellent data about the population it describes – people identified and assisted by these organisations – and as no evidence at all about the population it does not.

Characteristic false positives

  • Reading the dataset as prevalence. The commonest and most damaging error: an increase in records for a country is an increase in identification and assistance activity, and a country with few records may simply have no functioning referral system and no contributing organisation.
  • Contributor composition effects mistaken for real change. When a contributor joins, expands a programme or changes its intake criteria, every distribution in the file shifts. Time series that are not stratified by data source will show phantom trends at exactly those moments.
  • Treating exploitation-type flags as mutually exclusive. Many victims experience more than one form, and analyses that force a single category either drop those cases or assign them arbitrarily, systematically understating combined labour and sexual exploitation.
  • Under-counting caused by suppression read as real absence. Disclosure control removes small cells, so a corridor with genuine but low-volume activity can appear as zero. Absence in this file is never evidence of absence in the world and frequently is evidence of distinctiveness.
  • Registration year treated as event year. Cohorts built on registration date place victims in the wrong period, which corrupts any attempt to relate trafficking patterns to external events such as a conflict, a border closure or a legal change.
  • Concatenated fields miscounted. Multi-value strings analysed as single categorical values produce long tails of spurious categories and undercount every individual sector, and this error survives into published charts because the output still looks plausible.
  • Self-report distortion in duration and means fields. Survivors recall a period of coercion, often after trauma and often in an interview where disclosure has consequences for their immigration status or their safety, and both under-disclosure and compression of timelines are well documented.
  • Nationality-level inference sliding into stigma. Because citizenship is present and identification systems differ enormously by nationality, it is easy to produce a chart that appears to show which nationalities are trafficked and actually shows which nationalities the assistance system recognises.

None of these make the source unusable. They make it a source that requires corroboration before an assertion built on it goes into a product, which is true of every source and admitted by few.

Ageing

The dataset itself does not go stale so much as accumulate: old records remain valid descriptions of the cases they describe. What ages is applicability. Trafficking routes, recruitment methods and control mechanisms change with migration policy, labour-market conditions, conflict and technology, so a control-mechanism profile from a decade ago may no longer describe current cases in the same corridor – the shift of recruitment to social platforms and of payment to digital channels is a visible example that the older records cannot show. Corridor patterns age faster than mechanism patterns; the fact that debt bondage and document retention dominate in a given sector is durable, while the specific citizenship-destination pairs are not. A stale use of this data looks like a corridor claim drawn from a decade-old cohort, or a screening tool built on a control-mechanism distribution that predates the shift of recruitment online. Refresh your extract at each release, and re-run any operational profile against the most recent three to five years of registrations rather than the full file.

What this source feeds

A source is only worth what it lets you conclude. These are the disciplines that collect through it, the mission domains it serves and the data points it yields — every one is a tag, so you can follow any thread from here into the rest of the library.

Collected by these intelligence disciplines

Serves these mission domains

Yields these data points

How each sector uses Counter-Trafficking Data Collaborative

The same dataset is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The records are shared — the constraints, thresholds and outputs are not.

🎖 Military and defence

The relevance is narrow and real. In stabilisation, peacekeeping and force-protection contexts, trafficking exists around deployed forces – in labour supply to bases through recruitment agencies, in the contractor supply chain, and in the sexual exploitation that has repeatedly accompanied deployments. The means-of-control and recruiter-relation distributions give you a screening profile for contracted labour: retained passports, unexplained deductions, accommodation controlled by the employer, recruitment fees. Use it to inform contractor vetting and troop education, apply the ILO indicator mapping to labour supplied to your installations, and route any identified case to civilian protection actors rather than handling it as a security matter.

🕵 National intelligence

Treat this as a base-rate and structure source for HUMINT and OSINT work on trafficking networks, not as a targeting input – it contains no perpetrators, no identifiers and no locations finer than a country. Its analytical contribution is knowing what normal looks like for a corridor, so that a case with an unusual control profile or an unusual recruitment relationship stands out as worth pursuing. It is also the right source for briefing policymakers on mechanism, because it is open, citable and defensible in a way that agency reporting on the same subject usually is not. Stratify everything by contributor before drawing a conclusion about a country.

👮 Law enforcement

Two concrete uses. First, element analysis: the means-of-control distribution tells you which coercion mechanisms most commonly appear in cases like the one in front of you, which informs what evidence to look for and which statutory element is most likely provable, particularly in jurisdictions where the means element carries the case. Second, victim identification training: the recruiter-relation data is the strongest available corrective to the stranger-abduction model, and officers who expect intimate partners and family members as recruiters identify victims that officers expecting a stranger do not. What it will not do is give you a suspect, a location or a lead – it is deliberately stripped of everything that would.

🔍 Private investigation and corporate security

For corporate investigators and supply-chain due diligence, the sector and control-mechanism data is directly usable. It tells you which labour sectors recur in identified cases and which coercion indicators to build into audit protocols and worker interviews, and it maps cleanly onto the frameworks that regulators use. If you are assessing a supplier in agriculture, construction, fishing, domestic services or hospitality, the base rates here are the honest starting point. Do not use nationality patterns from this dataset in any screening of individuals – that is discriminatory in most jurisdictions and analytically wrong, since the pattern reflects identification systems rather than risk.

📰 Journalism and OSINT media

This is the dataset that lets a journalist write about trafficking with real numbers and without inventing them. The correct story is almost never how many – that figure does not exist here – and almost always what happens: who recruits, how control is maintained, how long it lasts, which sectors. The recruiter-relation and means-of-control fields carry a story that contradicts the dominant public narrative and is well supported. The mandatory discipline is stating in the copy that these are identified victims assisted by particular organisations. A chart from this file captioned as global trafficking is a factual error, and specialists will say so.

🌍 NGO, humanitarian and human rights

For service providers and advocacy organisations this is the shared evidence base and the strongest argument for contributing to it. Operationally, the control-mechanism and recruitment data supports screening tool design and caseworker training, and the sector distributions help target outreach. For advocacy, it supplies mechanism evidence that survives scrutiny. Two cautions. Keep survivor voice in the product – a categorical dataset describes experiences it cannot convey, and survivor-led organisations have been clear about the harm of purely statistical framing. And when working with communities represented in the data, be alert that nationality-level statistics are readily weaponised against migrant communities.

🎓 University and research

The most useful open microdata in the field, and it should be analysed as what it is: administrative case data with a non-random, institution-driven selection mechanism and post-hoc statistical disclosure control. The productive research questions are conditional and mechanistic – co-occurrence of control methods, sector and age structure, recruitment relationship by exploitation type – rather than prevalence. Model the contributor as a fixed effect at minimum. Treat suppression as informative missingness rather than missing at random, since it is applied precisely to distinctive combinations. Cite the release version and download date, because the file is not stable across releases and reproducibility depends on it.

Playbook: working Counter-Trafficking Data Collaborative end to end

A repeatable sequence from first pull to finished product. Each phase states what you are trying to establish, not merely what to click — the objective is a defensible chain of reasoning, not a completed checklist.

Phase 1 — Decide whether you are asking a prevalence question

If your question is how many, stop and use a different source – this file cannot answer it and any answer you construct will be wrong. If your question is what happens, to whom, by what mechanism, in what sector, then this is the best open dataset available. Getting this decision right at the start prevents the single most common failure with the source.

Phase 2 — Download the file and its matching codebook together

Take both, record the download date, and store them as a versioned pair. Column names, code values and disclosure-control parameters have changed between releases, so a file without its codebook is uninterpretable and a comparison across releases without both is unsound.

Phase 3 — Profile contributor composition before anything else

Cross-tabulate the data source field against year, country of exploitation and exploitation type. This single table tells you where every subsequent finding comes from and will usually reveal that an apparent geographic or temporal pattern is really a contributor pattern. Do it first and keep it beside you.

Phase 4 — Expand the concatenated and flag fields properly

Split multi-value strings into normalised attributes and treat exploitation-type flags as non-exclusive. Verify the expansion by checking that per-record flag counts are plausible and that the sector vocabulary matches the codebook exactly. Errors introduced here are invisible downstream because the resulting charts still look reasonable.

Phase 5 — Distinguish suppressed from missing in your model

Encode the two states separately and quantify both. Then check whether suppression correlates with the variables you care about – it almost always will, because it targets distinctive combinations – and state that dependency in your limitations rather than treating the values as missing at random.

Phase 6 — Restrict to a defensible analytic window

Pick a recent multi-year window rather than the full file for any operational profile, and justify the choice. Older cohorts describe a recruitment and control environment that predates the migration of recruitment onto social platforms and of payment onto digital rails, and mixing them flattens a real change.

Phase 7 — Build the mechanism profile

Produce the co-occurrence structure of the means-of-control flags, separately for labour and sexual exploitation and separately by contributor. This is the highest-value output the dataset supports: it tells you which coercion methods travel together, which is what screening tools and prosecutorial strategies need to know and what no prevalence estimate can supply.

Phase 8 — Map to the legal and standards frameworks

Translate the control mechanisms into the ILO indicators of forced labour and the act-means-purpose elements of the Palermo Protocol, and handle the minority-status fields against the rule that the means element is not required for child victims. This mapping is what makes the data usable by lawyers, regulators and supply-chain auditors rather than only by researchers.

Phase 9 — Construct corridors with a minimum cell size

Build citizenship-to-exploitation-country pairs only above a threshold you set and disclose, and remember that a case may involve several countries flattened into one field. Present corridors as identified-case corridors, and check each significant one against migration, labour-recruitment and enforcement sources before treating it as a real route.

Phase 10 — Corroborate before you assert anything about a country

Take any country-level finding to independent sources – national rapporteur reports, the US Trafficking in Persons Report, ILO and UNODC material, national referral mechanism statistics – and ask whether the picture is consistent. Where it is not, the usual explanation is a difference in identification practice, and that difference is itself the finding worth reporting.

Phase 11 — Write the product with the referral pathway attached

Any output that will be read by people who might encounter a victim should carry the relevant reporting route – the national anti-trafficking hotline or referral mechanism, the labour inspectorate, the child protection authority. This dataset exists because organisations wanted better responses, and a product that improves analysis without improving referral has taken from the collaborative without giving anything back.

Phase 12 — State the selection bias in the summary, not the appendix

The first paragraph of your product should say that these are identified victims assisted by specific organisations, that the distribution reflects identification systems, and that no prevalence claim is being made. If a reader can take a number out of your product and present it as a count of trafficking victims, you have written it badly.

The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.

What to pair it with

No single source carries a finding. These are the datasets that corroborate, extend or contradict this one — and a source that contradicts is worth more than one that agrees, because it is the only thing that will tell you when you are wrong.

Source Relationship What it adds
UNODC Global Report on Trafficking in Persons corroborates State-reported detection and conviction data compiled biennially, giving a criminal-justice view of the same phenomenon from a completely different collection mechanism.
US Trafficking in Persons Report extends Country-by-country narrative and tier ranking of government response, which supplies the institutional context that explains why a country's identification numbers look as they do.
ILO Forced Labour prerequisite The indicator framework and legal standards that the means-of-control fields map onto, and the global estimates that supply the population context this dataset deliberately does not.
Global Slavery Index (Walk Free) contradicts A prevalence estimate built by an entirely different method. Where the two point in different directions for a country, the disagreement is informative about identification capacity and should be examined rather than reconciled away.
Polaris Project and the US national hotline prerequisite The contributor behind the US-heavy portion of the data, publishing its own analyses that explain what its rows mean and how its intake works.
IOM prerequisite The largest contributor and the operator of the collaborative, whose programme footprint determines most of the dataset's geographic shape.
National referral mechanism statistics corroborates Where they exist, national identification statistics provide the local denominator that lets you judge how much of a country's caseload this dataset represents.

Legal, ethical and operational constraints

The published file contains no personal data in the ordinary sense – it has been through disclosure control precisely so that it does not – but three legal considerations still apply. First, re-identification. Attempting to link these records to other datasets to identify individuals would in most jurisdictions constitute processing of personal data without a lawful basis, would breach the terms under which the data is published, and would violate the consent basis on which survivors provided the information. Do not do it, and do not build systems that could do it accidentally by joining on rare attribute combinations. Second, the survivor consent chain. The records originate with people who disclosed traumatic experiences to obtain assistance, and the ethical constraint that flows from that outlives the technical anonymisation: use the data to improve responses, not to produce material that sensationalises or stigmatises. Third, downstream harm. Trafficking data is routinely used in immigration policy debates in ways that harm migrants, and nationality-level statistics from this file can be lifted out of context to support restrictive measures. In most jurisdictions you have no legal duty to anticipate that, but you have a professional one. If your work touches identified individuals rather than this dataset, entirely different obligations apply – victim confidentiality, non-refoulement, child protection duties and mandatory reporting – and they take precedence over any analytical objective.

Operational security

Downloading a public research file reveals almost nothing beyond your interest in the subject, and the file is static so there is no query pattern to leak. The exposure sits elsewhere. If you approach a contributing organisation for less aggregated data, you are disclosing your organisation, your research question and often your case context to an assistance provider with its own protection obligations and its own relationships with governments, including possibly the government you are investigating. That is usually fine and occasionally is not; think about it before the email rather than after. Within your own environment, treat any analysis derived from this data with the same handling as case material even though the source file is public, because analytical products that combine corridor patterns with a live case can become sensitive in ways the inputs were not. And be careful about publishing corridor-level findings at high granularity in a context where traffickers or complicit officials read the output – specificity that helps a caseworker can also tell a network which route is now being watched.

Two rules that hold regardless of jurisdiction. Collection that is lawful is not automatically proportionate, and a dataset assembled for one purpose does not carry consent for another. Where the records concern identifiable people, the question is not only whether you may hold the data but whether holding it serves the purpose you are accountable for.

Is it earning its place?

Sources accumulate. Feeds get added during an incident and are never reviewed again, and a decade later the pipeline is carrying dead weight that nobody dares remove. These are the measures that show whether Counter-Trafficking Data Collaborative is contributing anything, and they are worth baselining now so the answer is available later.

  • Proportion of your analyses that are stratified by contributor. If it is not effectively all of them, you are producing findings about programme footprints and labelling them findings about trafficking.
  • Suppression rate in the cells your analysis depends on, reported explicitly. A conclusion resting on cells that are heavily suppressed has an uncertainty you have not stated.
  • Number of country-level findings you have corroborated against an independent source, versus those resting on this file alone. The second set is where your errors will be.
  • Release version currency: how far behind the latest publication your working extract is, and whether any live product is running on a superseded release.
  • Share of products carrying the identified-victims caveat in the body text rather than in a footnote, which is the honest measure of whether the limitation is actually reaching readers.
  • Whether operational outputs – screening tools, training material, audit protocols – derived from this data have been reviewed by practitioners with survivor-contact experience before deployment.
  • Count of downstream uses where a referral pathway was included in the product, as a check that the analysis is feeding the response system rather than only the reporting system.

Beware of volume. Indicator counts rise easily and say almost nothing. Unique contribution — findings this source produced that no other source in your stack would have — is the measure that matters, and it is usually far lower than anyone expects.

Tradecraft notes

The distinctions that separate a competent analyst from a fast one:

  • The data source field is the most important variable in the file. Every geographic, temporal and demographic pattern needs to be checked against it before it is believed, because most of them are contributor effects wearing a substantive disguise.
  • Identification bias is not a caveat, it is the structure of the dataset. Male labour-exploitation victims, victims criminalised on first contact, and victims in sectors no inspector enters are missing in a way that is patterned rather than random, and the pattern differs by country.
  • Suppressed cells carry information. A cell removed by disclosure control is a cell that was distinctive, which means the map of suppression is a rough map of unusual corridors – the opposite of the intuition that suppression indicates nothing was there.
  • Recruiter relationship is the field that changes minds. Recruitment overwhelmingly by known people – partners, family, friends – rather than strangers is the single most policy-relevant finding the dataset supports, and it is also the finding most consistently ignored in public discussion.
  • The means element and the minority-status fields interact legally. For child victims, the means element is not required to establish trafficking in most jurisdictions, so an analysis that filters on means-of-control flags will systematically drop child cases and produce a distorted picture.
  • Duration figures are recollection under duress, not measurement. Treat them as ordinal at best, and never build a model whose conclusions are sensitive to the difference between eight months and twelve.
  • Never publish a corridor at a granularity the publisher suppressed. If your aggregation reconstructs a small group that disclosure control removed, you have defeated the protection regardless of your intention, and you have done it with public data which makes it worse rather than better.
  • Compare this dataset's country picture against the criminal-justice picture deliberately. Where identified victims are numerous and prosecutions are absent, or the reverse, the gap is the story – it usually indicates a system that assists but does not prosecute, or one that prosecutes without identifying.
  • Keep the survivor perspective structurally in the workflow. Categorical data about coercion is easy to analyse in a way that reads as clinical and lands as dehumanising, and the organisations that supplied it have said clearly that this matters to whether survivors keep consenting.

Questions analysts actually ask

How many victims are in the dataset?

Read the current figure off the portal rather than from any secondary description, including this one – the file grows with each release and contributor. More importantly, the count is a count of records contributed, not a count of victims in the world, and quoting it as the latter is the error the publishers most want you to avoid.

Can I identify anyone from this data?

No, and you must not try. The file has been through statistical disclosure control specifically so that no combination of published attributes corresponds to an identifiable individual. Attempting re-identification breaches the terms of use, breaches data protection law in most jurisdictions, and violates the basis on which survivors disclosed.

Why are so many values blank?

Two different reasons that you must keep apart. Some values were never recorded by the assisting organisation. Others were removed by disclosure control because the combination would have been too distinctive. Encode them as separate states, because the second kind is informative and the first kind is not.

Can I compare countries?

Only with heavy qualification. A country's record count is a function of whether a contributing organisation operates there, how well its identification system works, and how many people it assisted – not of how much trafficking occurs. Country comparison from this file is a comparison of identification and assistance capacity, and should be labelled as such.

Is there perpetrator data?

Only the recruiter-relationship flags, which tell you the relationship category and nothing else. There are no names, no networks, no locations and no financial information. For network analysis you need criminal-justice records, investigative reporting and financial intelligence; this file supplies the victim-side base rates that make those investigations interpretable.

How does it differ from the Global Slavery Index?

Fundamentally. This is case-level data on identified victims collected through assistance provision. The Global Slavery Index is a modelled prevalence estimate built from population surveys and extrapolation. They answer different questions, use incompatible methods, and disagreement between them for a country is expected rather than a sign that one is wrong.

Can I use it for supply-chain due diligence?

Yes, for the sector and mechanism dimensions. It tells you which labour sectors recur in identified cases and which coercion indicators to screen for, and those map onto the frameworks regulators use. It cannot tell you anything about a specific supplier, a specific facility or a specific country's risk level in a way that would support a sourcing decision on its own.

How often is it updated?

Periodically and not on a fixed schedule. Treat it as a versioned research dataset: check quarterly, record the release you are using, and re-run any live product when a new release appears rather than assuming continuity, because column sets and code values have changed between versions.

What should I do if my analysis surfaces a suspected live case?

Nothing in this dataset can surface a live case – it contains no identifiers. If separate work brings you into contact with a suspected victim, route it to the national anti-trafficking referral mechanism or hotline, or to the child protection authority for a minor, and do not attempt to investigate or approach. Victim safety and the survivor's own decisions take precedence over the investigation.

Standards, formats and interoperability

What this source speaks natively, and what it has to be translated into before a partner can consume it. Work that arrives in a recognised format is easier to defend, easier to hand over and easier to automate against:

  • The UN Protocol to Prevent, Suppress and Punish Trafficking in Persons (the Palermo Protocol) and its act-means-purpose definitional structure, which the exploitation and control fields are designed to support.
  • ILO indicators of forced labour, onto which the means-of-control flags map almost directly and which is the standard vocabulary for supply-chain and labour-inspection use.
  • ISO 3166 country codes for citizenship and country of exploitation.
  • Statistical disclosure control practice, specifically k-anonymity-style generalisation and suppression, which governs what the published file can and cannot contain.
  • The International Classification of Crime for Statistical Purposes, which is how UNODC-side trafficking statistics are structured and therefore the join point for criminal-justice comparison.
  • Sustainable Development Goal indicator 16.2.2 on trafficking victims, which is the reporting framework several contributors' national counterparts feed.
  • CSV with a published codebook as the actual technical interface, with no API, no schema registry and no version negotiation – plan accordingly.

References

Primary documentation and authoritative references for this source. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.

  1. Counter-Trafficking Data Collaborative — IOM and partners. The hub itself: datasets, visualisations, documentation and the current terms of use. Read the methodology pages before the charts.
  2. CTDC global dataset download — IOM and partners. The download point for the global victim-level file and its codebook. Take both together and record the version.
  3. International Organization for Migration — IOM. The largest contributor and operator. Its counter-trafficking programme documentation explains how the cases behind the rows are generated.
  4. Polaris Project — Polaris. Operator of the US national hotline and a contributor. Its own reporting is the best guide to what hotline-derived records represent and how they differ from case-management records.
  5. National Human Trafficking Hotline — Polaris. The US referral pathway. Include it, or the equivalent national mechanism, in any product that could be read by someone who encounters a victim.
  6. UNODC human trafficking — United Nations Office on Drugs and Crime. The Palermo Protocol framework and the Global Report series, which supply the criminal-justice counterpart to this victim-side data.
  7. Trafficking in Persons Report — US Department of State. Annual country-by-country assessment of government response, the standard reference for the institutional context behind identification numbers.
  8. ILO forced labour and modern slavery — International Labour Organization. The indicator framework the control fields map to, plus the global estimates that answer the prevalence question this dataset cannot.
  9. Walk Free Global Slavery Index — Walk Free. The prevalence estimate built by an incompatible method. Useful precisely because disagreement with the case data is diagnostic.

Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.

Put it into practice

The Quantus Intel threat intelligence platform operationalises this source: it loads the released file against its codebook as a versioned dataset, keeps contributor and suppression state as first-class dimensions, maps control mechanisms onto the ILO indicators, and refuses to let a case-level distribution be presented as prevalence.. Browse the full source catalogue, or follow any tag above into the rest of the library.

Leave a Reply