OpenCorporates: Intelligence Source Guide
OpenCorporates is the largest open database of companies, assembled by scraping and importing official registers across well over a hundred jurisdictions. Its distinguishing feature is not size but provenance: every fact carries the source URL and the date it was retrieved.
OpenCorporates is the largest open database of companies, assembled by scraping and importing official registers across well over a hundred jurisdictions. Its distinguishing feature is not size but provenance: every fact carries the source URL and the date it was retrieved.
At a glance
| Source | OpenCorporates |
|---|---|
| Category | Corporate, Ownership & Legal Records › Corporate Registries |
| Homepage | https://opencorporates.com/ |
| Machine interface | https://api.opencorporates.com/v0.4/companies/search |
| Format | JSON |
| Access | Free registration — API key at no cost |
| Disciplines | Corporate Intelligence |
| Mission domains | Financial Crime, Anti-Money Laundering, Corruption & Governance |
Largest open company registry. — as catalogued in the platform’s own source registry.
OpenCorporates is a UK company that collects company records from official government registers around the world and republishes them in one normalised structure with a consistent identifier scheme, a web interface and a REST API. The unit of record is a company in a jurisdiction, addressed by a jurisdiction code and the company number the register itself uses – so a Delaware corporation and an Irish company are addressed the same way even though the underlying registers share nothing. Around each company sit the attributes the register published: current and previous names, incorporation and dissolution dates, company type, current status, registered address, industry codes, and where the source register discloses them, officers. Alongside companies there are officer records, filings where they have been captured, and control or ownership statements drawn from the beneficial-ownership disclosures a small number of jurisdictions publish. The architectural decision that defines the product is that nothing is asserted without a source: each datum is associated with the register page it came from and a retrieval timestamp, which means you can always answer the question of where this came from and when. It is a copy of official data, and it is honest about being a copy.
The job OpenCorporates does that nothing else does at this price is cross-jurisdictional search. National registers answer questions about their own jurisdiction and refuse everything else; BRIS federates across the EU but will not search by person; commercial providers do this well and charge accordingly. OpenCorporates lets you take a name – of a company or, crucially, of a person – and ask whether it appears anywhere in a hundred-plus registers, which is the first move in almost every corporate investigation that crosses a border. The second job is historical. Registers overwhelmingly publish the present state and quietly overwrite the past; OpenCorporates has been capturing snapshots for well over a decade, so it frequently holds a company's previous name, a former registered address or a director who has since resigned, which the register itself no longer shows. For CORPINT work, and for the money-laundering and corruption domains in particular, that historical layer is often the whole finding: the company was called something else when the contract was signed, and only a scraped archive remembers. The third job is machine access. Most registers offer no API at all, and OpenCorporates provides one uniform interface over all of them.
Who publishes it, and why that matters
OpenCorporates is a private UK company that has always presented itself as mission-driven, arguing publicly for company data as public infrastructure and campaigning for open registers – most visibly during the period when the UK opened Companies House data and when the EU debated beneficial ownership access. That advocacy has been genuine and consequential. It has also had to be funded, and the funding model has changed repeatedly: the API and bulk data have moved between open, free-on-application for public-benefit users, and paid commercial tiers, with the specific terms and thresholds revised more than once. The corporate and commercial arrangements around the company have also changed over its lifetime. The practical consequences for you are three. First, do not rely on any pricing, quota or licence description you remember or read in an old blog post – check the current terms on the site before you plan a project around them. Second, if your use is journalistic, academic or non-profit, ask directly rather than assuming you are excluded, because those categories have historically been accommodated. Third, treat the possibility of terms changing again as a real risk to any pipeline you build, and keep your own copy of what you collect, subject to the licence you obtained it under.
Provenance is the first question to ask of any dataset and the one most often skipped. Who collects it, what their incentive is, whether they publish a methodology, and whether they correct the record when they get something wrong all bear directly on how much weight a finding drawn from it can carry.
What a record actually contains
The fields you will be working with, what each one means, and whether it is something you can pivot on. Read the meanings carefully — more analysis is wrecked by misreading a field than by failing to find one, and a field that looks like an observation is often an inference.
| Field | Type | What it means | Pivot value |
|---|---|---|---|
name |
string | The company name as the source register published it, in the register's own language and script. Not cleaned, not translated, and including or excluding legal-form suffixes according to the register's convention. | Cross-jurisdiction name search; but expect suffix and transliteration variance to defeat exact matching. |
company_number |
string | The identifier used by the source register. It is unique only within its jurisdiction, and formats vary from short integers to long alphanumeric strings with meaningful internal structure. | The register of record, national tax identifiers where derived from it, court and procurement records that quote the number. |
jurisdiction_code |
string | A short code for the register's jurisdiction, using country codes with sub-national extensions for federal systems – so US states, Canadian provinces and similar are addressed separately. Combined with the company number this is the primary key. | Jurisdiction-level statistics, the correct national register, and the applicable company law. |
incorporation_date |
timestamp | Date of incorporation as the register states it. Where a register has migrated systems or re-registered entities, this can post-date actual formation, and some registers do not publish it at all. | Timeline construction, age-based risk scoring, and comparison with dates asserted in contracts and filings. |
dissolution_date |
timestamp | Date of dissolution or removal where published. Absence does not mean the company is alive – many registers do not publish this, and some never remove dead entities. | Insolvency and liquidation records, gazette notices, and the question of who the successor or liquidator is. |
current_status |
string | The register's own status string, passed through largely unaltered – Active, Dissolved, Struck Off, In Liquidation, Good Standing and dozens of national variants. It is not a harmonised enumeration and should not be treated as one. | Jurisdiction-specific interpretation of what the status means legally; insolvency proceedings; whether a legal person still exists. |
company_type |
string | The national legal form as the register expresses it. Determines disclosure obligations and therefore what else should exist for this entity. | Expected filing set at the register; comparability decisions when analysing across jurisdictions. |
registered_address |
string | The address on the register, usually available both as a single string and as parsed components. Frequently a formation agent, an accountant or a registered-agent service rather than a place of business. | Address clustering across the whole dataset, which is one of OpenCorporates' genuinely strong capabilities – and one of its strongest false-positive generators. |
previous_names |
array | Former names with the dates they applied, where the register published them or where a snapshot captured the change. This is one of the highest-value fields in the entire dataset because registers routinely overwrite names. | Historical contracts, sanctions listings under an old name, litigation filed against a predecessor name. |
officers |
array | Directors, secretaries and other officers where the source register discloses them, with position, start and end dates where available. Coverage is entirely determined by the register: some publish full officer data, some publish none. | Officer search across jurisdictions – but note officers are not resolved into global person entities, so the same human appears as many separate records. |
source |
object | The provenance block: publisher, the URL on the official register, and the retrieval timestamp. This is what separates OpenCorporates from an undocumented aggregation and is the field you cite. | The original register page, for verification and for ordering documents. |
retrieved_at |
timestamp | When this record was last taken from the register. The most important field for judging reliability, and the one most often ignored. A record retrieved four years ago describes a company as it was four years ago. | Staleness assessment; deciding whether to re-verify at the register before relying on the record. |
identifiers |
array | Other identifiers associated with the entity where available – tax numbers, LEI, register-specific codes. Coverage is patchy but where present these are the cleanest join keys to other datasets. | GLEIF LEI records, tax and VAT validation services, financial market datasets. |
branch |
enum | Whether the record is a branch registration of a foreign company rather than a domestically incorporated entity. Confusing these is a common and consequential error about who the legal person actually is. | The home jurisdiction record of the parent legal person, which is a different company record entirely. |
Coverage — and what is not in it
Coverage is broad and deeply uneven, and the unevenness is the thing to understand. The dataset spans well over a hundred jurisdictions, with sub-national granularity in federal systems so that individual US states, Canadian provinces and Australian and other regional registers are separate jurisdictions. Depth within a jurisdiction is entirely determined by what the official register publishes and how accessible it is. Jurisdictions with open bulk data – the United Kingdom is the standout – are covered richly, with officers, filings and frequent refreshes. Jurisdictions where the register charges per lookup, requires a captcha, publishes only in an image-based interface or forbids automated access are covered thinly or not at all, and Delaware is the notorious example of a heavily used jurisdiction that publishes very little. Update cadence therefore varies from daily to years, and the retrieved_at field is the only honest statement of it. Historically the archive runs back well over a decade for early-covered jurisdictions, which is where the previous-names and former-officers value comes from. Officer coverage is much narrower than company coverage. Ownership and control data is narrower still, limited to the handful of jurisdictions that publish beneficial ownership openly.
Known blind spots
Absence of evidence here is not evidence of absence. These are the conditions under which OpenCorporates will not show you something that is nevertheless real:
- The dataset is a copy with a date on it, and for many jurisdictions that date is old. A company can be dissolved, renamed or restored without OpenCorporates knowing for months or years, and nothing in the record announces that except the retrieval timestamp that most users never read.
- Officer records are not resolved into people. The same individual appears as a separate officer record for every company and every jurisdiction, with no global person identity, so counting directorships or asserting that two companies share a director requires your own matching and carries your own error rate.
- Where a register does not publish officers, OpenCorporates has none, and an officer search returning nothing is a statement about register transparency rather than about the person. This is exactly the failure mode in the secrecy jurisdictions where you most want an answer.
- Beneficial ownership is largely absent. Control statements exist only for the small number of jurisdictions that publish them openly, and the EU picture became more restrictive after the Court of Justice judgment of November 2022, so the coverage that did exist has in places contracted.
- Financial statements and accounts are not in the dataset in any usable form. Knowing that a company exists tells you nothing about its size, turnover, solvency or whether it trades at all, and a great many registered companies never trade.
- Corporate group structure is not modelled. There is no reliable parent and subsidiary graph, and inferring one from shared addresses, shared officers or similar names produces plausible structures that are frequently wrong in exactly the cases designed to be confusing.
- Coverage bias tracks register openness, not economic importance, so the dataset systematically under-represents the jurisdictions most used for opacity. An analyst who treats it as a world census of companies will draw conclusions about geography that are really conclusions about transparency policy.
- Status strings are the register's own vocabulary and are not comparable across jurisdictions. Good Standing, Active, Registered and Current are not equivalent legal states, and normalising them into a boolean throws away the distinction that matters.
- Non-corporate vehicles are mostly missing. Trusts, most partnerships, foundations and unincorporated associations rarely appear because most registers do not publish them, and these are disproportionately the vehicles used where secrecy is the objective.
Write the blind spot into the product. A statement that something “was not observed in OpenCorporates” is defensible; a statement that it “did not happen” is not, and the difference is what survives cross-examination.
Access, licensing and what you may do with it
Access model: Free registration — an account or API key, at no cost
The web interface is browsable without an account and is genuinely usable for individual lookups, including company, officer and address searches. Programmatic access requires an API key, obtained by registering, and the terms attached to that key depend on who you are and what you are doing: there have historically been free tiers for public-benefit, journalistic and academic use obtained by application, and paid tiers for commercial use. Bulk data has been available separately, historically under an open licence for qualifying users and on commercial terms otherwise. All of these arrangements have been revised more than once. The operational advice is therefore procedural rather than factual: read the current terms on the site, state your use case honestly when you apply, and get the answer in writing before you design around it. If your work is investigative journalism, academic research or non-profit accountability work, apply and explain that – the answer has historically been more generous than the public pricing page suggests, and asking costs nothing.
Licence
OpenCorporates has long distinguished between the underlying facts, which it argues should be open, and the compiled database, which it licenses. Open licensing has historically been applied to much of the company data with attribution and share-alike conditions, which matters enormously if you intend to build a product on top: a share-alike condition can propagate to your derived database. Commercial use, redistribution and bulk reuse have been governed by separate terms. Because these arrangements have changed, the only responsible position is that you must read the licence currently attached to whatever you obtain and keep a copy of it with the data. Two practical rules survive every version of the terms. Attribute, both because it is required and because provenance is the point of the source. And do not assume that because the underlying facts are public records in their home jurisdiction, the compiled database carries the same freedom – the compilation is a separate work with separate rights, and several national registers additionally restrict downstream reuse of their own data.
Rate limits and fair use
Rate limits are attached to the API key tier and have changed over time, so treat any specific number you have seen as unreliable and read the current documentation. The design guidance is stable regardless of the numbers. Cache aggressively, because company records change slowly and re-fetching the same entity daily is wasted quota. Use the company endpoint rather than search when you already have a jurisdiction and number, because search is expensive and imprecise. Where you need to process thousands of entities, ask about bulk data rather than iterating the API – bulk is the intended route for that shape of work and iterating search endpoints at volume is both slow and likely to breach fair use. Set a descriptive user agent identifying your organisation with a contact address. And build backoff into your client from the start, because a key that gets throttled mid-run leaves you with a partially collected dataset that is worse than none.
Licensing changes, and it changes without warning. A dataset that was free for research this year may not be free for commercial or evidential use next year. Confirm the current terms before you build a dependency on it, and record the terms you relied on alongside the data — the licence in force at the time of collection is part of the provenance.
Collecting it
How OpenCorporates is actually pulled, in the order you would set it up. Prefer the bulk or export interface over per-item lookups wherever one exists: it is kinder to the publisher, faster for you, and gives a reproducible snapshot rather than a series of point-in-time answers you cannot reconstruct later.
| Method | Format | Cadence | Notes |
|---|---|---|---|
| Company lookup by jurisdiction and number | JSON | on demand; re-check before relying on a record | The cheapest and most precise call. Use it whenever you already know the jurisdiction code and company number, which after the first pass is most of the time. |
| Company and officer search | JSON | on demand | The discovery route, and the one that justifies the source. Officer search across jurisdictions is the capability no single register offers. Expect to filter results manually because name matching is inherently noisy. |
| Registered address search | JSON|HTML | on demand | Powerful for finding company clusters at a single address, and the fastest way to generate a false network if you do not first establish whether the address belongs to a formation agent. |
| Bulk data | bulk|CSV | periodic snapshots | The correct route for anything involving more than a few thousand entities, for population-level analysis, or for building your own index. Availability and licensing depend on your tier and use case – apply and get it in writing. |
| Control and ownership statements | JSON | as published by the small number of open jurisdictions | Where beneficial ownership disclosures are open, they surface here. Coverage is narrow, so treat a null result as uninformative rather than as evidence of no declared owner. |
| Targeted re-verification at the register | HTML | before any reliance | Not a collection method for the dataset but a necessary companion to it. For any finding that matters, go to the official register the source block names and confirm the current position. |
Ingesting it into the platform
Every step below is idempotent and cursor-based: interrupt one and it resumes from where it stopped rather than duplicating rows or losing progress. Collection is recorded per source, so a feed that quietly stops publishing shows up as a stale timestamp instead of silently thinning your coverage.
- Register the source and its tier — Add OpenCorporates in sources.php with the specific access tier and licence you obtained recorded against it, because what you may lawfully do with the data downstream depends on that and nobody will remember in a year.
- Key on jurisdiction plus company number — In ingest.php, build the primary key from the jurisdiction code and company number together, never from the name. Store the OpenCorporates URL separately as a convenience reference rather than as the identity.
- Carry retrieved_at through as a first-class field — Propagate the retrieval timestamp into the platform record and surface it wherever the record is displayed. A company record without its retrieval date is an undated claim, and treating it as current is the most common misuse of this source.
- Keep previous names as searchable aliases — Index former names alongside current ones so that a search for the name on a 2016 contract finds the company that now trades under a different one. This is where much of the source's unique value sits and it is lost if names are overwritten on update.
- Normalise but do not flatten status — Store the register's raw status string and, separately, your own interpretation with the jurisdiction recorded. Never replace the national term with a boolean, because the legal meaning of dissolved versus struck off versus in liquidation differs and downstream users need it.
- Identify agent addresses before clustering — Count companies per normalised address across your ingested set, flag high-count addresses as probable service providers, and exclude them from link-analysis.php by default. Without this step address clustering produces enormous meaningless components.
- Screen entities and officers against sanctions data — Run resolved names and identifiers through sanctions.php, and screen officer names too – the person is frequently the listed party while the company is not. Record which screening list version was used.
- Cross-reference against the register of record — For entities that reach a decision point in a case, use correlate.php against BRIS or the relevant national register and log the comparison result, so the case file shows the copy was checked against the original rather than relied on alone.
Registered sources and their last-collected state are listed in sources.php, and the scheduled chain that keeps them current is in automation.php.
How it is wrong, and how to tell
Every dataset is wrong in characteristic ways. Knowing which ways is the difference between using a source and being used by one, and it is the part of source evaluation most often skipped because it is the part that takes work.
Quality here is best understood as fidelity to the register at a point in time, and on that measure the source is strong: the scraping and import pipelines are mature, the provenance model is explicit, and errors of transcription are uncommon relative to the volume. What varies enormously is freshness, and freshness is the dominant quality dimension for company data because the underlying facts change constantly. A record refreshed last week from a jurisdiction with open bulk data is close to authoritative; a record retrieved four years ago from a register that has since been redesigned is a historical artefact wearing the same clothes. The retrieved_at field makes this judgeable, which is a design decision worth respecting – most aggregators do not let you make it. The second quality limitation is inherited rather than introduced: OpenCorporates cannot be more accurate than the register, and registers are filing-based, so a fictitious director accepted by a registry appears here as a director. The third is structural: because officers are not resolved into people, any analysis about individuals is your analysis, with your error rate, and should not be attributed to the source.
Characteristic false positives
- Stale records read as current. Nothing in a returned record shouts that it was last checked in 2019; you have to look at retrieved_at, and analysts under time pressure do not. This produces confident statements that a dissolved company is active, which is the single most common error made with this source.
- Officer name matching creates people who do not exist. Two records for John Smith in different jurisdictions may be one person or two hundred, and any tooling that merges on name alone will fabricate a director with an implausible portfolio of companies – which then gets published.
- Address clustering merges unrelated companies. Formation agents and registered-agent services host thousands of companies at one address, and a cluster built without excluding them looks like a corporate network and is a client list.
- Branch registrations are mistaken for separate companies. A foreign company registered as a branch is the same legal person as its parent, and treating it as a subsidiary invents a corporate veil that does not exist – or misses that liability reaches the parent directly.
- Similar names across jurisdictions get merged. Companies with near-identical names in different countries are usually unrelated, and occasionally are deliberately named to be confused with a well-known entity, so name similarity is a lead and never a link.
- Absence of officers is read as absence of directors. The company has directors; the register simply does not publish them. Writing that no directors are recorded without saying that the register does not disclose them misleads the reader about what was checked.
- Status strings are mistranslated across jurisdictions. Good Standing in one US state, Active in another and Registered elsewhere reflect different legal tests, and comparing them as though they were the same category produces false conclusions about compliance.
- Previous names are read as evidence of concealment. Companies rename for entirely ordinary reasons – rebranding, acquisition, group restructuring – and a name change is only suspicious in combination with other facts such as timing relative to litigation or a sanctions listing.
None of these make the source unusable. They make it a source that requires corroboration before an assertion built on it goes into a product, which is true of every source and admitted by few.
Ageing
Company records age at wildly different rates depending on which jurisdiction they came from, and the dataset makes no attempt to hide this. For a well-covered register with open bulk data, records may be days old; for a hard-to-scrape jurisdiction, years. The half-life of the facts themselves is short: registered addresses change, directors resign, companies are struck off for failing to file, and dissolution can happen quietly and without any external announcement. Names change less often but change decisively. Identifiers – jurisdiction code and company number – are effectively permanent and are the only part of the record you should treat as durable. A stale record has no distinguishing appearance whatsoever; it is well formed, plausible and wrong, and the only detector is the retrieval timestamp followed by a check at the register. Operationally, set a staleness threshold appropriate to your risk – a few months for anything decision-bearing, longer for background – and treat any record older than that as requiring re-verification at the source before it can support a conclusion.
What this source feeds
A source is only worth what it lets you conclude. These are the disciplines that collect through it, the mission domains it serves and the data points it yields — every one is a tag, so you can follow any thread from here into the rest of the library.
Collected by these intelligence disciplines
Serves these mission domains
Yields these data points
How each sector uses OpenCorporates
The same dataset is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The records are shared — the constraints, thresholds and outputs are not.
🎖 Military and defence
In defence procurement and supply chain security, the recurring question is whether a supplier is what it says it is and who stands behind it. OpenCorporates answers the first quickly across borders and the second only partially. Its genuine strengths for this work are the ability to search a director's name across jurisdictions, which surfaces the same individuals appearing behind nominally unrelated suppliers, and the historical name archive, which catches a supplier that traded under a different name when it was previously excluded or investigated. Its limits are equally clear: no beneficial ownership for most jurisdictions, no financial data to assess whether a supplier can actually deliver, and thin coverage in exactly the secrecy jurisdictions used to obscure foreign control. Use it as the fast first pass in vendor vetting, then escalate to register documents, ownership data and commercial providers for anything critical.
🕵 National intelligence
For CORPINT this is the everyday tool, and the discipline is in knowing what it can and cannot bear. It is excellent for entity discovery across jurisdictions, for spotting the reuse of directors and addresses across a structure, and for recovering the historical state of a company at the time of a transaction under investigation – which is frequently the whole point in a corruption or sanctions-evasion case. It is not an ownership source and must not be quoted as one. The provenance model is what makes it usable in finished intelligence: every assertion can be attributed to a named official register with a date, which is a much stronger position than citing an aggregator that will not say where its data came from. Build collection around the historical archive specifically – the current state is available from the register, but the state as at a past date usually is not, and that asymmetry is where the analytical value concentrates.
👮 Law enforcement
Investigators use this to map corporate structures fast and cheaply at the intelligence stage, before the point where formal process becomes proportionate. Officer search across jurisdictions frequently identifies the network of companies around a subject that no single register would reveal. Address search identifies the professional enabler – the formation agent or accountant whose address recurs – who is often a more productive line of enquiry than any individual company. What it cannot do is produce evidence: this is a third-party copy of a register, and a court will want the certified extract from the register itself. Treat every OpenCorporates finding as a lead to be confirmed through the register or through process, record the retrieval date, and be careful about the officer-matching step, because merging identically named people is an analytical act performed by you and is disclosable as such.
🔍 Private investigation and corporate security
For due diligence, asset tracing and pre-litigation work the practical value is breadth per unit cost. A subject's name searched across jurisdictions returns directorships that a country-by-country manual search would take weeks to assemble, and previous names catch the company that quietly rebranded after a judgment. The professional discipline is to treat every hit as provisional until confirmed at the register, and to be explicit in client reporting about the difference. Where a report will support a decision to litigate, transact or extend credit, the OpenCorporates record establishes where to look and the register extract establishes what is true. Be careful with negative findings in particular: telling a client that a subject holds no directorships when the relevant registers simply do not publish officers is the kind of error that ends engagements.
📰 Journalism and OSINT media
This is one of the workhorse tools of cross-border investigative journalism, and it has earned that place. Reporters use it to establish that the company in a document is a real registered entity, to find the other companies a named individual is behind, and to recover the name a company was trading under at the time of the events being reported. Two publication disciplines matter. First, cite the register and the retrieval date rather than the aggregator alone, because the register is the authority and the date is what makes the claim falsifiable. Second, be honest with readers about what an officer match does and does not prove – name coincidence is common, and asserting that two companies share a director on the basis of an identical name string is a correction waiting to happen. Verify the individual through date of birth, address or a second document before printing.
🌍 NGO, humanitarian and human rights
Corporate accountability, tax justice and supply chain organisations rely on this source heavily because it is affordable and cross-border, and because its licensing has historically accommodated public-benefit users who should apply rather than assume. It supports the standard workflow of moving from an alleged harm to a named legal person to the wider group around it, and its historical archive is particularly valuable when the entity involved has been restructured since the events at issue. The coverage caveat needs to be stated in published work: thin coverage in secrecy jurisdictions means an absence of findings there is not evidence of absence, and reports that fail to say so invite an easy rebuttal. Pair it with beneficial ownership data where available and with Aleph-style document collections for the material that registers never contained.
🎓 University and research
For research in economics, law, political economy and network science this is one of the few large firm-level datasets available without a commercial licence, and it has supported substantial published work. Its properties as a sample must be handled explicitly: inclusion is a function of register openness, so the dataset is a biased sample of world firms in a way that correlates with governance quality, tax policy and transparency legislation – which are frequently the very variables under study. Any panel construction has to reckon with variable refresh rates across jurisdictions, meaning apparent changes in a firm's attributes are often changes in when it was last scraped. Bulk data rather than the API is the right access route for research, licensing terms should be secured and cited, and the retrieval timestamps should be treated as data rather than metadata.
Playbook: working OpenCorporates end to end
A repeatable sequence from first pull to finished product. Each phase states what you are trying to establish, not merely what to click — the objective is a defensible chain of reasoning, not a completed checklist.
Phase 1 — Decide what question needs a copy and what needs the original
OpenCorporates is for discovery, breadth and history. The register is for the current, authoritative, citable state. Starting an investigation at the register is slow and starting a court filing at the aggregator is negligent. Know which mode you are in at every step and label your findings accordingly.
Phase 2 — Fix the jurisdiction before you search
The jurisdiction code plus company number is the identity. Work out the likely jurisdiction from the legal form, the address or the document you started with, and search within it. Global name searches are useful for discovery and terrible for identification, because the same name exists everywhere.
Phase 3 — Read retrieved_at before you read anything else
Establish how old the record is before forming any view of what it says. A record from a jurisdiction refreshed weekly supports different conclusions from one last touched years ago. Make this the first thing you look at, and record it in your notes alongside the fact you are taking from the record.
Phase 4 — Mine the previous names deliberately
Pull the full name history and compare it against the dates in your source documents. A contract signed with a differently named company, a sanctions listing under a former name, or a rebranding that coincides with litigation are findings that exist only because someone archived the old value. This is the source's least used and highest value feature.
Phase 5 — Run the officer search and then discipline it
Search the individual's name across jurisdictions to build the candidate set, then reduce it using every discriminator available – date of birth where published, address, nationality, co-occurrence with other known entities, and the plausibility of the timeline. Write down which discriminators you used, because the reduction is your analysis and will be challenged.
Phase 6 — Establish whether the address is an agent
Before treating a shared registered address as a relationship, count how many companies sit there and look at what they have in common. Hundreds of unrelated companies means a service provider, which is a lead about the enabler rather than a link between the companies. Three companies with overlapping officers is a real cluster.
Phase 7 — Map the structure only as far as the evidence reaches
Build the entity map from confirmed facts – shared officers with verified identity, ownership statements where they exist, branch relationships where declared. Mark inferred edges as inferred and keep them visually distinct in link-analysis.php. Structures assembled from name similarity and shared addresses look authoritative and are frequently fiction.
Phase 8 — Screen everything before going deeper
Push companies and officers through sanctions and PEP screening early. A hit reshapes the legal posture of the enquiry, changes what you may do and whom you must tell, and finding it after three weeks of network mapping is a wasted three weeks.
Phase 9 — Confirm the load-bearing facts at the register
Take every fact your conclusion actually rests on to the official register named in the source block, and record the confirmation with its own timestamp. This is the step that converts a research finding into something defensible, and it is the step most often skipped under deadline.
Phase 10 — Interrogate the gaps as data
Where the picture is thin, establish whether that reflects the entity or the register. A jurisdiction publishing no officers, no accounts and no ownership tells you something real about the structure's design. Write that up as a finding about opacity rather than leaving a blank space that reads as nothing to see.
Phase 11 — Preserve what you collected under the licence you have
Store the retrieved records with their provenance blocks and the licence terms under which you obtained them. Terms have changed before and will change again, and a case reopened in two years needs the data you had and the basis on which you held it.
Phase 12 — State coverage limits in the finished product
Any report built on this source should say which jurisdictions were searched, that coverage depends on register openness, that officer matching was performed by the analyst, and what the retrieval dates were. Omitting this does not make the report stronger; it makes it easier to attack.
The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.
What to pair it with
No single source carries a finding. These are the datasets that corroborate, extend or contradict this one — and a source that contradicts is worth more than one that agrees, because it is the only thing that will tell you when you are wrong.
| Source | Relationship | What it adds |
|---|---|---|
| EU Business Registers (BRIS) | corroborates | The authoritative live check for EU and EEA entities, returning the official identifier and a link to the register of record – the natural verification step for any European finding here. |
| UK Companies House | extends | Open bulk data, a documented API, officers, charges and full filing history. For UK entities go here directly rather than through an aggregator; the register gives more and is current. |
| OpenSanctions | extends | Screens companies and officers against consolidated sanctions, PEP and watchlist data, which is where a corporate structure turns into a compliance or enforcement finding. |
| OCCRP Aleph | extends | Searches leaked and scraped document collections for the same entity names, frequently supplying the ownership and control information that registers never held. |
| ICIJ Offshore Leaks Database | extends | Offshore entities, intermediaries and officers from the major leak investigations, covering exactly the jurisdictions where register coverage is weakest. |
| GLEIF | corroborates | Independently verified legal entity identifiers with published parent relationships, providing a group-structure signal that company registers do not offer. |
| LittleSis | extends | Adds the interpersonal and institutional relationships around corporate officers – board interlocks, donations, advisory roles – that a register cannot see. |
| CourtListener | extends | US litigation involving the entity, which routinely discloses ownership, contractual relationships and financial detail that no register requires to be filed. |
| OpenOwnership | extends | Beneficial ownership data and the data standard for it, addressing the dimension that is largely missing from register-derived company data. |
Legal, ethical and operational constraints
Two distinct legal questions apply. The first is the licence, which governs what you may do with the compiled database and has changed more than once – read the terms attached to your access, keep them, and remember that share-alike conditions can propagate to a derived database you build and distribute. The second is data protection. Officer records name natural persons and frequently include partial addresses, dates of birth and nationality, and in the EU and UK that is personal data even though the source register publishes it. Holding it in bulk engages your obligations around lawful basis, purpose limitation, retention and subject rights, and the fact that the underlying register is public does not by itself supply the basis. Separately, several national registers impose their own restrictions on downstream reuse of their data, which persist through the aggregator. For investigative use the answers are usually available – legitimate interests, journalistic exemptions, or a statutory function – but they need to be identified in advance and documented, not assumed. And be careful about publishing officer identifications derived from name matching, because misidentifying a private individual as a director of a company under investigation is both a defamation and a data protection problem.
Operational security
API queries are authenticated to your key and therefore attributable to your organisation by definition; the operator can see exactly which companies and which people you looked up and when. For most work this is unremarkable. For sensitive enquiries it is not, and there is no anonymous programmatic route – the choice is between attributable API use and unauthenticated web browsing, which is slower, still logged, and observable at the network layer. Where the subject of an enquiry is sensitive, consider pulling broader slices via bulk data and querying locally, so that the operator sees a general acquisition rather than a targeted interest. Also consider the second-order exposure: some registers publish or notify search activity, so following an aggregator hit through to the register can be more revealing than the original query. And bear in mind that a distinctive pattern of lookups is itself intelligence about your investigation's direction to anyone who obtains those logs, lawfully or otherwise.
Two rules that hold regardless of jurisdiction. Collection that is lawful is not automatically proportionate, and a dataset assembled for one purpose does not carry consent for another. Where the records concern identifiable people, the question is not only whether you may hold the data but whether holding it serves the purpose you are accountable for.
Is it earning its place?
Sources accumulate. Feeds get added during an incident and are never reviewed again, and a decade later the pipeline is carrying dead weight that nobody dares remove. These are the measures that show whether OpenCorporates is contributing anything, and they are worth baselining now so the answer is available later.
- Median and 90th-percentile age of retrieved_at across the records your cases actually rely on, which measures the staleness risk you are carrying rather than the staleness of the dataset overall.
- Proportion of load-bearing findings confirmed at the official register, which should approach one for anything that leaves your organisation and is the clearest indicator of methodological discipline.
- Officer-match precision measured by how often an identification made from name matching survives verification against a second discriminator, tracked as your error rate rather than the source's.
- Number of unique jurisdictions actually queried per investigation versus the number relevant, which exposes the habit of searching only the jurisdictions you are comfortable with.
- Share of investigations where a previous name recovered from the archive changed the analysis, since that is the capability least replaceable by any other source.
- Count of addresses flagged as formation agents and excluded from clustering, as a direct measure of whether the dominant false-positive mechanism is being controlled.
- Rate at which a register check contradicted the aggregated record on a material field, tracked per jurisdiction so you learn which jurisdictions you can trust the copy for.
- Quota consumption per finding, which tends to reveal that expensive search calls are being used where cheap direct lookups would do.
Beware of volume. Indicator counts rise easily and say almost nothing. Unique contribution — findings this source produced that no other source in your stack would have — is the measure that matters, and it is usually far lower than anyone expects.
Tradecraft notes
The distinctions that separate a competent analyst from a fast one:
- The retrieval timestamp is part of the fact. Never quote a company attribute from this source without it, internally or externally, because the same string means different things depending on whether it was captured last week or four years ago.
- Officer records are strings, not people. The dataset does not claim two identically named officers are the same human, and if you make that claim it is yours to defend. Reduce candidate sets with independent discriminators and document the reduction.
- A thin record in a secrecy jurisdiction is a finding. Where the register publishes almost nothing, say so explicitly rather than leaving the section blank, because the absence is a designed property of the structure and belongs in the analysis.
- Search the history, not just the present. The comparative advantage of this source over any register is that it remembers, and analysts who query only current state are using the least distinctive thing about it.
- Establish the register's own transparency profile before interpreting coverage. Whether officers, accounts and ownership are published is a property of the jurisdiction, and knowing it turns every gap from an unknown into a known constraint.
- Treat branch registrations with care, because a branch is the same legal person as its parent. Getting this wrong changes who you sue, who is liable, and where the assets are.
- Cross-check dissolution against a second source before writing that a company is dead. Registers often lag, aggregators lag further, and asserting that a company has ceased to exist when it is trading is a serious error in due diligence work.
- Use bulk data for anything at scale and the API for anything targeted. Iterating search calls over thousands of names is slow, breaches fair use and produces worse results than a local index would.
- Keep the provenance block with the record forever. When a finding is challenged in two years, the register URL and retrieval date are what let you reconstruct exactly what was known and when.
Questions analysts actually ask
Is OpenCorporates free?
The website is browsable without charge and there have historically been free or application-based tiers for public-benefit, journalistic and academic users, alongside paid commercial tiers. The specific arrangements have changed more than once, so check the current terms rather than relying on anything you remember. If your use is non-commercial, apply and explain it rather than assuming you must pay.
Can I use it as evidence in court?
Not directly. It is a third-party copy of a public register, and a court will want the register's own certified extract. Use OpenCorporates to find the entity and understand its history, then obtain the authenticated document from the register named in the source block. The aggregator's value in litigation is often the historical state, which the register may no longer show at all.
Why does it show a company as active when the register says dissolved?
Because the record has not been refreshed since the dissolution. Check retrieved_at, which will usually explain it immediately. This is the most common complaint about the source and it is not really a defect – it is a copy with an honest date on it, and the date is there to be read.
Can I find the beneficial owner?
Usually not. Beneficial ownership is published openly by only a small number of jurisdictions, and the European picture became more restrictive after the Court of Justice judgment of November 2022. Where control statements exist they surface here; where they do not, absence tells you nothing about who actually controls the company, and you should say so in your write-up.
Are the officer records reliable?
They are as reliable as the register that published them, which means they are declarations rather than verified identities. More importantly, they are not resolved into people: the same name across two jurisdictions may be one person or many, and deciding is your analytical work, with your error rate attached.
Why is coverage so poor for some jurisdictions?
Because coverage follows register openness. A register that publishes bulk data is covered richly; one that charges per lookup, blocks automation or publishes only images is covered thinly or not at all. This creates a systematic bias against exactly the secrecy jurisdictions where you most want data, and it should be stated in any analysis of geographic patterns.
Should I use the API or bulk data?
Targeted lookups and interactive investigation, the API. Anything involving thousands of entities, population-level analysis or your own index, bulk data. Iterating the search endpoint over a large list is slow, wasteful of quota and likely to breach fair use, and it produces worse matching than a local index with your own logic.
How should I cite it?
Cite the underlying register and the retrieval date, and mention OpenCorporates as the route. That is both what the provenance model supports and what makes your claim checkable by someone else. Citing the aggregator alone invites the response that you relied on a copy without checking the original.
Can I build a commercial product on this data?
Possibly, but only after reading the current licence carefully. Open licences with share-alike conditions can propagate to your derived database, some national registers restrict downstream reuse of their own data regardless of what the aggregator permits, and the commercial terms have been revised more than once. Get written confirmation before you build.
Standards, formats and interoperability
What this source speaks natively, and what it has to be translated into before a partner can consume it. Work that arrives in a recognised format is easier to defend, easier to hand over and easier to automate against:
- Jurisdiction codes follow country codes with sub-national extensions for federal systems, which is what allows US states and Canadian provinces to be addressed as separate registers rather than collapsed into a country.
- The Legal Entity Identifier maintained through GLEIF appears in identifier data where available and is the cleanest join to financial market datasets and to published parent relationships.
- The Beneficial Ownership Data Standard developed through OpenOwnership is the reference model for control and ownership statements, and is the format to target if you are consolidating ownership data from multiple sources.
- The FollowTheMoney model used by Aleph and OpenSanctions maps company records with jurisdiction, registration number and incorporation date, and is the practical bridge into investigative graph tooling.
- Industry classifications appear as the source register published them, which means SIC, NAICS and NACE codes coexist in the dataset and are not interconverted for you.
- The European Unique Identifier is the parallel official identifier for EU and EEA entities, and holding both it and the OpenCorporates key gives you a clean path to the register of record.
- The platform exports resolved corporate entities and their relationships in STIX 2.1, MISP, CSV, JSON and JSONL, so structures built here move into case and sharing workflows without re-keying.
References
Primary documentation and authoritative references for this source. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.
- OpenCorporates — OpenCorporates Ltd. The service itself. Start with the current terms, coverage pages and the register list, all of which change more often than the data model does.
- OpenCorporates API reference — OpenCorporates Ltd. The authoritative description of endpoints, parameters and response structure. Read it rather than relying on remembered field names, which have shifted between API versions.
- European e-Justice Portal — European Commission. The authoritative live check for EU and EEA entities and the route to the register of record for verification of European findings.
- UK Companies House — Companies House. The best-documented open company register, and the benchmark for what register transparency can look like when a government decides to do it.
- GLEIF — Global Legal Entity Identifier Foundation. The LEI system, including the golden copy files and the parent relationship data that supports group structure work.
- OpenOwnership — OpenOwnership. The Beneficial Ownership Data Standard and current analysis of ownership transparency regimes, covering the dimension registers largely omit.
- OpenSanctions — OpenSanctions. Consolidated sanctions and PEP data with a documented entity model, and the natural screening layer over any corporate dataset.
- ICIJ Offshore Leaks Database — International Consortium of Investigative Journalists. Offshore entities and intermediaries from the major leaks, covering the jurisdictions where register-derived coverage is thinnest.
- OCCRP Aleph — OCCRP. Document and dataset search for the same entities, useful for finding the ownership and contractual material that never reached a register.
- EUR-Lex — Publications Office of the European Union. For the EU legal instruments governing company disclosure and beneficial ownership access, which determine what open company data can contain.
Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.
Put it into practice
The Quantus Intel threat intelligence platform operationalises this source: it keeps the retrieval timestamp and register URL bound to every company attribute, indexes previous names as first-class aliases, and flags formation-agent addresses before they can turn an address search into a fictitious network.. Browse the full source catalogue, or follow any tag above into the rest of the library.