LittleSis: Intelligence Source Guide
LittleSis is a free, editable database of relationships between powerful people, companies, agencies and non-profits, run by the Public Accountability Initiative. Its unit of value is the typed, cited edge: who sits on whose board, who funded whom, who moved from agency to lobbying firm.
LittleSis is a free, editable database of relationships between powerful people, companies, agencies and non-profits, run by the Public Accountability Initiative. Its unit of value is the typed, cited edge: who sits on whose board, who funded whom, who moved from agency to lobbying firm.
At a glance
| Source | LittleSis |
|---|---|
| Category | Corporate, Ownership & Legal Records › Investigative Leak Archives |
| Homepage | https://littlesis.org/ |
| Machine interface | https://littlesis.org/api |
| Format | REST |
| Access | Open — no account required |
| Disciplines | Corporate Intelligence |
| Mission domains | Corruption & Governance, Election Security & PSYOP, Financial Crime |
Free database mapping relationships between powerful people, companies and organisations; open REST API. — as catalogued in the platform’s own source registry.
LittleSis is a relationship graph maintained as a public wiki with an open read API. The two core object types are entities – which are either people or organisations, with subtypes such as business, public company, political candidate, elected representative, lobbyist, government body, school, philanthropy and industry trade association – and relationships, which connect two entities with a typed category. The category set is small and deliberate: position, education, membership, family, donation, transaction, lobbying, social, professional, ownership, hierarchy and a generic catch-all. Each relationship can carry a start date, an end date, a currency amount for financial ties, a description, and most importantly a set of references – citations to the documents or pages that support it. The name is a deliberate inversion of Big Brother: the project's premise is that the surveillance of ordinary people is well developed and the mapping of elite networks is not. Content is contributed by volunteer editors and by project staff, seeded in bulk from public filings – securities disclosures, campaign finance records, non-profit tax returns, board listings – and then extended by hand from reporting and primary documents.
Company registers tell you that a person is a director. Litigation tells you they were sued. LittleSis tells you they went to school with the regulator, sat on a foundation board with the counterparty, and gave to the campaign of the legislator who wrote the exemption. It is a source about affiliation rather than about legal fact, and affiliation is the layer that most institutional datasets deliberately exclude. For CORPINT work on influence, and for the corruption and election domains specifically, that makes it a discovery instrument: you arrive with one name and leave with an interlocking directorate, a revolving-door path, a donor cluster or a think tank that connects two organisations that appeared unrelated. The second job it does uniquely is preserve the human labour of investigative research. Journalists and researchers who assembled a network by hand for a specific story can deposit it where the next person can find it, with citations attached, which is why a LittleSis entity often carries relationships that exist in no structured dataset anywhere and were only ever documented in a single article.
Who publishes it, and why that matters
LittleSis is run by the Public Accountability Initiative, a small non-profit research organisation in the United States, funded by foundation grants and donations. PAI does its own investigative research – on fossil fuel interests, financial institutions, police foundations, university governance and similar – and LittleSis is both the tool that supports that research and a public good in its own right. Two implications follow. First, coverage reflects PAI's research agenda and that of its contributor community, which is progressive, US-centred and focused on corporate and political power; this is not concealed and the project is explicit about its perspective, but it means depth is uneven in a way that correlates with what activists and journalists have chosen to investigate. Second, it is small. There is no service level, the API surface is modest, and continuity depends on grant funding for a handful of staff. The counterweight is that the project is genuinely open – the software and much of the data are published, the edit history is visible, and references are attached to claims – so you can audit what you are relying on in a way that no commercial relationship database permits.
Provenance is the first question to ask of any dataset and the one most often skipped. Who collects it, what their incentive is, whether they publish a methodology, and whether they correct the record when they get something wrong all bear directly on how much weight a finding drawn from it can carry.
What a record actually contains
The fields you will be working with, what each one means, and whether it is something you can pivot on. Read the meanings carefully — more analysis is wrecked by misreading a field than by failing to find one, and a field that looks like an observation is often an inference.
| Field | Type | What it means | Pivot value |
|---|---|---|---|
entity_id |
int | Stable identifier for a person or organisation. This is what you should store, because names in this dataset change as editors correct and disambiguate them. | All relationships attached to the entity, its aliases, its references, and its position in the graph. |
primary_ext |
enum | Whether the entity is a Person or an Org. Determines which further extensions and fields are meaningful and how relationships should be interpreted. | Person entities resolve to individuals in other datasets; Org entities resolve to companies, agencies or non-profits. |
extensions |
array | Subtypes layered on an entity – business, public company, political candidate, elected representative, lobbyist, government body, school, philanthropy, industry trade association and others. An entity can hold several at once. | Sector filtering, and the strongest available signal of what kind of thing you are looking at before you read any relationship. |
name |
string | The display name, plus aliases where editors have recorded them. Person names are stored in a structured form with prefix, first, middle, last and suffix components where known. | Name-based lookup in registers, filings and sanctions data – with the usual caution that a name is not an identity. |
relationship_id |
int | Identifier for a single typed edge between two entities. The edge, not the entity, is the unit of value in this dataset, and it is what carries dates, amounts and references. | The two entities it joins, its references, and the surrounding subgraph. |
category_id |
enum | The relationship type: position, education, membership, family, donation, transaction, lobbying, social, professional, ownership, hierarchy or generic. A small controlled vocabulary, which is exactly what makes the graph analysable. | Filtering a network to the kind of tie that matters – board interlocks are not the same evidence as a single donation. |
start_date |
timestamp | When the relationship began, recorded with partial precision where only a year or a month is known. Frequently absent, and its absence is the most common quality problem in the dataset. | Temporal filtering of a network to the period that matters, which is the difference between a real finding and an anachronism. |
end_date |
timestamp | When the relationship ended. Very often missing even where the relationship has plainly ended, because recording a departure requires someone to notice and edit. | Currency assessment; determining whether a board seat still exists before writing that two organisations are connected. |
is_current |
enum | An explicit flag for whether the relationship is ongoing, which may be true, false or unknown. The unknown case is common and should not be collapsed into either of the others. | Confidence weighting; deciding which edges may be presented as present-tense facts. |
amount |
int | Monetary value for donation and transaction relationships, where an editor or an import recorded it. Usually present for campaign finance imports and often absent for hand-entered ties. | Ranking of financial relationships by materiality; comparison against campaign finance and non-profit filing data. |
description |
string | A free-text label for the relationship – job title, board role, nature of the transaction. Not controlled, so titles vary in wording, but often the most informative single field on the edge. | Role seniority assessment; distinguishing an advisory position from an executive one. |
references |
array | Citations supporting the entity or relationship, typically URLs to filings, articles or official pages. This is the field that determines whether a claim is usable, and edges without it should be treated as unverified. | The underlying primary source, which is what you actually cite and what you should read before relying on the edge. |
updated_at |
timestamp | When the record last changed. In a volunteer-maintained dataset this is a proxy for attention rather than for accuracy – a stale record may be correct and a recently edited one may be wrong. | Prioritising which relationships to re-verify; detecting bursts of editing around a news event. |
summary |
string | A prose description of the entity written by editors, which frequently contains context and sourcing that exists nowhere in the structured fields. | Leads for further research; identification of the specific investigations that produced the entity's relationships. |
Coverage — and what is not in it
Coverage is deep, narrow and purposive. The dataset is overwhelmingly American, concentrated on corporate executives and directors, financial institutions, energy companies, political donors and candidates, federal and state officials, lobbying firms, think tanks, foundations, universities and industry associations. Within those areas it can be extraordinarily rich – a major bank or fossil fuel company will have hundreds of relationships with dates and references, mapping boards, philanthropy, lobbying and revolving-door movements over years. Outside them it thins fast: small and mid-sized businesses, most sectors of the economy, and almost all non-US entities are absent or represented by a bare stub. Temporally, seeded data from public filings can reach back a decade or more, while hand-entered material clusters around the periods of active investigation. The update rhythm is bursty rather than steady: an entity gets attention when someone is researching it, acquires a batch of new relationships, and then may sit untouched for years. That pattern means recency of editing tells you about researcher interest, not about the currency of the underlying facts.
Known blind spots
Absence of evidence here is not evidence of absence. These are the conditions under which LittleSis will not show you something that is nevertheless real:
- The graph records what somebody chose to research. An entity with no relationships is an entity nobody has worked on, and reading that as evidence of an unconnected actor inverts the meaning of the data completely.
- End dates are systematically under-recorded. Board seats, executive roles and memberships are entered when they begin and frequently never closed out, so the network as displayed contains a large stock of relationships that have quietly ended – which makes any present-tense claim risky without checking.
- Coverage outside the United States is thin to the point of unusability for most purposes. Foreign subsidiaries, non-US executives and international institutions appear mainly where they intersect with American investigations.
- There is no verification workflow in the sense a compliance function would recognise. Edits come from volunteers and staff, references are encouraged rather than enforced, and an unreferenced relationship carries no more evidentiary weight than an assertion by an anonymous contributor.
- Relationship categories flatten important distinctions. Position covers everything from chief executive to unpaid advisory board member, so an edge in the graph tells you a tie exists without telling you whether it conveys control, income or nothing at all.
- Financial amounts are present only where an import or an editor captured them, so the absence of a figure on a donation or transaction edge says nothing about its size and should never be treated as small.
- The project has an explicit orientation towards accountability research on corporate and political power, which shapes which networks get built. Actors of comparable significance may be absent because they were not of interest to the contributor community.
- Person entities are not linked to authoritative identifiers, so distinguishing two people with the same name relies on editor diligence, and merged or duplicated person entities do occur in a dataset this size.
- Sourcing decays. References are URLs, and a substantial share of any decade-old citation set now points at dead pages, which means the audit trail behind an old relationship may no longer be followable.
Write the blind spot into the product. A statement that something “was not observed in LittleSis” is defensible; a statement that it “did not happen” is not, and the difference is what survives cross-examination.
Access, licensing and what you may do with it
Access model: Open — no account required
The website is fully open for browsing and searching, and entity pages display relationships grouped by category with their references visible, which is the fastest way to understand what the data actually looks like before writing any code. There is a documented read API returning JSON, requiring no key, covering entity lookup, entity search and the relationships attached to an entity. The project has also historically published bulk database exports, which is the right route for network analysis at any scale and avoids repeated crawling of a small non-profit's infrastructure. Editing requires an account and is subject to community norms; if you are going to consume the data heavily it is worth contributing corrections back, particularly end dates, since that is the dataset's weakest field and the one you will most often find yourself researching anyway. The visual network mapping tool the project maintains is useful for producing publication-quality diagrams from the underlying graph, and works from the same data.
Licence
The project publishes its content under an open licence in the Creative Commons attribution and share-alike family, and the application code is open source. Confirm the exact terms on the site before you build on them, because the specific licence version and its application to bulk exports versus page content are details you should read rather than assume. Two practical consequences hold under any version of these terms. Attribution is required, and it is also good practice for a different reason: a claim sourced to a volunteer-edited wiki should be presented as such rather than laundered into an unattributed assertion in your report. And share-alike conditions can propagate to derived databases, so if you merge LittleSis relationships into a graph you intend to distribute, the licence status of that graph needs to be worked out in advance. Separately, the references attached to relationships point at third-party material with its own rights; the citation is open, the cited article is not.
Rate limits and fair use
No formal published quota should be assumed, and the correct posture is to treat this as a small non-profit's infrastructure rather than as a service. If you need more than a few hundred entities, use the bulk exports rather than walking the API entity by entity – crawling a graph through per-entity relationship calls generates a large number of requests very quickly and is the classic way a well-meaning researcher becomes a problem. Where you must use the API, cache aggressively, respect the fact that entities change slowly, set a descriptive user agent identifying your organisation with a contact, and introduce delays between requests. If your project genuinely needs continuous access at volume, contact the organisation and explain what you are doing; a small project would generally rather help a legitimate researcher than block an unexplained crawler, and the conversation frequently produces better data access than you would have engineered.
Licensing changes, and it changes without warning. A dataset that was free for research this year may not be free for commercial or evidential use next year. Confirm the current terms before you build a dependency on it, and record the terms you relied on alongside the data — the licence in force at the time of collection is part of the provenance.
Collecting it
How LittleSis is actually pulled, in the order you would set it up. Prefer the bulk or export interface over per-item lookups wherever one exists: it is kinder to the publisher, faster for you, and gives a reproducible snapshot rather than a series of point-in-time answers you cannot reconstruct later.
| Method | Format | Cadence | Notes |
|---|---|---|---|
| Entity search | JSON | on demand | Name-based lookup returning candidate entities. Expect ambiguity on common names and use the entity type and extensions to disambiguate before pulling relationships. |
| Entity detail | JSON | on demand; entities change slowly | Full record for a single entity including its extensions, aliases and summary. Cache these; there is rarely a reason to re-fetch an entity within the same week. |
| Relationships for an entity | JSON | on demand | The core call. Returns the typed edges with dates, amounts and references. This is where the analytical value is and where you should spend your attention rather than on entity metadata. |
| Bulk database export | bulk|CSV | periodic | The correct route for any network analysis, centrality measurement or systematic study. Also gives you a fixed snapshot you can cite, which crawling never does. |
| Reference resolution | HTML | per relationship you intend to rely on | Not a collection method for the dataset but a necessary companion to it: fetch and read the cited source behind any edge that will carry weight in your product. |
| Web page capture | HTML | at the point of use | Capture the entity page as you found it, with the date. The wiki changes, and a claim you relied on in March may have been edited by June. |
Ingesting it into the platform
Every step below is idempotent and cursor-based: interrupt one and it resumes from where it stopped rather than duplicating rows or losing progress. Collection is recorded per source, so a feed that quietly stops publishing shows up as a stale timestamp instead of silently thinning your coverage.
- Register the source with its provenance model — Add LittleSis in sources.php marked as a community-edited secondary source. That classification should follow the data everywhere, because it determines how any downstream product may present a relationship drawn from it.
- Key on entity identifier, never on name — In ingest.php store the numeric entity identifier as the join key. Names in a wiki are corrected, disambiguated and merged over time, and a pipeline keyed on names will silently lose or duplicate entities.
- Preserve category, dates and currency flag on every edge — Import relationships as typed edges retaining the category, start and end dates, and the current flag with its unknown state intact. An edge stripped of its type and dates is worse than no edge, because it looks like a fact.
- Import references as required provenance — Attach the citation URLs to the relationship record and mark edges with no references distinctly. Analysts should be able to filter to referenced edges only, and by default they should be.
- Resolve organisations against registry data — Run resolve-everything.php to match Org entities to corporate registry records, storing the resolution and its confidence separately from the LittleSis name so that the two can be compared rather than conflated.
- Screen people and organisations — Push person and organisation names through sanctions.php for sanctions and PEP screening. Influence networks and sanctions exposure intersect more often than either dataset suggests on its own.
- Load the graph with edge weights that reflect evidence — In link-analysis.php, weight edges by whether they are referenced, whether they are dated, and whether they are current. An unweighted import makes an unreferenced social tie look identical to a documented board seat.
- Attach to the case with the source characterised — When relationships enter a case in cases.php, record that they came from a community-edited source, note the capture date, and flag which ones have been independently verified against the underlying reference.
Registered sources and their last-collected state are listed in sources.php, and the scheduled chain that keeps them current is in automation.php.
How it is wrong, and how to tell
Every dataset is wrong in characteristic ways. Knowing which ways is the difference between using a source and being used by one, and it is the part of source evaluation most often skipped because it is the part that takes work.
This is a secondary source assembled by volunteers and staff, and it should be judged as such – not dismissed, because much of it is carefully researched and cited, but never treated as authoritative in the way a register or a court filing is. The reliable core consists of relationships imported from public filings, which inherit the accuracy of those filings, and hand-entered relationships carrying specific references to primary documents. The unreliable margin consists of unreferenced edges, undated edges and edges whose currency is unknown, and that margin is not small. The single largest systematic weakness is the missing end date: because entering a departure requires an editor to notice it, the graph accumulates relationships that have ended and still present as live, which biases every network measure towards over-connection. The right posture is to use the graph for discovery with enthusiasm and for assertion with discipline: follow its edges to find what to investigate, then verify anything load-bearing against the reference it cites or against an independent primary source, and cite that rather than the wiki.
Characteristic false positives
- Ended relationships shown as current. The dominant failure mode. A director who left in 2016 remains connected to the company in the graph, and a network built without date filtering describes an organisation that has not existed for years.
- Position edges conflated across wildly different roles. Chief executive, non-executive director, unpaid advisory board member and former intern can all appear as a position relationship, so treating edge count as a measure of influence overstates ties that convey nothing.
- Name-based entity duplication and conflation. Two people with the same name may have been merged, or one person may exist as two entities, and neither error announces itself. Check the summary and references before treating a person entity as a single individual.
- Unreferenced edges read as documented. The interface shows relationships whether or not they carry citations, and an analyst scanning a network page will not naturally distinguish them, so a claim traceable to nothing can propagate into a report looking exactly like one traceable to a filing.
- Missing amounts read as small amounts. A donation edge with no figure means nobody recorded a figure, not that the donation was minor, and ranking relationships by recorded amount systematically demotes exactly the ties that were hand-entered from reporting.
- Hub entities distort network metrics. Large institutions – major banks, universities, trade associations – accumulate hundreds of edges and dominate any centrality calculation, connecting everyone to everyone at short path lengths in a way that means very little.
- Coverage bias read as a finding. A dense cluster around one industry and a sparse one around another usually reflects who has been researched rather than a real difference in interconnection, and comparative claims across sectors are the most common misuse of this dataset.
- Dead references treated as verification. An edge with three citations that all now return errors is an edge with no auditable basis, and the presence of reference count in the interface makes that easy to miss.
None of these make the source unusable. They make it a source that requires corroboration before an assertion built on it goes into a product, which is true of every source and admitted by few.
Ageing
Ageing here works differently from a machine-collected source, because the record does not decay – attention does. A relationship entered in 2015 with a start date and no end date stays exactly as it was while the world moves on, and the graph therefore drifts steadily away from reality in one specific direction: towards showing too many connections. Person entities age as people change roles; organisation entities age as companies merge, rename or dissolve; and reference URLs age fastest of all, so that the audit trail behind an old edge frequently no longer resolves. A stale record is invisible on inspection – it is well formed and plausible – and the two available detectors are the presence or absence of an end date and the last-updated timestamp, neither of which is conclusive. Operationally, apply an explicit temporal filter to any network you build, treat an edge with no end date and no recent update as a hypothesis about the present rather than a statement about it, and re-verify against a current source before any published claim in the present tense.
What this source feeds
A source is only worth what it lets you conclude. These are the disciplines that collect through it, the mission domains it serves and the data points it yields — every one is a tag, so you can follow any thread from here into the rest of the library.
Collected by these intelligence disciplines
Serves these mission domains
Yields these data points
How each sector uses LittleSis
The same dataset is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The records are shared — the constraints, thresholds and outputs are not.
🎖 Military and defence
The application here is narrow but real: understanding the influence environment around defence contractors, and mapping the revolving door between the department, the services and industry. Retired officers taking board seats and consulting roles with suppliers is a documented pattern and precisely the kind of relationship the dataset is built to record, with dates and references. For supply chain and foreign influence questions it also helps identify the think tanks, trade associations and philanthropic vehicles through which a contractor engages with policy. The limitations are severe for anything operational: no foreign coverage to speak of, no classified or restricted material, and a US-centric, accountability-oriented perspective. Use it to characterise the civil-industrial influence picture around a programme or a supplier, and never as an authoritative record of a named individual's affiliations.
🕵 National intelligence
For CORPINT and for work on corruption and elite networks, this is a discovery source with an unusually good property: its claims come with citations, so a lead can be traced to a primary document rather than trusted. It fills the layer that registers and filings deliberately omit – the informal, social and philanthropic ties through which influence actually moves – and it preserves network research that would otherwise exist only inside a single published article. The analytical discipline is to treat every edge as a hypothesis with an attached source, to filter by date before drawing any picture, and never to let a wiki-sourced relationship enter finished intelligence without the underlying reference being read. Its US concentration means it is a domestic influence tool rather than a foreign one, and analysts working non-US targets will find it useful mainly where those targets intersect with American institutions.
👮 Law enforcement
In corruption, procurement fraud and public integrity investigations the value is in identifying non-obvious relationships early: the shared board, the foundation that connects a contractor to an official, the donation that precedes a decision. Because the source is open and volunteer-maintained, it produces leads rather than evidence, and every lead needs independent development through the primary records the reference points at – campaign finance filings, non-profit returns, corporate disclosures. Two cautions specific to this work. The dataset is a public wiki with a visible edit history, so an investigator making a large number of queries or edits around a subject may be observable. And an unreferenced relationship about a named individual is an allegation from an anonymous source; treating it as intelligence without corroboration is a route to a wrong and potentially defamatory conclusion.
🔍 Private investigation and corporate security
For due diligence and reputational work this is the fastest free way to understand who a subject is connected to beyond their formal directorships. Board interlocks, philanthropic affiliations, industry association roles and political giving all appear in one place with citations, which would otherwise require assembling from four separate datasets. The professional constraint is that a client report is a product you are accountable for, and a community-edited source cannot be the last word on a named individual. Use it to generate the list of relationships to check, verify each one against the underlying filing or article, and cite the primary source in the report. Be particularly careful with undated position edges, since telling a client that a subject currently sits on a board they left in 2018 is exactly the error that costs an engagement.
📰 Journalism and OSINT media
Reporters have used this heavily for a decade and it is one of the better free tools for the specific task of finding the connection you did not know to look for. Its citation model means an edge usually leads somewhere publishable, and the network mapping tool produces graphics that can go into a story. The publication rules are firm: never cite LittleSis as the source of a fact, cite the document it points to, and check that document yourself. Also apply a date filter before publishing any network diagram, because the most common correction in this genre is describing a person as currently affiliated with an organisation they left years ago. Consider contributing back what your reporting establishes – the dataset exists because journalists and researchers deposited their working networks, and it degrades if everyone only takes.
🌍 NGO, humanitarian and human rights
Accountability, environmental and community organisations are the core constituency, and much of the data reflects exactly the kind of research these organisations do – mapping the interests behind a policy position, the funders behind a campaign, or the interlocking boards of institutions in a sector. It is free, editable and open-licensed, which makes it a genuine commons rather than a product, and depositing your own research into it multiplies its value to everyone working the same issue. The discipline is the same as for any advocacy use of open data: a claim in a report will be attacked at its weakest link, and an unreferenced wiki edge is the weakest link available. Verify, cite the primary document, apply date filters, and be honest in published work that coverage reflects research attention rather than a systematic survey.
🎓 University and research
For research on elite networks, corporate governance, interlocking directorates and the political economy of influence, this is one of very few open relational datasets with typed, dated and cited edges. It has supported published work in sociology, political science and network analysis. The sampling properties must be stated explicitly in any paper: entities enter the dataset through the research interest of activists and journalists, so the network is not a sample of any defined population, and comparative structural claims across sectors are confounded by attention. Missing end dates bias connectivity upward, and hub entities dominate centrality measures. Use bulk exports and report the snapshot date for reproducibility, treat reference presence as a data quality variable, and consider that the dataset is also an interesting object of study in itself as an instance of collaborative investigative infrastructure.
Playbook: working LittleSis end to end
A repeatable sequence from first pull to finished product. Each phase states what you are trying to establish, not merely what to click — the objective is a defensible chain of reasoning, not a completed checklist.
Phase 1 — Decide that you are doing discovery, not verification
Fix the mode before you start. This source is excellent at telling you what to look at and poor at telling you what is true. Every workflow that goes wrong with it begins with an analyst treating a displayed relationship as an established fact rather than as a pointer to a document they have not yet read.
Phase 2 — Resolve your subject to the right entity
Search by name, then confirm using the entity type, the extensions, the summary text and the existing relationships that you have the right person or organisation. Common names produce multiple candidates and occasionally a conflated entity holding two people's relationships, which is the worst starting point available.
Phase 3 — Read the relationship categories before the relationships
Look at what kinds of edge exist for this entity – positions, donations, memberships, lobbying – because the distribution tells you what has been researched and what has not. An entity with fifty position edges and no financial edges has been mapped structurally and not financially, and that gap is your next task.
Phase 4 — Apply a date filter immediately
Before you look at the network, restrict it to the period that matters for your question. Undated edges and edges with no end date should be set aside into a separate bucket rather than included by default, because they are the mechanism by which this graph over-connects.
Phase 5 — Separate referenced from unreferenced edges
Split the relationships into those carrying citations and those that do not. The referenced set is your working material. The unreferenced set is a list of things somebody believes, which may be worth investigating and must never be reported.
Phase 6 — Read the references, do not count them
Open the cited documents for every edge that will carry weight. Check that the source says what the edge claims, that it is still live, and that it is a primary record rather than another aggregator. This is the step that converts a wiki claim into a citable finding and it is the step most often skipped.
Phase 7 — Expand one hop at a time and stop at hubs
Traverse outward from your subject deliberately, and stop expanding through very high-degree entities – major banks, large universities, big trade associations – because everything connects to everything through them and the resulting network says nothing. Note the hub, do not traverse it.
Phase 8 — Reconcile organisations against the register
Take every organisation that matters out to corporate registry data and establish which legal person it actually is. Wiki entities are names and concepts; registry entries are legal persons, and the mapping between them is frequently one to many.
Phase 9 — Look for the revolving door explicitly
Query for position edges that move a person between a government body and a regulated entity, and check the dates on both sides. This is the single most productive pattern in the dataset and the one it is best equipped to record, because the transitions are usually documented and datable.
Phase 10 — Screen the resulting population
Run the people and organisations you have collected through sanctions and PEP screening. Influence mapping and compliance exposure overlap more than either activity assumes, and finding the overlap late in a project is expensive.
Phase 11 — Capture what you used, when you used it
Save the entity and relationship pages as you found them with the retrieval date. This is a live wiki with an edit history; a claim you relied on may be corrected, expanded or removed, and your record needs to show what the data said at the time you acted on it.
Phase 12 — Give something back and state your sourcing
Where your own research established an end date, a correction or a new documented relationship, contribute it. In your finished product, cite the primary documents rather than the wiki, and say that a community-edited relationship database was used for discovery. Both practices make the work more defensible, not less.
The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.
What to pair it with
No single source carries a finding. These are the datasets that corroborate, extend or contradict this one — and a source that contradicts is worth more than one that agrees, because it is the only thing that will tell you when you are wrong.
| Source | Relationship | What it adds |
|---|---|---|
| OpenCorporates | prerequisite | Resolves organisation entities to registered legal persons with jurisdictions and identifiers, converting a wiki concept into something you can research in official records. |
| SEC EDGAR | corroborates | Proxy statements and beneficial ownership filings are the authoritative source for public company directorships and executive compensation, and are what most position edges should trace back to. |
| ProPublica Nonprofit Explorer | corroborates | Non-profit tax returns disclose officers, directors, key employees and grants, which is the primary record behind foundation and think tank relationships. |
| CourtListener | extends | Litigation involving the same people and organisations, which frequently documents relationships under oath that a wiki records as an inference. |
| OpenSanctions | extends | Screens the mapped population against sanctions and PEP data, which turns an influence network into a compliance-relevant picture. |
| OCCRP Aleph | extends | Adds document-level evidence and international coverage for entities that appear in the wiki only as US-facing stubs. |
| Wikidata | corroborates | Independent, differently sourced structured claims about the same people and organisations, with identifiers that help disambiguate entities with common names. |
| Public Accountability Initiative | prerequisite | The operating organisation and its published investigations, which explain why particular networks in the dataset are dense and others are empty. |
| GLEIF | corroborates | Verified legal entity identifiers and published parent relationships, useful for pinning organisation entities to the right corporate group. |
Legal, ethical and operational constraints
The dataset is almost entirely about identifiable natural persons and their affiliations, which makes it personal data under most data protection regimes even though it is published and even though much of it derives from official filings. If you are collecting or holding it from a jurisdiction with a comprehensive data protection law, you need a lawful basis, a purpose, a retention limit and an answer on subject rights, and the fact that the source is a public wiki does not supply any of those. The second exposure is defamation and accuracy. Relationships here are assertions by volunteer editors, and an unreferenced claim that a named individual is connected to a controversial organisation is exactly the kind of statement that generates liability if you republish it as fact. Verify against the primary source and cite that. Third, the content licence has attribution and probable share-alike conditions that propagate to derived works. Fourth, if you use the data to inform decisions about individuals – employment, credit, tenancy – you may fall within regulation designed for consumer reporting, in which case an unverified community dataset is a poor foundation. And ethically, mapping private individuals who merely appear near a powerful network is not what the resource is for; keep the focus on public figures and institutional roles.
Operational security
Read access is anonymous in the sense that no key is required, but requests are still logged with source address and timing, and a concentrated pattern of lookups on one entity is legible to anyone with access to those logs. Editing is far more exposing: contributions are attributed to an account and the edit history is public and permanent, so an edit that corrects or extends a subject's record announces both your interest and, over time, your research direction to anyone watching that page. Some entities in this dataset are actively monitored by the people and organisations they describe. If your interest is sensitive, use bulk data and query locally so no per-entity request is made, browse over a neutral network path, and do not edit. If you do want to contribute – which is good practice – consider doing it after your work is published rather than during it, and separate the contributing account from any identity that links to a live investigation.
Two rules that hold regardless of jurisdiction. Collection that is lawful is not automatically proportionate, and a dataset assembled for one purpose does not carry consent for another. Where the records concern identifiable people, the question is not only whether you may hold the data but whether holding it serves the purpose you are accountable for.
Is it earning its place?
Sources accumulate. Feeds get added during an incident and are never reviewed again, and a decade later the pipeline is carrying dead weight that nobody dares remove. These are the measures that show whether LittleSis is contributing anything, and they are worth baselining now so the answer is available later.
- Proportion of edges you rely on that carry at least one live reference, checked rather than counted, which is the clearest single indicator of whether your use of the source is disciplined.
- Share of relationships in your working set with both a start and an end date, since undated edges are the mechanism by which this graph misleads and tracking them quantifies the risk you are carrying.
- Verification yield: how often reading the cited reference confirms the edge as stated, tracked over time to calibrate how much trust the dataset earns in the areas you actually work in.
- Number of leads that originated here and were subsequently confirmed in a primary source, which is the honest measure of the source's contribution rather than a count of edges imported.
- Ratio of hub-mediated to direct connections in networks you build, because a network whose paths run mostly through large institutions is describing membership of the economy rather than a relationship.
- Coverage of your target population, measured as the share of entities of interest that exist in the dataset with more than a stub, which tells you whether this source is relevant to your beat at all.
- Corrections contributed back per quarter, which measures whether your organisation is maintaining the commons it depends on or simply extracting from it.
Beware of volume. Indicator counts rise easily and say almost nothing. Unique contribution — findings this source produced that no other source in your stack would have — is the measure that matters, and it is usually far lower than anyone expects.
Tradecraft notes
The distinctions that separate a competent analyst from a fast one:
- Treat every edge as a citation, not a fact. The valuable thing this dataset gives you is a pointer to a document; the relationship label is a summary of that document written by someone you do not know.
- The absence of an end date is not evidence that a relationship continues. It is evidence that nobody recorded an ending, and in a volunteer dataset that is the default state rather than a meaningful signal.
- Filter by date before you draw anything. A network diagram without a temporal filter mixes relationships from different decades into a single picture that never existed at any moment in time.
- Read the relationship description, not just the category. Position covers everything from chief executive to honorary adviser, and the free-text description is usually the only place the actual role appears.
- Do not traverse hubs. Large banks, universities and trade associations connect everyone to everyone, and a path running through one is not a relationship – note the hub, record it as context, and stop.
- Check whether an entity is one person. Conflated and duplicated person entities exist, and inheriting a conflation into your analysis produces an individual with an implausible portfolio of affiliations that will be picked apart immediately.
- Recency of editing is a measure of attention, not accuracy. A record edited last week may be wrong and one untouched for five years may be right; the last-updated field tells you who has been researching, not what is true.
- Cite the primary document in your product, always. Naming a community-edited wiki as the source of a claim about a named individual invites the entire finding to be dismissed, and the underlying filing usually says it better anyway.
- Contribute end dates when you find them. It is the field the dataset most needs, you will already have researched it, and it improves the resource for everyone including your future self.
Questions analysts actually ask
Can I trust what LittleSis says?
Trust the references, not the edges. Relationships imported from public filings and those carrying citations to primary documents are generally sound; unreferenced edges are assertions by volunteer editors. The right workflow is to use the graph to find what to check and then check it, citing the underlying document in whatever you produce.
Why does it show someone on a board they left years ago?
Because recording a departure requires an editor to notice and act, while recording an appointment happens naturally when news reports it. Missing end dates are the dataset's most systematic weakness, and any network you build should treat an undated, unclosed relationship as a hypothesis about the present rather than a statement of it.
Does it cover companies outside the United States?
Barely. Coverage is overwhelmingly American, with foreign entities appearing mainly where they intersect with US investigations. For international influence mapping you will need other sources, and an absence here tells you nothing at all about a non-US entity.
Is there an API and does it need a key?
There is a documented read API returning JSON for entity lookup, search and relationships, and it does not require a key. For anything beyond a few hundred entities use the bulk exports instead, because crawling a relationship graph entity by entity generates far more requests than a small non-profit's infrastructure should have to absorb.
How is this different from a corporate registry?
A registry records legal facts about a company – who is formally a director, where it is registered. This records affiliations, including informal ones a registry never captures: philanthropic boards, think tank memberships, political donations, professional and social ties. They answer different questions and the registry is authoritative where they overlap.
Can I use it in a compliance or due diligence report?
As a discovery layer, yes. As a source of assertions about named individuals, no – verify each relevant relationship against the primary document and cite that. A regulated report resting on unverified wiki edges is difficult to defend, and stating that someone currently holds a role they left is the specific error most likely to occur.
Should I contribute to it?
If you use it, yes, and end dates are the most valuable thing you can add because they are the weakest field. Bear in mind that edits are public and attributed, so if your research is sensitive, contribute after publication rather than during, and keep the contributing identity separate from a live investigation.
Why do some major figures have almost no relationships?
Because nobody has researched them. Entities enter this dataset through the interest of journalists, activists and staff researchers, so a sparse record reflects attention rather than isolation. Never present an empty or thin entity as evidence that a person or organisation is unconnected.
Can I use the data commercially?
Read the current licence first. The content is published under an open Creative Commons style licence with attribution and probable share-alike conditions, which can propagate to a derived database you distribute. Attribution is required in any case, and it is also the honest way to present a claim that originated in a volunteer-edited source.
Standards, formats and interoperability
What this source speaks natively, and what it has to be translated into before a partner can consume it. Work that arrives in a recognised format is easier to defend, easier to hand over and easier to automate against:
- The relationship category vocabulary – position, education, membership, family, donation, transaction, lobbying, social, professional, ownership, hierarchy and generic – is a small controlled ontology and is what makes the graph analysable rather than merely browsable.
- Entity extensions function as a type system over people and organisations, distinguishing public companies, government bodies, political candidates, lobbyists, schools and philanthropies, and should be preserved through any import.
- Partial date precision is used throughout, allowing a year or a year and month to be recorded where the exact day is unknown, and flattening these into full dates fabricates precision that the source deliberately avoided.
- The FollowTheMoney model used by Aleph and OpenSanctions expresses directorships, ownership, membership and family ties as first-class entities, and is the natural target schema when merging this graph with document-derived investigative data.
- Campaign finance and non-profit tax filing formats are the upstream standards behind much of the seeded financial data, and understanding them is how you validate an imported donation edge.
- Creative Commons attribution and share-alike licensing governs reuse, with the practical consequence that derived databases you distribute may inherit the share-alike obligation.
- The platform exports entities and relationships in STIX 2.1, MISP, CSV, JSON and JSONL, so a verified influence network travels into the same case and sharing structures as any other relationship data.
References
Primary documentation and authoritative references for this source. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.
- LittleSis — Public Accountability Initiative. The database itself. Browse a few well-developed entities before writing any code – the interface makes the data model and its weak points immediately visible.
- LittleSis API — Public Accountability Initiative. The documented read interface for entity lookup, search and relationships. Read it for the current endpoint shapes rather than relying on remembered field names.
- LittleSis bulk data — Public Accountability Initiative. Database exports, which are the correct route for network analysis and the only reproducible basis for a published structural claim.
- Oligrapher — Public Accountability Initiative. The network mapping tool built on the same data, useful for producing publication-quality relationship diagrams from a filtered subgraph.
- Public Accountability Initiative — Public Accountability Initiative. The operating non-profit and its published investigations, which explain the shape of the dataset's coverage better than any documentation could.
- ProPublica Nonprofit Explorer — ProPublica. Non-profit tax filings listing officers, directors and grants – the authoritative record behind foundation, think tank and trade association relationships.
- SEC EDGAR — US Securities and Exchange Commission. Proxy statements and ownership filings, the primary source for public company directorships that position edges should be traced back to.
- OpenCorporates — OpenCorporates Ltd. The route from an organisation entity to a registered legal person with a jurisdiction and identifier, which is where influence mapping meets corporate fact.
- Wikidata — Wikimedia Foundation. Independent structured claims and stable identifiers for the same people and organisations, useful for disambiguating entities with common names.
- OpenSanctions — OpenSanctions. Sanctions and politically exposed person data for screening the population an influence map produces.
Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.
Put it into practice
The Quantus Intel threat intelligence platform operationalises this source: it imports typed relationships with their dates, currency flags and citations intact, weights edges by whether they are referenced and current, and keeps community-edited claims visibly separate from register and court records in the same case.. Browse the full source catalogue, or follow any tag above into the rest of the library.