CourtListener / RECAP: Intelligence Source Guide
RECAP is a free archive of United States federal court dockets and filed documents, built from copies contributed by people who paid PACER for them. It is the cheapest route to the primary documents of federal litigation, and its coverage is shaped by what somebody else bought.
RECAP is a free archive of United States federal court dockets and filed documents, built from copies contributed by people who paid PACER for them. It is the cheapest route to the primary documents of federal litigation, and its coverage is shaped by what somebody else bought.
At a glance
| Source | CourtListener / RECAP |
|---|---|
| Category | Corporate, Ownership & Legal Records › Court Records & Litigation |
| Homepage | https://www.courtlistener.com/ |
| Machine interface | https://www.courtlistener.com/api/rest/v4/ |
| Format | JSON |
| Access | Open — no account required |
| Disciplines | Legal Intelligence, Financial Intelligence |
| Mission domains | Financial Crime, Fraud & Identity, Corruption & Governance |
US federal & state dockets/opinions. — as catalogued in the platform’s own source registry.
PACER is the electronic records system of the United States federal courts. It holds the docket of every federal civil, criminal, bankruptcy and appellate case, along with the documents filed in them, and it charges per page for access. RECAP is the mechanism by which those paid-for pages become free. A browser extension, maintained by the Free Law Project, sits alongside a user's PACER session; when that user retrieves a docket or purchases a document, a copy is uploaded to the archive and published on CourtListener at no cost to anyone who wants it afterwards. The result is a growing public mirror of federal court records, structured as dockets containing docket entries containing documents, with parties, attorneys, case numbers, nature of suit, judges and the full text of everything that has been contributed. Scanned documents are put through optical character recognition so the text is searchable. Alongside the passive archive there is an active retrieval path, which lets an authorised user supply their own PACER credentials so that a specific document is purchased and simultaneously added to the public archive. The name is PACER reversed, which tells you most of what you need to know about the project's politics.
Opinions tell you what a judge concluded. Dockets tell you what happened, and the documents tell you what the parties said, showed and were forced to disclose. For FININT and fraud work that difference is decisive: a complaint sets out a transaction structure in detail, an exhibit reproduces the contract, a receivership filing lists assets, a bankruptcy schedule enumerates creditors, a deposition excerpt names intermediaries, an indictment describes a laundering typology, and a forfeiture pleading identifies accounts. None of this appears in a published opinion, and much of it appears nowhere else in the open record because the parties produced it only under legal compulsion. RECAP is also the only free route to the procedural shape of a case – the pace, the motions, the sealing decisions, the withdrawal of counsel – which is frequently how you detect that something changed without any public announcement. The second distinctive job is temporal: dockets update as cases move, so with alerting this becomes a monitoring capability rather than a lookup, and litigation involving a target becomes something that reaches you rather than something you must remember to check.
Who publishes it, and why that matters
The archive is run by the Free Law Project, a small US non-profit, and the project's origin is a Princeton research effort concerned with public access to court records. That lineage matters because the whole design is an argument: federal court records are public by law, PACER charges for them anyway, and the archive exists to break the cost barrier by making each purchase serve everyone. The organisation has been an active participant in litigation and advocacy about PACER fees, and the fee structure itself has been the subject of a significant class action about whether the judiciary was charging more than the statute permitted. Two consequences for you. First, the incentives are aligned with your access rather than against it, but the resources are those of a small non-profit dependent on grants, memberships and donations, so heavy automated consumption is a real cost to a real organisation. Second, the archive's continued growth depends on volunteer behaviour – the extension has to be installed and used – so coverage is a function of a community's habits and can shift with browser policy changes, PACER interface changes and the project's own capacity.
Provenance is the first question to ask of any dataset and the one most often skipped. Who collects it, what their incentive is, whether they publish a methodology, and whether they correct the record when they get something wrong all bear directly on how much weight a finding drawn from it can carry.
What a record actually contains
The fields you will be working with, what each one means, and whether it is something you can pivot on. Read the meanings carefully — more analysis is wrecked by misreading a field than by failing to find one, and a field that looks like an observation is often an inference.
| Field | Type | What it means | Pivot value |
|---|---|---|---|
docket_id |
int | The archive's identifier for one case in one court. A dispute that moves from district to appellate court produces separate dockets, and treating them as one case or as unrelated cases are both errors. | Docket entries, parties, attorneys, the originating court record, and any opinions that issued from the case. |
docket_number |
string | The court's own case number, encoding the year, case type and a sequence, and prefixed by an office or division code in many districts. This is the identifier you quote and the one PACER understands. | The case in PACER itself, related cases in the same district, references in other filings and in press coverage. |
court |
string | Which federal court holds the case. Determines the applicable local rules, the sealing practice, the electronic filing conventions and, in practice, how complete the archive's coverage of that court is likely to be. | Court-level coverage assessment, comparison against federal caseload statistics, the correct PACER interface. |
case_name |
string | The caption. Abbreviated, inconsistent across the life of a case, and naming only the lead parties on each side, so it is a human label rather than a data field about who was involved. | Party records, which are the reliable route to the full set of litigants. |
parties |
array | Structured party records with role – plaintiff, defendant, intervenor, third-party defendant, creditor – and, where captured, the attorneys and firms appearing for each. Far richer and far more reliable than the caption. | Corporate entity resolution against registry data, counsel networks, identification of repeat litigants across unrelated matters. |
date_filed |
timestamp | When the case was filed, or for an entry, when the document was entered on the docket. Court dates, usable for timelines – but distinguish filing from entry from service, which can differ by days. | Correlation with corporate events, regulatory action, share price movement and news coverage. |
date_terminated |
timestamp | When the case closed, where the court has recorded it. Its absence usually means the case is live, but not always, because docket capture in the archive may simply predate the closure. | Case duration analysis; determining whether a monitored matter is still active before relying on the last entry. |
nature_of_suit |
string | The civil case category recorded at filing – contract, torts, civil rights, intellectual property, forfeiture and so on. Assigned by the filing party, sometimes inaccurately, and the basis of essentially all federal caseload statistics. | Population-level comparison against Federal Judicial Center statistics; filtering large docket sets to a case type. |
entry_number |
int | The sequential number of an entry on the docket. Gaps are informative: a missing number normally means a sealed or restricted entry rather than a data error, and gaps cluster where cases are sensitive. | Detecting sealing; measuring how much of a case is actually visible to you. |
description |
string | The docket text for an entry, written by clerk staff. Terse, abbreviated and non-standard across districts, but it is the only description of entries whose documents nobody has purchased. | Identifying which entries are worth buying; reconstructing case events where no document is available. |
is_available |
enum | Whether the document for this entry is actually in the archive. This is the single most important field in the dataset, because a docket can be fully captured while almost none of its documents are. | Purchase decisions; honest assessment of what your review of a case did and did not cover. |
document_text |
string | Extracted text of a contributed document or attachment, from the native PDF text layer or from optical character recognition of a scan. Quality varies accordingly, and exhibits – which carry most of the substance – are frequently images with poor recognition. | Named entities, account numbers, addresses, dates, transaction detail – the substance that dockets point at and opinions omit. |
assigned_to |
string | The judge assigned, and often separately the referred magistrate judge. Linked to the platform's people records where the judge is in the biographical data. | Judge profile, other cases before the same judge, recusal and assignment patterns. |
pacer_doc_id |
string | The identifier PACER uses for a specific document, which is how a request for the real thing is addressed. It is the bridge between the free archive and the authoritative source. | Purchasing the document from PACER; verifying that an archived copy matches the official record. |
Coverage — and what is not in it
Coverage is demand-driven and that is the defining property of this source. A case is in the archive because somebody with the extension installed looked at it – a journalist, a lawyer, a researcher, an activist – so the archive is dense in litigation that attracts attention and sparse in litigation that does not. High-profile civil cases, major criminal prosecutions, patent and antitrust litigation, cases involving well-known companies, and anything covered by the press are well represented. Routine debt collection, ordinary bankruptcy, prisoner petitions, immigration-adjacent civil matters and unremarkable contract disputes are represented poorly or not at all. Within a captured case, coverage is layered: the docket sheet is often complete because retrieving it is cheap, while individual documents are present only where somebody paid for them, so a case can look thoroughly archived and contain almost no filings. Geographically it covers the federal system – district, bankruptcy and appellate courts – and nothing else; state courts, which handle the great majority of American litigation, are outside it entirely. Update rhythm is continuous but uneven: an actively watched docket refreshes as people look at it, and a dormant one may not have been touched in years.
Known blind spots
Absence of evidence here is not evidence of absence. These are the conditions under which CourtListener / RECAP will not show you something that is nevertheless real:
- Nobody looked, nothing exists. The archive contains what volunteers purchased, so absence of a case is overwhelmingly a statement about attention rather than about litigation – and attention is systematically lower for exactly the routine, unglamorous cases that reveal patterns of conduct over time.
- Sealed and restricted material is absent by design and shows up only as gaps in the entry numbering, which means the most sensitive parts of a sensitive case are precisely the parts you cannot see and cannot easily measure.
- A captured docket is not a captured case. The docket sheet may be complete while the documents behind almost every entry are unpurchased, and an analyst who reviews the entries and concludes they have read the case has read a table of contents.
- State courts are entirely outside the federal system this covers, and they handle most American litigation including nearly all family, landlord and tenant, small claims and state criminal matters. A clean search here says nothing about a person's litigation history.
- Exhibits and attachments are less likely to be purchased than main filings because they cost more pages, so the documents most likely to contain contracts, financial records and correspondence are the ones most likely to be missing.
- Optical character recognition on scanned filings fails unevenly, and failure correlates with handwriting, poor photocopies, tables and non-standard layouts – which describes a great deal of what is filed as an exhibit in financial cases.
- Party strings are not resolved to real-world entities, so a corporate defendant appears as whatever the filing party typed, with subsidiaries, trading names and misspellings all separate and unlinked.
- Docket text is written by clerks in district-specific abbreviations with no controlled vocabulary, so text-mining docket descriptions across courts produces results that reflect local clerical conventions as much as case events.
- Cases can be captured once and never again, leaving a stale snapshot that presents as current. A docket last touched in 2021 will not show a 2024 judgment, and nothing about it announces that it is out of date except a modification timestamp.
Write the blind spot into the product. A statement that something “was not observed in CourtListener / RECAP” is defensible; a statement that it “did not happen” is not, and the difference is what survives cross-examination.
Access, licensing and what you may do with it
Access model: Open — no account required
Reading and searching the archive requires nothing at all – the docket and document search on CourtListener is open, and archived documents are free to download. Programmatic access to dockets, entries and documents runs through the same authenticated REST API as the rest of the platform, using a free account token, with the same versioning and quota considerations. The distinctive access route is contribution: installing the browser extension makes your own PACER usage add to the archive automatically, which is both the ethical price of using the resource and, in practice, the most reliable way to get the specific documents you need into it. There is also a retrieval path that lets you supply PACER credentials so that a document is purchased on your behalf and simultaneously published, which is the cleanest way to fill a gap in a case you are working on. PACER itself requires its own account with billing details, charges per page with a modest quarterly waiver for light users, and applies caps that do not extend to every kind of request – check the current fee schedule rather than relying on remembered figures.
Licence
Federal court filings are public records and judicial opinions are United States government work not subject to copyright, so the underlying material is freely usable, quotable and republishable as a matter of American law. The complications are elsewhere. Documents authored by private parties – expert reports, exhibits reproducing third-party works, attached publications – can contain separately copyrighted material that being filed in court does not release. The compiled archive, its metadata and the platform's own enhancements are governed by the site's terms of use, which contemplate research, journalism and public interest use rather than wholesale commercial redistribution. And the practical constraint that catches people is data protection rather than copyright: filings contain large volumes of personal information about parties and third parties, and holding that in bulk outside the United States engages obligations that American public-record status does not answer. Read the current terms, attribute the archive, and treat a decision to republish a filing verbatim as an editorial judgement rather than an automatic entitlement.
Rate limits and fair use
The API applies per-account quotas that have changed over time, and the archive is served by a small non-profit whose infrastructure costs are real. Practical etiquette matters here more than at most sources. Use bulk exports for corpus-scale work rather than paginating the API. Fetch dockets by identifier rather than repeatedly searching. Use docket alerts to learn about changes instead of re-polling a case on a schedule – that is exactly what they exist for and it is dramatically cheaper for both sides. When downloading documents, take what you need rather than mirroring a whole court. Identify yourself in your user agent with a contact address. And if your requirement is genuinely large or continuous, talk to the project about supported access rather than engineering around a limit; the organisation is approachable and would rather have the conversation than the traffic.
Licensing changes, and it changes without warning. A dataset that was free for research this year may not be free for commercial or evidential use next year. Confirm the current terms before you build a dependency on it, and record the terms you relied on alongside the data — the licence in force at the time of collection is part of the provenance.
Collecting it
How CourtListener / RECAP is actually pulled, in the order you would set it up. Prefer the bulk or export interface over per-item lookups wherever one exists: it is kinder to the publisher, faster for you, and gives a reproducible snapshot rather than a series of point-in-time answers you cannot reconstruct later.
| Method | Format | Cadence | Notes |
|---|---|---|---|
| Docket and document search | JSON|HTML | on demand; archive grows continuously as contributions arrive | The discovery route. Full-text search across contributed documents is the capability that distinguishes this from a docket index, and it will surface cases whose captions gave no clue they were relevant. |
| Docket retrieval by identifier | JSON | on demand | Once you have a docket, pull it and its entries directly. Check the availability flag on every entry so you know what fraction of the case you actually have rather than assuming. |
| Document download | bulk | per document | Contributed filings are downloadable as PDFs with extracted text. Preserve the original PDF alongside the text, because exhibits frequently carry meaning in layout and stamps that extraction discards. |
| Credentialed retrieval from PACER | JSON | per request, at PACER cost | The route for obtaining an entry nobody has bought. It costs money, it fills the public archive at the same time, and it is the honest way to close a gap in a case you are working. |
| Docket alerts | JSON|HTML | event-driven | The monitoring mechanism for live litigation. Set them on every case that matters to an open investigation, because a dispositive filing landing unnoticed is the standard way litigation monitoring fails. |
| Bulk exports | CSV|bulk | periodic | For statistical work across dockets, nature-of-suit distributions or party analysis at scale. Also the only reproducible basis for a published claim about the corpus. |
Ingesting it into the platform
Every step below is idempotent and cursor-based: interrupt one and it resumes from where it stopped rather than duplicating rows or losing progress. Collection is recorded per source, so a feed that quietly stops publishing shows up as a stale timestamp instead of silently thinning your coverage.
- Register the archive with its coverage caveat — Add RECAP in sources.php with the demand-driven coverage limitation recorded as a source-level note, so that every downstream product inherits the caveat rather than relying on an analyst to remember it.
- Model dockets, entries and documents as three layers — In ingest.php keep the hierarchy intact. Flattening entries into dockets loses the availability flag, and losing the availability flag is how a case with two purchased documents ends up described as fully reviewed.
- Store the availability state explicitly — Record for every entry whether a document exists in the archive, so the platform can report coverage of a case as a proportion rather than a binary. This single field is what makes an honest confidence statement possible.
- Extract parties as entities and resolve them — Push structured party records into the entity model, then run resolve-everything.php against corporate registry data, keeping the original filing string and the resolved entity as separate values with a confidence score.
- Mine documents for indicators — Run extracted document text through the platform's extraction so that addresses, phone numbers, email addresses, company names, account references and cryptocurrency addresses appearing in filings become searchable observables in ioc-workbench.php.
- Attach judges and counsel — Link assigned judges to person records and capture appearing attorneys and firms, because counsel networks are one of the most reliable relationship signals in litigation data and are almost never exploited.
- Set alerts on case entities — For every docket relevant to an open case, configure an alert so new filings arrive in alerts.php. Litigation moves on the court's schedule, not yours, and manual re-checking always decays.
- Build the timeline and record the gaps — Push filing, entry and termination dates into timeline.php alongside corporate and financial events, and record in cases.php both what the docket showed and which entries were unavailable or sealed.
Registered sources and their last-collected state are listed in sources.php, and the scheduled chain that keeps them current is in automation.php.
How it is wrong, and how to tell
Every dataset is wrong in characteristic ways. Knowing which ways is the difference between using a source and being used by one, and it is the part of source evaluation most often skipped because it is the part that takes work.
The archive is a faithful copy of what was purchased – documents are the court's own files, dockets are the court's own sheets, and the fidelity of an individual record is not in serious doubt. Text extraction quality varies with the document: native electronic filings extract cleanly, scanned exhibits do not, and the failure is silent. Metadata is generally sound because it comes from PACER rather than from parsing. The reliability question that actually matters is completeness, and it must be assessed case by case rather than assumed. For any docket you rely on, the honest measure is the proportion of entries whose documents are present, and it is common for that proportion to be low even in cases that feel well covered. A second and subtler completeness issue is currency: a docket captured once is a snapshot, and unless somebody has looked at it since, it stops at the date of that look. The right posture is that a positive record is strong evidence – this document was filed, this is its text – and an absence is no evidence at all, neither of a case not existing nor of a filing not having been made.
Characteristic false positives
- A captured docket reads as a complete case. Entry descriptions are present for everything while documents are present for a fraction, and reviewing the descriptions produces a confident narrative built on clerk shorthand rather than on filings.
- Party names are matched to the wrong legal person. Filing parties type company names inconsistently, subsidiaries appear where parents are meant, and matching a defendant string to a registry entry without corroboration attributes litigation to an entity that was never sued.
- The same dispute is counted several times. District, appellate and related-case dockets, plus consolidated and member cases in multidistrict litigation, mean naive counting of matters overstates litigation volume substantially and unevenly across case types.
- Nature of suit is trusted as a classification. It is selected by the filing party at the outset, is frequently inaccurate, and does not update when the case changes character, so filtering by it produces a sample that reflects filing habits as much as subject matter.
- Allegations in complaints are quoted as facts. A complaint is one side's pleading, drafted to survive a motion to dismiss rather than to be true, and the difference between an allegation and a judicial finding is the most consequential distinction in this entire dataset.
- Optical character recognition invents plausible values. A misrecognised digit in an account number, a date or a docket reference produces something that looks right and is wrong, and financial exhibits are exactly the material where recognition performs worst.
- Gaps in entry numbering are read as data errors. They usually indicate sealed or restricted filings, which is a finding about the case rather than a defect in the archive, and treating them as noise discards a real signal about what the court decided to protect.
- Stale dockets present as current. A case captured three years ago shows no subsequent activity, which looks identical to a case in which nothing has happened, and only the modification timestamp or a fresh retrieval distinguishes them.
None of these make the source unusable. They make it a source that requires corroboration before an assertion built on it goes into a product, which is true of every source and admitted by few.
Ageing
Two clocks run here and they behave differently. The documents do not age: a filing made in 2017 remains exactly what it was, and its evidential value is stable. The docket ages continuously while a case is live, and the archive's copy of it ages from the moment it was last contributed. That is the failure mode to design against, because a stale docket is indistinguishable from a quiet one. Practically, treat any docket in an open matter as reliable only up to its last capture date, and re-retrieve or set an alert before drawing a conclusion about the current state of a case. Party and counsel information ages too – firms substitute, attorneys withdraw, corporate parties merge or change name mid-case – and a filing from three years ago describes a configuration that may no longer exist. The entity resolutions you attached also age, since the company you matched a defendant string to may since have been dissolved or renamed. Re-resolve at analysis time rather than trusting an enrichment captured at ingest, and always report the capture date alongside any statement about what a docket shows.
What this source feeds
A source is only worth what it lets you conclude. These are the disciplines that collect through it, the mission domains it serves and the data points it yields — every one is a tag, so you can follow any thread from here into the rest of the library.
Collected by these intelligence disciplines
Serves these mission domains
Yields these data points
How each sector uses CourtListener / RECAP
The same dataset is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The records are shared — the constraints, thresholds and outputs are not.
🎖 Military and defence
Federal litigation is one of the few open windows into the defence industrial base that does not depend on a company choosing to disclose. Bid protests, contract disputes, False Claims Act actions, export control prosecutions and trade secret cases generate filings that describe supply chains, subcontractor relationships, programme failures and technical specifications in detail that procurement records never contain. Employment and whistleblower filings can reveal internal problems at a supplier well before any formal notification reaches the customer. For counter-proliferation and sanctions work, criminal dockets and civil forfeiture actions set out procurement networks and financial routing with a precision that only compelled disclosure produces. The obvious limits: classified and sensitive matters are sealed or handled outside the Article III system, so what is visible is the unclassified periphery, and coverage depends on someone having bought the documents.
🕵 National intelligence
For FININT and LEGINT this is a primary collection route rather than a supporting one. Indictments, forfeiture complaints, receivership filings and civil fraud pleadings describe financial structures, corporate layering, intermediaries and typologies in enough detail to constitute intelligence on their own, and they carry the weight of having been drafted for a court. Dockets also expose relationships that no register does – who retained which counsel, which entities were sued together, which third parties were subpoenaed. Two disciplines govern the work. First, distinguish rigorously between what a party alleged and what a court found, and label them differently in every product. Second, treat coverage as a variable to be measured, not assumed – if a conclusion rests on the absence of litigation, that conclusion is unsupportable without an independent check against federal caseload statistics or PACER itself.
👮 Law enforcement
Investigators use the archive for subject workup, network expansion and monitoring, at a cost of nothing where somebody else has already paid. A subject's prior civil litigation frequently discloses assets, business associates, addresses and disputes; criminal dockets identify co-defendants and counsel; bankruptcy filings enumerate creditors and transfers under penalty of perjury. Alerts turn an ongoing prosecution or related civil matter into a monitored feed. Three constraints matter. Most criminal history is in state courts and is not here. An archived copy is not a certified record and will not serve where authentication is required – obtain it from the court. And filings contain third-party personal information whose handling is governed by your own rules regardless of its public status, which is a live issue when you ingest documents in bulk rather than reading them one at a time.
🔍 Private investigation and corporate security
Asset tracing, judgment enforcement and pre-litigation due diligence all run on this material. Federal dockets reveal prior judgments, liens litigated, bankruptcy schedules listing assets and creditors, receiverships enumerating property, and business relationships disclosed in contract disputes. Because the documents are free where they exist, the economics are transformative compared with buying everything from PACER. The professional cautions are identification and completeness. Attribute a case to your subject only with a second discriminator beyond a matching name, and never report a clean litigation history without stating that state courts and unpurchased documents are outside what was searched. Where a document you need is not in the archive, buy it – the cost is small, and it adds to the commons at the same time.
📰 Journalism and OSINT media
Court filings are the backbone of a great deal of American investigative reporting because they are public, quotable, privileged in most jurisdictions and full of documents that nobody would otherwise release. RECAP removes the cost barrier that used to make systematic docket work a resourced-newsroom activity. Full-text search across contributed documents surfaces stories that would never be found by searching case names. Docket alerts let one reporter follow dozens of matters. The two publication rules are simple and constantly broken: an allegation in a complaint is not a fact, and a document being in a public file does not settle whether personal information in it should be republished – redaction failures are common and the decision to amplify them is yours. Installing the extension is also the right thing to do, since a newsroom that draws heavily on the archive should be contributing to it.
🌍 NGO, humanitarian and human rights
Accountability and human rights organisations use federal dockets to document patterns across cases – the same corporate defendant, the same legal theory, the same practices challenged in different districts – and to obtain the primary documents behind matters they already know about. Immigration-adjacent civil litigation, civil rights actions and environmental enforcement all generate filings that support documentation to an evidentiary standard rather than an inferential one. The coverage caveat must be stated in published work: the archive contains what somebody paid for, so a pattern found in it may be a pattern in what attracts attention. For litigation-support work, the credentialed retrieval route is worth budgeting for, because filling gaps costs little per document and permanently benefits everyone working the same issue.
🎓 University and research
Empirical legal scholarship has been reshaped by the availability of docket data at this scale, and the corpus supports work on case duration, settlement behaviour, judicial assignment, party representation and the anatomy of litigation that published opinions cannot address. The sampling problem is severe and must be handled explicitly: inclusion depends on a volunteer having purchased the record, which correlates with case salience, media attention, party resources and researcher interest – frequently the same variables under study. The standard corrective is to benchmark against the Federal Judicial Center's case-level statistics, which provide the true population of federal filings and allow a coverage rate to be estimated per court, per year and per case type. Use bulk data for reproducibility, report the snapshot date, and treat the availability flag as a variable rather than filtering silently on it.
Playbook: working CourtListener / RECAP end to end
A repeatable sequence from first pull to finished product. Each phase states what you are trying to establish, not merely what to click — the objective is a defensible chain of reasoning, not a completed checklist.
Phase 1 — Establish what you can afford not to know
Decide at the outset whether your question tolerates incomplete coverage. Finding evidence of something is well served by an incomplete archive; proving the absence of something is not. If your conclusion depends on a negative, plan from the start to verify it against PACER or federal caseload statistics rather than against the archive.
Phase 2 — Search documents, not just case names
The full-text search over contributed filings is the capability that makes this different from a docket index. A target's name buried in an exhibit to an unrelated case is exactly the finding you will never reach by searching captions, and it is the most common way this source produces something genuinely new.
Phase 3 — Pull the docket and read the shape before the substance
Before opening documents, read the entry sequence: how fast the case moved, when counsel changed, where the gaps in numbering are, when sealing happened, whether there was a dispositive motion. The procedural shape frequently tells you what the case is really about faster than any individual filing.
Phase 4 — Measure your coverage of the case explicitly
Count how many entries have documents available and record it. A case where you can read eight of two hundred entries is a case you have sampled, not read, and writing that number into your notes is what stops it being forgotten by the time you draft.
Phase 5 — Go straight to the attachments
Main filings are often procedural; the contracts, financial records, correspondence and expert analysis are in the exhibits. Attachments are also the least likely to have been purchased, so this is usually where you decide to spend money.
Phase 6 — Buy what you need and contribute it
Where a decisive entry is unavailable, use the credentialed retrieval route. It costs a small amount, gives you the authoritative document, and adds it to the public archive at the same time. This is the intended economics of the whole system and using the archive without ever feeding it is a free ride.
Phase 7 — Resolve parties properly before building anything
Take the structured party records, normalise them, and match to corporate registry data with a recorded confidence. Do not build a network from raw party strings; subsidiaries, misspellings and trading names will produce a diagram that is mostly artefact.
Phase 8 — Exploit counsel as a relationship signal
Which firm appears for which party, and where the same counsel appears across nominally unrelated matters, is one of the most under-used signals in litigation data. Counsel relationships are stable, deliberate and disclosed, which makes them better evidence of association than shared addresses.
Phase 9 — Separate allegation, admission and finding at the point of reading
Tag every extracted fact with its status as you read it, not later. A complaint alleges. An answer admits or denies. A stipulation is agreed. A judgment finds. Products that lose this distinction are indefensible, and it is impossible to reconstruct once the notes are written.
Phase 10 — Set alerts before you move on
Any live docket relevant to an open matter gets an alert. Cases move on the court's schedule and the filing that changes your assessment will arrive on a day you were doing something else. This is the cheapest single improvement available to most litigation monitoring.
Phase 11 — Cross-check the entity picture against registers and sanctions data
Take resolved corporate parties out to registry and sanctions sources. A defendant that turns out to be a recently formed entity at a formation agent's address, or one whose officer is a listed person, changes the character of the whole matter and is not visible from the docket alone.
Phase 12 — Write the coverage statement into the product
State which courts were searched, that the archive is contribution-driven, how much of each key docket was available, that sealed material is excluded, and that state courts are outside scope. A reader who is not told this will assume completeness, and the resulting misreading will be attributed to you.
The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.
What to pair it with
No single source carries a finding. These are the datasets that corroborate, extend or contradict this one — and a source that contradicts is worth more than one that agrees, because it is the only thing that will tell you when you are wrong.
| Source | Relationship | What it adds |
|---|---|---|
| PACER | prerequisite | The authoritative source. Everything in the archive originated here, and anything missing must be purchased here. Also the only place to confirm that a docket is current. |
| CourtListener | extends | The surrounding platform – case law, judges, oral argument and the citation graph – which supplies the doctrinal and biographical context around any docket. |
| Federal Judicial Center Integrated Database | corroborates | Case-level statistics for all federal civil, criminal and bankruptcy filings, which is the reference population for measuring how much of a court's caseload the archive actually contains. |
| OpenCorporates | prerequisite | Turns a corporate party string into a registered legal entity with a jurisdiction and identifier, which is the step that connects litigation to the rest of a corporate picture. |
| SEC EDGAR | corroborates | Public company filings disclose material litigation and give the company's own account of exposure, which can be compared against the docket record. |
| OpenSanctions | extends | Screens resolved parties and named individuals against sanctions and PEP data, which routinely reframes a civil dispute as a sanctions or corruption matter. |
| OCCRP Aleph | extends | Cross-references litigants against leaked and scraped document collections, supplying ownership and offshore context that federal filings reference but rarely establish. |
| Free Law Project | prerequisite | The operator, the browser extension, and the documentation of how the archive is built and funded – which is also how you understand its coverage bias. |
| United States Courts | prerequisite | Official material on court structure, electronic filing, sealing practice and the PACER fee schedule, all of which determine what a docket looks like and what it costs. |
Legal, ethical and operational constraints
Federal court records are public by law and the documents are generally free of copyright as government or public record material, which is why this archive can exist at all. The real constraints sit in three places. First, privacy: filings contain names, addresses, dates of birth, financial account references and details about children, health and immigration status, and although federal rules require redaction of certain identifiers, compliance is imperfect and redaction failures are routine. Holding a bulk corpus of this outside the United States engages data protection obligations – lawful basis, purpose limitation, retention, subject rights – that American public-record status does not answer for you. Second, use restrictions: in the United States, using court records to make employment, tenancy or credit decisions can bring you within consumer reporting regulation with accuracy and dispute obligations attached, and doing that from a scraped archive is a poor position to be in. Third, defamation and fair reporting: allegations in filings are privileged to report accurately in most American jurisdictions, but that privilege is narrower elsewhere, and repeating an unproven allegation as fact is actionable in many places. Decide these questions before you build a pipeline, not after a complaint arrives.
Operational security
Three separate exposures deserve thought. Reading the archive over the API is attributable to your account, and the operator can see which dockets and documents you retrieved – a concentrated pattern around one company or person is a clear statement of interest. Using PACER directly is considerably more exposing: it is an account with billing details tied to an identity, operated by the judiciary, and your query and purchase history exists there. Running the browser extension means your PACER activity is automatically published, which is admirable and is also a disclosure – if you retrieve a sealed-adjacent or unusual docket, the fact that somebody just pulled it becomes public. For sensitive pre-filing or pre-enforcement work, consider retrieving through a neutral path, using bulk data and searching locally, and thinking carefully about whether an alert on a specific docket – which creates a dated third-party record of exactly what you are watching – is a trade you want to make. Where the subject is well resourced and litigious, assume that interest in a docket is discoverable.
Two rules that hold regardless of jurisdiction. Collection that is lawful is not automatically proportionate, and a dataset assembled for one purpose does not carry consent for another. Where the records concern identifiable people, the question is not only whether you may hold the data but whether holding it serves the purpose you are accountable for.
Is it earning its place?
Sources accumulate. Feeds get added during an incident and are never reviewed again, and a decade later the pipeline is carrying dead weight that nobody dares remove. These are the measures that show whether CourtListener / RECAP is contributing anything, and they are worth baselining now so the answer is available later.
- Document availability rate for the dockets your cases actually rely on, reported as a proportion of entries, which is the single most honest measure of how much of a case you have read.
- Estimated coverage of the archive for the courts and case types you work in, benchmarked against federal caseload statistics, so that negative findings can be qualified rather than asserted.
- Number of documents purchased and contributed by your organisation, which measures whether you are consuming a commons or maintaining it.
- Party resolution precision, measured by how often a match from a filing string to a registered entity survives corroboration, tracked separately from source quality because it is your error.
- Alert responsiveness: elapsed time between a docket entry appearing and an analyst acting on it, which is the real test of whether litigation monitoring is functioning or decorative.
- Proportion of extracted document facts tagged with their evidentiary status – alleged, admitted, stipulated, found – audited on a sample of finished products.
- Count of findings sourced from an exhibit rather than a main filing, since exhibits are where the substance lives and a low count usually means the review stopped too early.
- Age of the most recent docket capture for live matters, tracked so that a stale snapshot is refreshed before it underpins a decision.
Beware of volume. Indicator counts rise easily and say almost nothing. Unique contribution — findings this source produced that no other source in your stack would have — is the measure that matters, and it is usually far lower than anyone expects.
Tradecraft notes
The distinctions that separate a competent analyst from a fast one:
- The availability flag is the most important field in the dataset. Read it before you read anything else, and never describe a case as reviewed without knowing what fraction of it you could actually open.
- Docket text is clerk shorthand, not a description of the case. It varies by district, uses local abbreviations and omits substance, so mining it across courts measures clerical convention as much as judicial activity.
- Buy the document when it matters. The cost of a filing is trivial against the cost of an analyst reasoning from an entry description, and the purchase adds permanently to the public record.
- Gaps in entry numbering are evidence, not error. They generally mark sealed filings, and where they cluster tells you which part of a case the court decided to protect – which is often the part you most wanted.
- Counsel is a stronger association signal than address or name similarity. Firms are retained deliberately, appearances are recorded formally, and the same counsel across unrelated matters is a real relationship rather than a coincidence.
- Distinguish district, appellate and consolidated dockets when counting. Multidistrict litigation in particular will inflate any naive count of matters by an order of magnitude if member cases are treated as independent.
- Nature of suit is a filing-party choice, not a classification. Use it to narrow a large set, never to define a population for analysis, and expect it to be wrong in a meaningful minority of cases.
- Keep the original PDF, not just the extracted text. Stamps, signatures, redaction boxes, exhibit markings and layout carry information that extraction discards, and financial exhibits in particular are unreadable as plain text.
- Contribute what you consume. This archive exists because people paid for documents and gave them away, and an organisation that draws on it heavily without installing the extension is taking a public good without maintaining it.
Questions analysts actually ask
Is RECAP a complete archive of federal court records?
No, and it does not claim to be. It contains what volunteers with the browser extension happened to retrieve from PACER. Coverage is dense in high-profile litigation and thin in routine cases, and within any case the docket sheet is far more likely to be present than the documents. Absence of a case here is not evidence that it does not exist.
Why can I see the docket entries but not the documents?
Because retrieving a docket sheet is cheap and buying each document is not. Somebody pulled the docket and did not buy the filings. The availability flag on each entry tells you which documents exist in the archive, and you can purchase the missing ones through the credentialed retrieval route, which also publishes them.
Does it cover state courts?
No. This is the federal system – district, bankruptcy and appellate courts. State courts handle the great majority of American litigation, including nearly all family, landlord and tenant, small claims and state criminal matters, and none of that is here. A search returning nothing says very little about a person's litigation history.
Can I use an archived document as evidence?
For investigative purposes, yes, and it is a true copy of what was filed. For proceedings requiring an authenticated record, obtain the document from the court itself. The archive is a mirror maintained by a third party, and while it is faithful, it is not the official record and does not come with certification.
How much does PACER cost?
It charges per page with a cap on most single documents and a modest quarterly threshold below which fees are waived. The specifics have changed and have been the subject of litigation about whether the judiciary was charging more than the statute allowed, so check the current fee schedule rather than relying on figures you remember.
What do gaps in the entry numbers mean?
Almost always sealed or restricted filings rather than missing data. That is analytically useful: it tells you the court decided something in the case should not be public, and where those gaps cluster is often the most sensitive part of the matter. Record them rather than ignoring them.
Should I install the extension?
If you use the archive, yes. It costs you nothing, publishes documents you were paying for anyway, and is the mechanism by which the resource exists. There is one consideration to weigh: your retrievals become public, so if the fact that somebody pulled a particular docket is itself sensitive, plan for that separately.
How do I know whether a docket is up to date?
Check when it was last contributed. A docket captured two years ago will show nothing since, and that looks identical to a case in which nothing has happened. For live matters, set an alert or re-retrieve before drawing any conclusion about the current state of the case.
Can I search inside the documents?
Yes, and this is the archive's most under-used capability. Full-text search runs over contributed filings including material recovered by optical character recognition, which means a name in an exhibit to an unrelated case is findable. Expect recognition failures on poor scans, so treat a null result as inconclusive for older or handwritten material.
Standards, formats and interoperability
What this source speaks natively, and what it has to be translated into before a partner can consume it. Work that arrives in a recognised format is easier to defend, easier to hand over and easier to automate against:
- The federal docket number format encodes year, case type and sequence with a district office prefix, and understanding it is what lets you construct references PACER will accept and match citations found elsewhere.
- Nature of suit codes are the classification underlying federal civil caseload statistics, which is what makes archive coverage measurable against the true population of filings.
- Federal Rules of Civil and Criminal Procedure govern filing, service and sealing, and are the reference for interpreting why a docket has the entries and the gaps it has.
- Federal privacy rules require redaction of specified personal identifiers in filings, which explains partially masked values and, by their frequent breach, the presence of material that should not be public.
- The Federal Judicial Center's Integrated Database is the standard population reference for federal civil, criminal and bankruptcy cases and the accepted way to estimate a docket sample's coverage.
- Optical character recognition over scanned filings is what makes the corpus searchable, and its known failure characteristics on handwriting, tables and poor reproductions should be treated as part of the data model.
- The platform exports entities, documents and relationships in STIX 2.1, MISP, CSV, JSON and JSONL, so litigation-derived observables and organisations move into shared case structures directly.
References
Primary documentation and authoritative references for this source. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.
- RECAP — Free Law Project. The project page: what the archive is, how the extension works, and the case for why federal court records should not sit behind a paywall.
- CourtListener RECAP archive — Free Law Project. The search interface over contributed dockets and documents, including full-text search across filings.
- CourtListener REST API documentation — Free Law Project. Endpoints for dockets, entries and documents, plus authentication and quota. The authoritative reference rather than remembered field names.
- RECAP API documentation — Free Law Project. The specific interface for docket and document retrieval, including the credentialed route for purchasing and publishing a filing.
- CourtListener bulk data — Free Law Project. Periodic exports, which are the correct basis for statistical work over dockets and for any reproducible published claim about the corpus.
- PACER — Administrative Office of the US Courts. The official system: registration, the current fee schedule, and the authoritative copy of every docket in the archive.
- Federal Judicial Center Integrated Database — Federal Judicial Center. Case-level data for the full federal caseload, and the standard instrument for measuring how much of a court's litigation the archive actually holds.
- United States Courts — Administrative Office of the US Courts. Court structure, electronic filing policy, sealing practice and published caseload statistics – the institutional context for every field in a docket.
- Free Law Project — Free Law Project. The operating non-profit, its funding, its open-source tooling and its advocacy on public access, which is also how you understand the archive's coverage bias.
- OpenCorporates — OpenCorporates Ltd. The practical route from a corporate party string in a filing to a registered legal entity with a jurisdiction and identifier.
Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.
Put it into practice
The Quantus Intel threat intelligence platform operationalises this source: it keeps the document availability flag attached to every docket entry so coverage is reported rather than assumed, resolves litigant strings against corporate registry data with confidence recorded, and turns filings into searchable observables without inventing anything the record does not contain.. Browse the full source catalogue, or follow any tag above into the rest of the library.