GovInfo / Official Gazettes: Intelligence Source Guide
GovInfo is the US Government Publishing Office’s authenticated archive of federal publications – statutes, regulations, the Congressional record, court opinions, hearings, budget documents and presidential papers – with a keyed JSON API over the whole corpus. It is the version of record, digitall…
GovInfo is the US Government Publishing Office's authenticated archive of federal publications – statutes, regulations, the Congressional record, court opinions, hearings, budget documents and presidential papers – with a keyed JSON API over the whole corpus. It is the version of record, digitally signed.
At a glance
| Source | GovInfo / Official Gazettes |
|---|---|
| Category | Corporate, Ownership & Legal Records › Legislation, Gazettes & Public Records |
| Homepage | https://www.govinfo.gov/ |
| Machine interface | https://api.govinfo.gov/ |
| Format | JSON |
| Access | Free registration — API key at no cost |
| Disciplines | Government Intelligence |
| Mission domains | Corruption & Governance, Election Security & PSYOP, Counter-Terrorism, Sanctions Evasion |
Government publications & legal notices. — as catalogued in the platform’s own source registry.
GovInfo is the Government Publishing Office's system for authenticating, preserving and distributing the official publications of all three branches of the United States federal government. It replaced the older FDsys and remains the government's own repository of record. Content is organised into collections, each a distinct publication series: the Federal Register and the Code of Federal Regulations; bills, resolutions and public laws; the Statutes at Large; the United States Code; the Congressional Record and its bound edition; committee hearings, reports and prints; the Congressional Serial Set; United States Courts opinions from participating federal courts; the Compilation of Presidential Documents; budget documents; the Congressional Directory; Government Accountability Office reports; and a long tail of agency publications. The distinguishing feature is authentication: PDFs carry a digital signature and a visible seal indicating that the document has not been altered since it was published by the issuing authority, which is what makes GovInfo the citable source in litigation rather than a convenience copy. The API is a JSON service requiring a free key issued through the government's shared API key service. It exposes the collection list, packages and their granules – a package being a publication unit such as a day's Congressional Record and a granule being an addressable piece within it such as a single speech or court opinion – along with published-date listings, per-package summaries and metadata, related-document links and a search service over the corpus.
This is where you go when the question is what the law actually says, what Congress actually said, or what a court actually held, and you need the answer in a form that survives challenge. Three capabilities are not available anywhere else at this price. First, authentication: a digitally signed PDF of a statute or rule is admissible and unarguable in a way that a copy from an aggregator is not, and for any work heading toward litigation, sanctions review or regulatory enforcement that difference is the whole point. Second, breadth across branches: the Federal Register covers the executive, but the legislative and judicial record – hearings where a company's conduct was examined under oath, committee reports explaining what a statute was meant to do, federal court opinions naming parties and describing conduct – is here and nowhere else in one place. Third, granule-level addressing: the API lets you retrieve a single court opinion, a single statute section or a single speech rather than a whole day's publication, which makes the corpus practically searchable rather than nominally available. The catalogue entry calls this Official Gazettes, and the honest qualification is that GovInfo is the United States gazette system and nothing else. It contains no foreign law. Analysts working on other jurisdictions need the relevant national gazette, and this guide points at several rather than pretending one service covers the world.
Who publishes it, and why that matters
The Government Publishing Office is a legislative branch agency whose statutory job is producing and distributing federal government publications, a function it has performed since 1861. GovInfo is its digital expression and is funded as core infrastructure rather than as a project, which makes it about as durable as anything in this catalogue. GPO also operates the Federal Depository Library Program, which distributes the same material to libraries nationwide, and that institutional commitment to permanent public access shapes the system: content is preserved rather than rotated, identifiers are stable, and the digital signature infrastructure exists specifically so that documents remain verifiable decades later. The incentive structure is archival and legal, not analytical, which explains the system's characteristic strengths and weaknesses. Metadata is rigorous and standards-based, and the API reflects a records-management view of the world – packages, granules, collections – rather than an investigator's view. Search is competent but not sophisticated, and the corpus rewards someone who knows the structure of federal publishing over someone who types keywords. The service has been affected by appropriations lapses in the past, and the API key infrastructure is shared with other federal services, so an outage there affects access here.
Provenance is the first question to ask of any dataset and the one most often skipped. Who collects it, what their incentive is, whether they publish a methodology, and whether they correct the record when they get something wrong all bear directly on how much weight a finding drawn from it can carry.
What a record actually contains
The fields you will be working with, what each one means, and whether it is something you can pivot on. Read the meanings carefully — more analysis is wrecked by misreading a field than by failing to find one, and a field that looks like an observation is often an inference.
| Field | Type | What it means | Pivot value |
|---|---|---|---|
packageId |
string | The identifier for a publication unit – a day's Congressional Record, a single bill version, a court opinion package, a CFR title. It is stable, human-legible and the primary key for retrieval and for citation in your own records. | All granules within the package, and the package's own summary metadata and related documents. |
granuleId |
string | The identifier of an addressable component within a package: a single speech, a statute section, an individual court opinion. This is the level at which analysis actually happens and the level most integrations forget exists. | The granule's text, its own metadata, and related granules across packages. |
collectionCode |
enum | Which publication series the package belongs to – Federal Register, bills, public laws, Congressional Record, courts opinions, hearings, reports and so on. It is the primary filter and each collection has different metadata richness. | The whole series, and the collection's documented metadata schema. |
dateIssued |
timestamp | The publication or issuance date of the package. For legislative material this is the date of the proceeding; for court opinions it is the date of the decision; for regulations it is the publication date. | Everything else issued the same day, which for hearings and the Congressional Record is often the relevant context. |
lastModified |
timestamp | When the record was last changed in the system. Used for incremental collection – you harvest by modification date rather than re-walking the corpus – and it tracks the record, not a revision of the underlying document. | none |
title |
string | The publication title, which for bills and hearings encodes a great deal – the bill number, the committee, the subject of the hearing. Titles here are more informative than in most corpora because they follow publishing conventions. | none |
governmentAuthor / committee |
array | The issuing body: the agency, the chamber, the committee or subcommittee, the court. For hearings and reports this identifies the committee whose jurisdiction produced the document, which is analytically important. | The committee's other hearings and reports, and its membership at the time. |
docClass / bill number |
string | Within legislative collections, the document class and number – the bill or resolution designation and its version. Bills exist in multiple versions and the version code is what distinguishes what was introduced from what passed. | Other versions of the same bill, the public law it became, and the committee reports on it. |
court / case identifiers |
string | In the courts collection, the deciding court, the docket number and the case designation. Participating federal courts contribute opinions with structured case metadata that supports party-name search. | The court's docket system, other opinions in the same case, and the named parties. |
congress / session |
string | The Congress number and session for legislative material. Bill numbers repeat every Congress, so a bill designation without a Congress number is ambiguous – a very common failure in cross-referencing. | The full legislative history for that Congress and the members serving in it. |
download links |
object | URIs for the available renditions – authenticated PDF, plain text, XML, and in some collections structured formats such as the modern legislative XML schemas. Different renditions suit different tasks and the XML is far better for analysis than the PDF. | none |
related documents |
array | Cross-references maintained by the system between related packages: a bill to its committee reports and to the public law, a rule to its Federal Register publication and its CFR codification. | The complete document family for a legislative or regulatory action, which is the point of the collection. |
authentication indicator |
enum | Whether the rendition carries a digital signature and visible seal certifying integrity. This is the field that distinguishes a citable version of record from a convenience copy, and it is the reason to use GovInfo rather than a mirror. | none |
suDocClass |
string | The Superintendent of Documents classification, the library classification scheme for US government publications, which persists across print and digital and links to depository library holdings. | Physical holdings in federal depository libraries and older material never digitised. |
Coverage — and what is not in it
The corpus spans all three branches of the United States federal government, with depth that varies sharply by collection. Regulatory material – the Federal Register and the Code of Federal Regulations – is comprehensive from the mid-1990s in digital form, with substantial retrospective digitisation extending the Federal Register back toward its 1936 origin. Legislative material covers bills, public laws, the Congressional Record, hearings, reports and the Serial Set, with modern Congresses fully covered and historical coverage extending back much further for some series through ongoing digitisation projects, in some cases into the nineteenth century. The United States Code and Statutes at Large give the codified and session-law forms of federal statute. The courts collection is the important qualification: it holds opinions from participating federal appellate, district and bankruptcy courts, and participation is not universal, so it is a large and useful sample rather than a complete corpus of federal case law. Presidential documents run from the Compilation series with earlier material in the predecessor weekly compilation. Update rhythm follows the underlying publications: the Federal Register daily on business days, the Congressional Record on days Congress sits, court opinions as they are issued and transmitted, hearings and reports weeks or months after the proceeding they record, and the codified Code of Federal Regulations on its annual revision cycle by title.
Known blind spots
Absence of evidence here is not evidence of absence. These are the conditions under which GovInfo / Official Gazettes will not show you something that is nevertheless real:
- It is United States federal only. There is no foreign law, no treaty text beyond what appears in US publications, and no state or local material, so an investigator working outside the United States gets nothing here and needs the relevant national gazette.
- The courts collection is not complete federal case law. Participation by courts is voluntary and uneven, so absence of an opinion means it may not have been contributed, may be unpublished, or may sit only in the court's own docket system.
- Hearings and committee reports appear long after the proceeding, sometimes many months, so this is an archival record of legislative activity rather than a monitoring source for it.
- Classified, restricted and sensitive material is simply not published, and closed committee sessions leave no transcript here. Silence about a subject is not evidence about the subject.
- Historical digitisation is uneven across collections and eras, which biases any quantitative study of publication volume over time and makes absence in older material uninformative.
- Search is document-oriented rather than entity-oriented. There is no normalised person or organisation index, so finding every mention of a company means text searching with variants and accepting imperfect recall.
- The system reproduces what was published, including errors. A misstatement in a hearing transcript or a typographic error in a statute is preserved faithfully, and corrections appear as separate documents rather than as amendments.
- Granule-level structure varies dramatically by collection. Some series are richly segmented and others are effectively a single PDF, so an integration that assumes uniform granularity will break on the collections that matter least to its author and most to someone else.
- Nothing here reflects the current legal state by itself. The Code of Federal Regulations edition is an annual snapshot and the United States Code has its own currency rules, so present-tense legal assertions need the continuously updated sources instead.
Write the blind spot into the product. A statement that something “was not observed in GovInfo / Official Gazettes” is defensible; a statement that it “did not happen” is not, and the difference is what survives cross-examination.
Access, licensing and what you may do with it
Access model: Free registration — an account or API key, at no cost
The API requires a key, obtained free through the shared federal API key service, and the key is passed with each request. A demonstration key exists for evaluation with tight limits; get a real one before building anything. The interaction model is: list collections, then either walk packages by published date within a collection or use the search service, then fetch a package summary, then enumerate its granules, then retrieve the rendition you want. That is more steps than a flat search API, and the payoff is precision – you can request exactly one opinion or one statute section rather than parsing a five-hundred-page PDF. Bulk access exists separately for the collections most in demand as datasets, particularly the structured legislative XML, and is the right route for anyone building a corpus rather than answering a question. The public website exposes the same content with browse-by-collection and search interfaces, and is much the faster route for a one-off lookup. The GPO maintains its API documentation and issue tracking in public, which is the most reliable place to confirm current endpoint behaviour.
Licence
United States government works are in the public domain domestically, so the overwhelming majority of this corpus can be copied, republished, indexed and commercialised without permission. Three exceptions matter in practice. Some publications reproduce copyrighted material with permission – submitted testimony, exhibits, articles inserted into the Congressional Record, standards incorporated by reference – and that material retains its own protection. Court opinions are public domain as government works, but headnotes and editorial apparatus added by commercial publishers are not, and none of that appears here anyway. And authentication is a property of the file GPO serves; a copy you re-host loses the digital signature and with it the evidentiary advantage, so redistributing an unsigned derivative and calling it authenticated is a misrepresentation. Cite the package identifier and the official citation. For legal use, link to or supply the authenticated PDF rather than a text extraction, because the signature is the point.
Rate limits and fair use
The key-based limits are applied by the shared federal API gateway and are generous for interactive and normal programmatic work while being genuinely restrictive for corpus harvesting. Plan accordingly: use published-date walks and modification-date incrementals rather than repeated searching, request only the metadata you need, and never enumerate granules for packages you have not already decided to use. Where a whole collection is required, use the bulk data route instead of the API – this is not merely etiquette, it is the difference between a job that finishes and one that spends a week hitting a limit. Cache everything: this is an archive, and archived documents do not change, so a package you fetched last year is still correct. Identify yourself in the user agent, handle throttling responses by backing off rather than retrying, and keep a single collection process rather than parallelising across many workers with the same key.
Licensing changes, and it changes without warning. A dataset that was free for research this year may not be free for commercial or evidential use next year. Confirm the current terms before you build a dependency on it, and record the terms you relied on alongside the data — the licence in force at the time of collection is part of the provenance.
Collecting it
How GovInfo / Official Gazettes is actually pulled, in the order you would set it up. Prefer the bulk or export interface over per-item lookups wherever one exists: it is kinder to the publisher, faster for you, and gives a reproducible snapshot rather than a series of point-in-time answers you cannot reconstruct later.
| Method | Format | Cadence | Notes |
|---|---|---|---|
| Collection listing and published-date walk | JSON | Daily for active collections | Enumerate packages issued in a date range within a collection. The reliable way to build and maintain a corpus incrementally without depending on search behaviour. |
| Modification-date incremental | JSON | Daily | Harvest packages changed since your last run. Catches late additions and corrections that a published-date walk alone will miss, which is common in hearings and court collections. |
| Package summary and granule enumeration | JSON | Once per package | Retrieve a package's metadata and the list of addressable components within it. This is where the corpus becomes analysable rather than merely downloadable, and it is the step most integrations skip. |
| Search service | JSON | On demand | Full-text and metadata search across collections with faceting. Good for targeted questions and for entity sweeps; not a substitute for systematic collection because recall is not guaranteed. |
| Bulk data download | bulk | Per release | Packaged structured data for the collections published as datasets, particularly legislative XML. The correct route for building a full local corpus and the only one that will finish. |
| Authenticated PDF retrieval | bulk | Per document needed | Fetch the digitally signed rendition for documents that will be relied on evidentially. Store the original file rather than an extraction, because the signature does not survive conversion. |
Ingesting it into the platform
Every step below is idempotent and cursor-based: interrupt one and it resumes from where it stopped rather than duplicating rows or losing progress. Collection is recorded per source, so a feed that quietly stops publishing shows up as a stale timestamp instead of silently thinning your coverage.
- Register each collection as its own source — sources.php holds the Federal Register, courts, hearings, bills and statutes collections separately with their own cadence and last-collected state, because they update on completely different rhythms and a single registration hides which one has stalled.
- Collect at package level, analyse at granule level — import.php stores the package as the unit of provenance and the granules as the units of content, preserving both identifiers. A finding cites the granule; the audit trail cites the package.
- Prefer structured renditions over PDF — Where XML or plain text renditions exist, ingest.php takes those for analysis and retains the authenticated PDF as the evidential copy. Parsing text out of a PDF when the publisher offers XML is self-inflicted error.
- Extract entities from hearings and opinions — enrich.php runs named-entity extraction over testimony, opinions and reports to surface companies, individuals and places, feeding org-profile.php and entity.php. The extraction is mechanical text processing over documents that already exist; no relationship in the platform originates from a language model.
- Link the document family — resolve-everything.php follows the related-document links to connect a bill to its reports, its public law and its codification, so a legislative action appears in timeline.php as one story rather than six unconnected records.
- Preserve authentication state — The platform records whether the retrieved rendition was digitally signed and stores the original file unaltered, so that export.php can supply an evidentially sound copy rather than a reformatted derivative.
- Run incremental harvests on modification date — cron.php harvests by last-modified rather than re-walking published dates, which catches corrections and late-added hearings without re-fetching a corpus that does not change.
- Hold the currency caveat with codified material — Code of Federal Regulations and United States Code packages are annual or periodic snapshots. The platform tags them with their edition so that any citation carries the edition date and no query presents them as current law.
Registered sources and their last-collected state are listed in sources.php, and the scheduled chain that keeps them current is in automation.php.
How it is wrong, and how to tell
Every dataset is wrong in characteristic ways. Knowing which ways is the difference between using a source and being used by one, and it is the part of source evaluation most often skipped because it is the part that takes work.
Metadata quality is high and structurally guaranteed rather than curated: it is generated from the publishing process by the body that produced the documents, and the digital signature makes integrity verifiable rather than asserted. Identifiers are stable, which is rarer than it should be. Where quality varies is in granularity and in retrospective material. Modern collections with structured XML are excellent; older scanned material may have optical character recognition errors in its text layer, which affects search recall in ways that are invisible – a name misrecognised in a 1950s hearing simply will not be found. Court opinion metadata depends on what the contributing court supplied and is less consistent than legislative or regulatory metadata. The deeper point about quality is one of scope rather than accuracy: the system faithfully reproduces official publications, and official publications are not neutral records of events. A hearing transcript records what was said in a public session, which is a curated performance; a committee report states a committee majority's account of a statute's purpose. The document is authentic; its contents are advocacy, testimony or law depending on what it is, and the analyst has to know which.
Characteristic false positives
- Bill numbers without a Congress: designations repeat every two years, so a reference to a bill number without the Congress will match the wrong bill and the mismatch looks entirely plausible.
- Annual snapshot mistaken for current law: a Code of Federal Regulations title from this system is the edition as of its revision date, and asserting it as today's requirement can be a year or more wrong.
- OCR gaps in historical material: scanned older documents have imperfect text layers, so a search that returns nothing over a historical period may be finding nothing or may be failing to read it, and the two are indistinguishable from the result.
- Courts collection treated as complete: an absent opinion may mean the court does not participate, the opinion was unpublished, or it exists only in the docket. Concluding no such ruling exists from a null result is unsound.
- Testimony read as fact: statements in hearings are made by witnesses with interests, sometimes under oath and sometimes not, and the authenticated document certifies that the statement was made, not that it is true.
- Committee report treated as law: reports explain what a committee thought a bill should do and carry interpretive weight in some contexts, but they are not the enacted text and are frequently at odds with it.
- Granularity assumptions: code written against a richly segmented collection breaks silently on collections delivered as monolithic documents, usually producing empty results rather than errors.
- Authentication lost in processing: extracting text and storing that as the record discards the digital signature, so the pipeline ends up holding an unverifiable copy of a verifiable document.
None of these make the source unusable. They make it a source that requires corroboration before an assertion built on it goes into a product, which is true of every source and admitted by few.
Ageing
Archived documents do not age; the inferences drawn from them do, and at very different rates by collection. A statute as enacted is permanently what it was, and a court opinion is permanently what the court held on that date, though its precedential force can be destroyed by a later decision that this system will not flag. Codified material ages on a schedule – the Code of Federal Regulations is revised annually by title, so a title's edition is up to a year behind, and a current-state question answered from it is unreliable by construction. Legislative material ages by relevance rather than accuracy: a bill from three Congresses ago is exactly as accurate as it ever was and probably died. The stale-record signature here is subtle. Nothing looks old. An authenticated PDF of a 2015 regulation looks as authoritative as one from last month, because it is equally authentic – it is simply no longer the law. The discipline is to carry the edition or issuance date into every citation and to route present-tense legal questions to continuously updated sources.
What this source feeds
A source is only worth what it lets you conclude. These are the disciplines that collect through it, the mission domains it serves and the data points it yields — every one is a tag, so you can follow any thread from here into the rest of the library.
Collected by these intelligence disciplines
Serves these mission domains
Yields these data points
How each sector uses GovInfo / Official Gazettes
The same dataset is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The records are shared — the constraints, thresholds and outputs are not.
🎖 Military and defence
The useful streams are authorisation and appropriation legislation, committee hearings and reports on defence programmes, and the statutory basis for authorities that operations rely on. Committee reports accompanying defence bills explain congressional intent behind programme decisions and frequently contain detail on force structure, procurement and oversight requirements that is not in the statute. Hearings put officials and contractors on the record about programme status, which is a useful independent check on internal reporting. For legal support to operations, authenticated statute and regulation text is the reference that will not be challenged. Remember that the publication lag on hearings and reports is substantial, that classified sessions leave no record, and that congressional documents describe the legislative view of a programme rather than its actual state.
🕵 National intelligence
For GOVINT this is the authoritative substrate on United States law and legislative process, and it is the right place to establish what an authority actually permits rather than what commentary says it permits. Hearings and reports are a genuine open-source collection target: they place officials, contractors and foreign-linked entities on the record, and the Serial Set and older material support long-horizon research on how a policy area developed. For counterintelligence and economic security, federal court opinions in the courts collection name parties and describe conduct with judicial findings behind them, which is a higher evidentiary grade than reporting. The scope limitation is absolute and worth internalising: this is US material only, so its silence about a foreign matter carries no information, and comparable work on another country requires that country's gazette and law reporting.
👮 Law enforcement
The operational value is evidentiary. When a case requires proving what a regulation or statute said on a particular date, the authenticated PDF from this system is what you put in the file, and it forecloses an argument the defence would otherwise get for free. Court opinions in the courts collection give you prior judicial treatment of similar conduct and, in specific matters, named parties and findings. Hearings occasionally surface conduct by companies and individuals in a form that is useful for lead development. Two practical cautions: the courts collection is not comprehensive so it cannot support a negative finding about case law, and codified material here is an edition, so a compliance-date question needs the continuously updated source and this system for the historical text.
🔍 Private investigation and corporate security
For due diligence, the courts collection and hearing transcripts are the highest-yield parts. A federal court opinion naming a subject is a hard, citable adverse finding, and testimony in a congressional hearing about a company's conduct is frequently the most detailed public account that exists. Statutes and regulations support the compliance dimension of a diligence report. The tradecraft is to search name variants and predecessor entities across collections, to record package and granule identifiers so a client or counsel can verify, and to be explicit that a null result in the courts collection is not evidence of a clean record, because participation is voluntary and coverage is partial. Everything is public domain, so findings can be reproduced in a client report without licensing exposure.
📰 Journalism and OSINT media
Congressional hearings, committee reports and the Congressional Record are underused by reporters who rely on transcripts services and wire coverage. The full record contains submitted statements, inserted materials and exchanges that no summary carries, and the Serial Set and older collections support historical accountability work that competitors will not do. Court opinions give you judicial findings rather than allegations. For any story that turns on what a law or rule actually says, the authenticated document ends the argument and protects the piece. Two habits pay off: cite by package identifier so the reference survives site redesigns, and check the publication lag before treating an absence as significant, because hearings appear months after they happen.
🌍 NGO, humanitarian and human rights
Advocacy organisations use this for legislative history – what Congress said a statute was meant to do, which matters in litigation and in regulatory advocacy – and for the hearing record, where affected communities' testimony is preserved and citable. For human rights and humanitarian work, the statutory and regulatory basis of sanctions, immigration and foreign assistance programmes is here in authenticated form, which matters when the question is exactly what a humanitarian exemption authorises. Court opinions support impact litigation research. The limitation to keep in view is that this is the record of formal federal process, which systematically over-represents institutional voices and under-represents everyone else, and that reading a hearing as a balanced account of a subject is a category error.
🎓 University and research
This is a foundational corpus for political science, law, history and public policy: complete-for-scope, authenticated, freely licensed, with stable identifiers and structured metadata across three branches and, for some series, more than a century. It supports text-as-data work on legislative language, quantitative analysis of legislative and regulatory output, and archival research in the Serial Set and historical hearings. The methodological hazards worth teaching are the uneven historical digitisation, OCR error in scanned material affecting recall in ways that correlate with era, the partial nature of the courts collection, bill number ambiguity across Congresses, and the temptation to treat publication volume as a measure of activity when it is a measure of publishing.
Playbook: working GovInfo / Official Gazettes end to end
A repeatable sequence from first pull to finished product. Each phase states what you are trying to establish, not merely what to click — the objective is a defensible chain of reasoning, not a completed checklist.
Phase 1 — Identify the collection before the query
Decide which publication series would contain the answer – statute, regulation, hearing, report, opinion, presidential document – and go there. Searching the whole corpus for a legal question returns a mixture of series with different authority levels, and sorting them afterwards is slower than choosing first. If you do not know which series holds what you need, read the collection list before writing anything.
Phase 2 — Establish whether you need current law or historical text
This system holds editions and issuances, not continuously current law. If the question is what the rule requires today, use the continuously updated codified sources and come back here for the authenticated historical text. If the question is what it required in 2017, this is exactly the right place and almost the only one.
Phase 3 — Get an API key and learn the package-granule model
Register a key through the federal key service rather than relying on the demonstration key, and spend an hour understanding packages and granules before building. The model is unusual and the whole efficiency of the source depends on retrieving granules rather than whole publications. Integrations that ignore granules end up parsing enormous PDFs for no reason.
Phase 4 — Run the entity sweep across the right collections
For a named company or person, search the courts, hearings, Congressional Record and Federal Register collections with name variants, former names and transliterations. Record package and granule identifiers for every hit. Expect imperfect recall in historical material and say so in the finding rather than implying exhaustiveness.
Phase 5 — Pull the legislative history properly
For a statute, assemble the bill versions, the committee reports, the relevant Congressional Record debate and the public law and Statutes at Large text. Use the related-document links rather than searching for each piece. This is the standard method for establishing legislative intent and it is entirely doable from this one system.
Phase 6 — Work the hearing record rather than the summary
Retrieve the full hearing package including submitted written statements and inserted materials, not just the oral exchange. The written submissions frequently contain the substantive detail – data, correspondence, internal documents – that the oral session only gestures at. Note who submitted what, because that is a map of the interests engaged.
Phase 7 — Check court opinions with an explicit completeness caveat
Search the courts collection for the parties and conduct at issue, and record what you find. Then state plainly in your product that the collection reflects participating courts and does not constitute complete federal case law, so a null result is not a clean record. Where the stakes justify it, go to the court's own docket system.
Phase 8 — Retrieve and preserve the authenticated rendition
For anything that may be relied on evidentially, download the digitally signed PDF and store it unaltered alongside your extracted text. Verify the signature. A pipeline that keeps only the text has thrown away the reason to use this source at all.
Phase 9 — Reconcile against the executive-branch record
Where a matter spans branches – a regulation implementing a statute, an agency action examined in a hearing – hold the Federal Register document, the statute and the hearing record side by side. Discrepancies between what an agency told Congress and what it published are among the most productive findings this corpus supports.
Phase 10 — Date-stamp every citation with its edition
Codified material carries an edition date, legislative material carries a Congress, and court opinions carry a decision date. Every citation in your product should carry the relevant one. This is what allows a reader in two years to know whether your statement was true when you made it.
Phase 11 — Go outside the United States deliberately
If your matter touches another jurisdiction, stop and route to that country's official gazette and law reporting rather than looking for it here. The European Union publishes through EUR-Lex, the United Kingdom through the Gazette and its legislation service, Canada through the Canada Gazette, and most states have an equivalent. Assuming a US system covers foreign law is a beginner's error with expensive consequences.
Phase 12 — Preserve identifiers, not URLs
Store package and granule identifiers and official citations rather than links. Government web architecture changes, and an identifier resolves in any future system while a URL is a bet on the current one. This is a five-minute decision at build time that determines whether your corpus is still usable in five years.
The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.
What to pair it with
No single source carries a finding. These are the datasets that corroborate, extend or contradict this one — and a source that contradicts is worth more than one that agrees, because it is the only thing that will tell you when you are wrong.
| Source | Relationship | What it adds |
|---|---|---|
| Federal Register API | extends | The daily executive branch publication with much better search and no key requirement for the regulatory portion of what GovInfo holds. Use it for monitoring and GovInfo for the authenticated archive. |
| Congress.gov | extends | The legislative information system, with bill status, sponsorship, amendments, votes and current legislative process – the live tracking layer over the documents GovInfo archives. |
| Electronic Code of Federal Regulations | supersedes | Continuously updated codified regulations. Supersedes the annual CFR editions here for any present-tense question about what the law currently requires. |
| PACER and federal court dockets | extends | The complete federal court docket system, including filings and cases whose opinions never reach the GovInfo courts collection. Necessary wherever completeness of case law matters. |
| EUR-Lex | corroborates | The European Union's official legal publication service, covering the Official Journal, treaties, legislation and case law. The European analogue and the correct route for EU legal questions. |
| The Gazette (United Kingdom) | corroborates | The UK official public record, carrying insolvency, corporate, honours and statutory notices with its own structured data interface. A genuinely different gazette tradition worth understanding for comparative work. |
| Canada Gazette | corroborates | Canada's official newspaper of record for regulations and statutory notices, structured similarly to the US Federal Register and useful for cross-border regulatory work. |
| USASpending.gov | extends | Turns appropriations legislation into observed spending, connecting what Congress authorised in documents held here with what was actually obligated. |
| Government Accountability Office | corroborates | Audit and evaluation reports on federal programmes, many held in this corpus, providing independent assessment of what agencies told Congress. |
Legal, ethical and operational constraints
Using this source is legally uncomplicated; the discipline is in how you represent what you retrieved. The corpus is public domain with narrow exceptions for third-party material reproduced within official publications, and there is no access restriction to respect. The substantive risks are three. Misrepresenting an annual edition as current law can produce advice that is materially wrong, with professional consequences for the adviser. Republishing testimony or court material that names individuals engages defamation and data protection considerations outside the United States even though the underlying publication is official – the fact of official publication is a strong but not automatic defence, and accuracy and retention obligations still apply under European rules. And presenting an unsigned derivative as an authenticated document is a misrepresentation with real consequences if it reaches a court. Where material is used evidentially, supply the signed original and be prepared to explain the authentication chain. For sensitive personal information appearing incidentally in older publications, apply the same proportionality reasoning you would to any other personal data regardless of its public status.
Operational security
The API key ties your requests to a registered identity and an email address, which is a meaningful difference from unauthenticated federal services. Query patterns are therefore attributable, and a sequence of searches for a named person or company is a disclosure to a federal system about the subject of your interest. For most research this is irrelevant. For sensitive investigations it is not, and the mitigations are to collect broadly and search locally, to use bulk data rather than targeted queries where feasible, and to think carefully before running a name search that you would not want associated with your organisation. The website is usable without an account for one-off lookups, which shifts the exposure from a registered key to a network path. Note also that key issuance is through a shared federal service, so the identity you register there is visible across the other federal APIs that use it.
Two rules that hold regardless of jurisdiction. Collection that is lawful is not automatically proportionate, and a dataset assembled for one purpose does not carry consent for another. Where the records concern identifiable people, the question is not only whether you may hold the data but whether holding it serves the purpose you are accountable for.
Is it earning its place?
Sources accumulate. Feeds get added during an incident and are never reviewed again, and a decade later the pipeline is carrying dead weight that nobody dares remove. These are the measures that show whether GovInfo / Official Gazettes is contributing anything, and they are worth baselining now so the answer is available later.
- Ratio of granule-level retrievals to package-level retrievals, which measures whether your integration is using the source's actual precision or brute-forcing whole publications.
- Coverage of your watched collections against their publication schedules, tracked separately per collection, because a stalled hearings harvest is invisible in an aggregate figure.
- Number of evidentially relied-upon documents held as verified authenticated PDFs rather than as text extractions – a compliance measure for anything heading toward litigation.
- Entity hits found in hearings and court opinions that were not present in commercial screening sources, which is the clearest measure of what this corpus adds.
- Proportion of legal citations in your products carrying an edition or Congress designation, since undated citations to codified material are the characteristic failure here.
- Rate at which historical searches return null, monitored over time, as an indirect indicator of OCR and digitisation gaps rather than of actual absence.
- API request volume against your key limit, watched as a leading indicator that a collection strategy needs to move to bulk data before it fails.
Beware of volume. Indicator counts rise easily and say almost nothing. Unique contribution — findings this source produced that no other source in your stack would have — is the measure that matters, and it is usually far lower than anyone expects.
Tradecraft notes
The distinctions that separate a competent analyst from a fast one:
- Learn the package and granule model properly before writing code against it. It is the difference between retrieving a single court opinion and downloading a thousand-page volume to find one, and almost every performance problem with this API comes from ignoring it.
- Never cite codified material without its edition date. The annual snapshot looks identical to current law and is not, and this single omission is responsible for more bad regulatory advice than any other habit.
- A bill number is meaningless without its Congress. Carry the Congress number in every legislative reference and reject inbound references that lack it rather than guessing.
- Search the written submissions in hearings, not only the transcript. The substance – internal documents, data, correspondence – is usually in the appended material that summaries and news coverage never reach.
- Treat a null result in the courts collection as uninformative and say so explicitly. Voluntary participation means absence carries no weight, and an investigator who implies otherwise is overselling.
- Keep the signed PDF. Extraction is for analysis; the signed original is for evidence, and a pipeline that discards it has thrown away the one thing this source has that mirrors do not.
- Assume OCR failure in older material and compensate with variant spellings, partial-string searches and manual review of adjacent pages. Historical recall here is real but it is not free.
- Store identifiers rather than links. Package and granule identifiers survive site rearchitecture; URLs do not, and the corpus has moved systems before.
- Know when to leave. This is a United States federal system, and the instinct to keep digging here for a foreign legal question wastes hours that the relevant national gazette would answer in minutes.
Questions analysts actually ask
Does GovInfo contain gazettes from other countries?
No. Despite the catalogue heading, GovInfo is exclusively United States federal material across the three branches. For other jurisdictions you need that country's official publication service – EUR-Lex for the European Union, the Gazette for the United Kingdom, the Canada Gazette for Canada, and national equivalents elsewhere. Assuming otherwise is a common and costly mistake in cross-border work.
What does authentication actually give me?
A digital signature and visible seal on the PDF certifying that the document has not been altered since GPO published it on behalf of the issuing authority. Practically, it means a court or regulator will accept the document without argument about provenance, and it means a copy you extract text from and re-host does not carry the same weight. Keep the signed file, not just the text.
Do I need an API key and does it cost anything?
Yes and no respectively. Keys are free and issued through the shared federal API key service in a few minutes. A demonstration key exists for evaluation but has limits low enough that it will not support real work. The key ties requests to a registered identity, which is worth remembering for sensitive research.
Is the courts collection complete federal case law?
No. It contains opinions from participating federal courts, and participation is voluntary and uneven across districts and circuits. It is a large, useful and free sample. It cannot support a negative finding, and for comprehensive case law research you need the federal docket system or a commercial legal database. Say which you used in any product that makes a claim about case law.
Should I use GovInfo or the Federal Register API for regulatory work?
Both, for different jobs. The Federal Register API is faster, keyless, has better search and is the right monitoring layer for daily regulatory activity. GovInfo is the authenticated archive, holds the codified Code of Federal Regulations editions and the pre-1994 material, and is what you cite when the document must be beyond dispute. Monitor with one, cite from the other.
How current is the Code of Federal Regulations here?
It is published as annual editions revised on a staggered schedule by title, so a given title can be up to about a year behind current law. For any present-tense compliance question use the continuously updated electronic Code. Use the editions here when you need to establish what a regulation said on a past date, which they do better than anything else.
Why is search returning nothing for a historical name I know appears?
Most likely optical character recognition failure in the scanned text layer, which is common in older material and correlates with era and print quality. Try variant spellings, shorter distinctive substrings and browsing by date around the expected document rather than searching. Assume historical recall is imperfect and never treat a null result over old material as evidence of absence.
Can I harvest the whole corpus through the API?
You can try and you will hit rate limits long before you finish. Use the bulk data route for collections available that way, particularly the structured legislative XML, and reserve the API for incremental updates and targeted retrieval. This is not just etiquette – it is the only approach that completes, and it is what the bulk service exists for.
What is the difference between a package and a granule?
A package is a publication unit – one day's Congressional Record, one bill version, one CFR title, one court opinion package. A granule is an addressable component within it – a single speech, a single statute section, an individual opinion. Analysis happens at granule level and retrieval at granule level is far cheaper. Building an integration that only understands packages is the most common design mistake with this API.
Standards, formats and interoperability
What this source speaks natively, and what it has to be translated into before a partner can consume it. Work that arrives in a recognised format is easier to defend, easier to hand over and easier to automate against:
- Digital signature and PKI-based document authentication, which is what distinguishes the version of record from a copy and is verifiable in standard PDF readers.
- Package and granule identifier schemes, stable across the system's lifetime and the correct thing to store in place of URLs.
- Legislative XML schemas used for bills, statutes and the Congressional Record, which make the legislative corpus genuinely machine-analysable rather than merely machine-readable.
- Superintendent of Documents classification, the long-standing library scheme for US government publications, linking digital records to physical depository holdings.
- Official legal citation conventions – Federal Register volume and page, United States Code title and section, Statutes at Large, public law numbers – which are what make a reference permanent.
- Metadata standards for archival description and preservation applied across collections, which is why identifiers and provenance are consistent where most corpora are not.
- The federal shared API key service, which governs authentication and rate limiting across this and other government APIs.
References
Primary documentation and authoritative references for this source. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.
- GovInfo — US Government Publishing Office. The system itself, with browse by collection, search and the authenticated documents. The fastest route for any one-off lookup and the place to understand what collections exist.
- GovInfo API documentation — US Government Publishing Office. The interactive API reference covering collections, packages, granules, published-date listings and search. Read the package-granule model here before designing an integration.
- GPO API GitHub repository — US Government Publishing Office. The public repository where GPO documents API changes and users raise issues. The most reliable place to confirm current behaviour and to see what has recently broken or changed.
- api.data.gov key signup — US General Services Administration. Where you obtain the free key required for the API, shared across several federal services. Note that the identity you register is visible to all of them.
- Government Publishing Office — US Government Publishing Office. The agency behind the system, including the Federal Depository Library Program and the institutional context for why authentication and permanent access are designed in.
- Congress.gov — Library of Congress. Legislative status, sponsorship, votes and process. The live tracking layer that complements GovInfo's archival documents, and the right place to see where a bill currently stands.
- Electronic Code of Federal Regulations — Office of the Federal Register / GPO. Continuously updated codified regulations, which supersede this system's annual editions for any question about what the law requires now.
- Federal Register — Office of the Federal Register. The daily executive branch publication with faster search and no key. Use it to monitor; use GovInfo for the authenticated archival copy.
- United States Courts — Administrative Office of the US Courts. The federal judiciary's own site, including access to the docket system that holds the case material the opinions collection only partially reflects.
- EUR-Lex — Publications Office of the European Union. The European Union's official legal publication service. The correct destination when the legal question is European rather than American, and a useful comparison of gazette design.
- The Gazette — The Stationery Office / UK Government. The United Kingdom's official public record, carrying insolvency, corporate and statutory notices with structured data access. A different and instructive model of what a gazette can publish.
- Canada Gazette — Government of Canada. Canada's official newspaper of record for regulations and statutory notices, useful for cross-border regulatory comparison and for matters spanning the two countries.
Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.
Put it into practice
The Quantus Intel threat intelligence platform operationalises this source: it registers each GovInfo collection as its own source, harvests at package level and analyses at granule level, preserves the authenticated PDF alongside extracted text, and carries the edition date into every legal citation it exports.. Browse the full source catalogue, or follow any tag above into the rest of the library.