Onion / Hidden Service: Data Point Intelligence Guide
An onion address is a public key, not a location. Everything you can learn about a hidden service comes from what its operator publishes, misconfigures or reuses.
An onion address is a public key, not a location. Everything you can learn about a hidden service comes from what its operator publishes, misconfigures or reuses.
Understanding the Onion / Hidden Service as an intelligence artifact
An onion address is the identity of a Tor hidden service, derived directly from the service's public key. Version 3 addresses are 56 base32 characters plus the onion suffix, encoding an ed25519 public key, a checksum and a version byte. Because the address is the key, it cannot be transferred, spoofed or seized in the way a domain name can, and it resolves through the Tor network's distributed hash table rather than through DNS. There is no registrar, no WHOIS and no authoritative directory.
Version 2 addresses were sixteen characters and have been deprecated and removed from the network, so any v2 reference in older reporting is now dead. Vanity addresses are generated by brute force to produce a readable prefix, which is itself an operator signal. Onion sites frequently publish signed mirror lists to counter phishing clones, and those signatures are a practical authentication mechanism for analysts confirming they are on the real service.
Why it matters
Hidden services host the operational surface of a large part of organised cybercrime: ransomware leak sites, marketplaces, forums, credential shops and command and control. The address is the stable identifier for a criminal service across content changes and mirror rotation. It anchors monitoring, victim identification from leak site postings, and long-running actor tracking. Operator mistakes around hidden services are also one of the most productive de-anonymisation avenues available to lawful investigators.
What analysts actually look for
These are the concrete, observable signals that carry weight in this area of work:
- Address version and checksum validity, which immediately filters typos, deprecated v2 references and fabricated addresses in reporting.
- Vanity prefixes, which indicate deliberate brand building and can be correlated with the same operator's other services.
- TLS certificates and self-signed certificates presented by the service, which sometimes leak the operator's clearnet hostname.
- Server banners, error pages, default pages and software version strings that fingerprint the hosting stack and its configuration.
- Content fingerprints such as favicon hashes, CSS assets, template structure and analytics identifiers that link mirrors and clearnet twins.
- Cryptocurrency addresses, contact handles, PGP keys and Jabber or session identifiers published on the service itself.
- Uptime and downtime patterns, which correlate with operator timezone, hosting problems and law enforcement action.
- Cross-signed mirror lists and PGP-signed announcements, which authenticate the genuine service against clone and phishing copies.
Where the data comes from
Authoritative and openly available collection points. Always confirm licensing and terms before operational or commercial use:
- Ahmia — Open search engine indexing Tor hidden services with a published blacklist policy and API access
- Tor Project Metrics and Relay Search — Network-level statistics, relay and bridge data, and historical consensus information
- dark.fail — Maintained list of PGP-verified onion mirrors used to distinguish genuine services from phishing clones
- Ransomware.live — Aggregated ransomware leak site postings with victim listings and onion addresses per group
- ThreatFox (abuse.ch) — Community indicator feed that includes onion command and control addresses linked to malware families
- Shodan and Censys — Clearnet scan data for locating misconfigured hosts serving the same content or certificates as a hidden service
- VirusTotal — Onion URLs referenced in submitted samples, linking hidden services to specific malware and campaigns
A working method
A repeatable sequence beats ad-hoc searching. This is a practical starting workflow:
- Validate the address — Check length, character set and checksum, and confirm the version before spending any time on collection or attribution.
- Confirm authenticity — Verify against PGP-signed mirror lists or the service's own announcements, since phishing clones of criminal services are extremely common.
- Collect safely — Access through isolated, non-attributable infrastructure with scripting disabled, and archive content with timestamps and hashes as you go.
- Extract on-page identifiers — Pull cryptocurrency addresses, PGP keys, contact handles, email addresses and usernames, which are the artifacts that actually pivot.
- Fingerprint the service — Record certificates, favicon hashes, headers, asset paths and template structure to enable correlation with other services and clearnet hosts.
- Hunt for clearnet leakage — Search scan data for hosts presenting identical certificates, favicons or content, which is the classic hidden service configuration failure.
- Monitor over time — Track uptime, content changes and new victim postings on a schedule, since the timeline of a leak site is intelligence in itself.
How this connects across the intelligence taxonomy
Intelligence work does not respect neat boundaries. The mission domain you are working, the disciplines you practise, and the data points you pivot on are one connected system. These are the direct relationships for this entry — every link is also a tag, so you can follow any thread across the whole library.
Collected by these disciplines
- Criminal Intelligence — Intelligence Supporting Criminal Investigation
- Dark Web Intelligence — Hidden Services and Closed Criminal Venues
- Threat Actor Intelligence — Tracking Adversary Groups Over Time
- Cryptocurrency Intelligence — Tracing Value on Public Ledgers
- Human Intelligence — Information from People, Ethically Obtained
- Cyber Intelligence — Adversary Activity in Networks and Systems
- Financial Intelligence — Following Value Through the Financial System
- Social Media Intelligence — Intelligence from Social Platforms and Networks
Investigated in these domains
Pivots to these data points
- File Hash — Cryptographic fingerprint of a file, used for malware identification.
- CVE / Vulnerability — Common Vulnerabilities and Exposures identifier for a known flaw.
- SSL/TLS Certificate — A digital certificate binding a public key to an identity.
- Malware Family — A named class of related malicious software.
- File / Document — A file or document artifact — malware sample, leaked document, image, or email attachment.
- Data Breach — A known data breach or leak incident with exposed records.
Inside the platform: where Onion / Hidden Service lives
The Quantus platform is 204 pages behind a 147-item sidebar organised into six working groups: Command (24 items), Dashboards (15), Threat Theaters (14), Intelligence Domains (15), Investigate (34), and Administration (45). This entry is not a page in isolation — it is a thread running through several of them.
The modules that matter most here:
ioc-type.php?t=onion— Onion / Hidden Service profiledatapoint.php?dp=dp_onion— Data point hubioc.php— Cyber Crime dashboardransomware.php— Ransomware dashboarddomain.php?d=darkweb— Dark Web Intel dashboarddomain.php?d=drugs— Drug Trafficking dashboardsearch.php— Advanced search, filter and pivotcorrelate.php— Correlation graphcases.php— Case management
Each dashboard is local-first: it renders from the platform’s own database rather than depending on a live third-party call, so it still works when an upstream API is unreachable or rate-limited. Heavy aggregates are cached with a hard query time cap and degrade to the last good value instead of hanging the page.
Automation, playbooks and AI skills
Analysis that only happens when someone remembers to run it is not a capability. The platform ships a 30-step automation pipeline (cron.php) that collects, ingests, resolves, enriches, correlates and scores on a schedule — 25 seeders, 11 resolvers and 7 enrichment runners, all idempotent and cursor-based so a run can be interrupted and resumed without duplicating or losing work.
AI skills that apply
The 16 one-click operations in ai-skills.php are deterministic jobs, not free-text generation. The ones that matter here:
- Enrichment Runner
- Enrichment → Local
- Correlate Infrastructure
- Summarise (Copilot)
- Generate Report
Alerting closes the loop: rules in alerts.php fire on new indicators matching a saved query, so a first sighting in this area raises a notification rather than waiting to be noticed at the next review.
Feeds, data sources and the API
The collection layer runs a feed registry of free, machine-readable sources — bulk blocklists and trackers (Maltrail, IPsum, FireHOL, the full abuse.ch corpora, phishing databases, Emerging Threats, Spamhaus, DigitalSide, ThreatView), authoritative government feeds (CISA KEV, OFAC, UN and EU sanctions lists), and reference datasets (RIR allocations, ip-to-ASN and geolocation tables, MITRE ATT&CK, EPSS). collect.php pulls them server-side on a schedule; feeds.php and source-catalog.php show what is registered, what it covers and when it last ran.
Anything the platform holds is reachable programmatically. The REST API in api.php exposes 11 endpoints — status, stats, search, lookup, recent, export, bulk_check, top_threats, by_category, categories, check — and export.php streams 18 formats in bounded chunks, so a million-row export neither exhausts memory nor times out:
STIX 2.1, MISP, OpenIOC 1.1, CEF (ArcSight), LEEF 2.0 (QRadar), Zeek/Bro intel, Snort/Suricata rules, Palo Alto EDL, BIND RPZ, hosts blackhole, iptables, CSV, JSON, NDJSON/JSONL, XML.
That covers the CTI standards (STIX 2.1, MISP, OpenIOC), SIEM ingestion (CEF, LEEF, Zeek), detection engines (Snort/Suricata), and direct enforcement (Palo Alto EDL, BIND RPZ, hosts, iptables) — so intelligence developed here can be actioned in the tools you already run, without a manual reformatting step. A TAXII 2.1 server and a MISP/RSS feed are also served for pull-based sharing.
Use cases
Three ways this entry earns its keep in day-to-day work:
- Triage under time pressure. An artifact or report lands and you need a defensible read in minutes, not days. Validate the address is the first move; the platform pre-computes the enrichment so the analyst spends the time on judgement rather than lookups.
- Building the picture. A single indicator is rarely the story. Collect safely turns one artifact into a network — shared infrastructure, repeated selectors, the same operator behind different names — via the correlation graph and the cross-entity link engine.
- Producing something actionable. Analysis that ends in a document nobody can use is wasted. Monitor over time feeds the case file, the detection rule, the block list or the referral — with sourcing attached so the recipient can verify it.
Case management (cases.php), watchlists, saved searches and scheduled reports mean the work persists between sessions and survives an analyst leaving the team.
How each sector uses Onion / Hidden Service
The same entry is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The underlying artifacts are shared — the constraints, outputs and thresholds are not.
🎖 Military and defence
Hidden service monitoring supports force protection and defensive planning rather than targeting. Ransomware leak sites publish victims including defence suppliers, and early sight of a posting naming a contractor in your supply chain drives an immediate assessment of what capability is at risk. Criminal marketplaces trade access to networks and stolen credentials that include military-adjacent organisations. Collection must run through authorised, non-attributable infrastructure under a written policy, because accessing hidden services from mission networks is both an operational security failure and potentially a policy breach. Products feed supply chain risk assessments, counter-intelligence referrals and network defence advisories, with the address treated as the stable identifier across content changes.
🕵 National intelligence
For national intelligence, onion addresses are stable identifiers for criminal and hostile services whose content, mirrors and clearnet fronts rotate constantly. Requirements-driven collection tracks named ransomware operations, marketplaces and forums against priority questions such as targeting of national infrastructure and proliferation of specific capabilities. Fusion combines on-page artifacts including cryptocurrency addresses, PGP keys and handles with financial, infrastructure and human reporting. Collection tradecraft is itself sensitive: the fact and method of monitoring, and any non-attributable infrastructure used, will usually be classified above the content. Products should state collection date and archive hashes, because hidden service content is ephemeral and unverifiable after the fact.
👮 Law enforcement
In law enforcement, the onion address anchors a case that may run for years across content changes and mirror rotation. Investigators preserve postings with timestamps and hashes, identify victims from leak site listings for notification, and pursue de-anonymisation through lawful means: operator configuration errors, financial tracing, undercover engagement under authorisation, and international cooperation. Legal process is essential and varies by jurisdiction. Purchasing goods, downloading certain categories of material, or engaging operators requires specific authorisation, and unauthorised action can taint an entire prosecution. Record every access with operator, time, method and the authority relied upon, and preserve collection in an evidentially sound archive from the start.
🔍 Private investigation and corporate security
Corporate security uses hidden service monitoring for breach discovery, executive protection and supply chain risk. Finding your organisation or a supplier listed on a leak site is frequently the first indication of a breach and starts the regulatory clock. What a private actor may not lawfully do is extensive: purchasing data, engaging operators, negotiating on behalf of a client without proper legal structure, downloading stolen data, or accessing non-public areas of a service. Those actions can constitute handling stolen goods, unauthorised access or sanctions violations. Restrict activity to passive viewing of public pages through isolated infrastructure under written policy, and route anything beyond that to counsel and law enforcement.
📰 Journalism and OSINT media
Reporting on hidden services demands verification discipline and source protection. Confirm you are on the genuine service rather than a phishing clone by checking PGP-signed mirror lists and the operator's own announcements, since clones of criminal services are extremely common and reporting from one produces a false story. Archive with timestamps and hashes because the content will change. Consider carefully whether to name victims listed on leak sites, since publication can compound the harm to organisations and individuals whose data was stolen. Never pay for access or data. Protect your own operational security, since an operator who identifies a journalist may retaliate against them or their sources.
🌍 NGO, humanitarian and human rights
Civil society organisations encounter hidden services in two ways: as victims whose data appears on leak sites, and as researchers documenting trafficking, exploitation and repression. Both require a victim-centred approach. Where personal data of vulnerable people is published, the priority is notification, support and mitigation rather than collection, and any documentation must minimise further circulation of the material. Staff exposed to abusive content need psychological support and rotation, which should be planned rather than improvised. Access through isolated infrastructure under a written policy, take legal advice before any collection involving illegal content, and coordinate with law enforcement through an established channel rather than acting alone.
🎓 University and research
Research on hidden services covers marketplace economics, ransomware ecosystems, censorship circumvention and network measurement. Ethics review is mandatory and non-trivial: research may involve illegal content, may require interaction with criminal actors, and can affect the safety of Tor users who rely on the network for protection. Measurement studies must avoid techniques that degrade user anonymity, and the Tor Project has published research safety guidance for exactly this reason. Publish methodology, collection dates and crawl parameters, since hidden service availability changes hourly and results are otherwise unreproducible. Aggregate and anonymise before publication, and do not republish victim data recovered from leak sites.
Playbook: working Onion / Hidden Service end to end
A repeatable sequence, from the moment the requirement lands to the moment a product is delivered and the case is closed out. Each phase states what you are trying to establish, not merely what to click — the point is a defensible chain of reasoning, not a checklist.
Phase 1 — Authorise before you collect
Establish written authority for hidden service access covering what may be viewed, what may not be downloaded, who may operate the capability and what triggers escalation to legal or law enforcement. Passive viewing of public pages is generally lawful; purchasing, downloading certain material and engaging operators frequently is not. A good output is a signed policy plus a per-operation authorisation record. Stop and escalate immediately on encountering child sexual abuse material, which requires a defined reporting route and must not be collected.
Phase 2 — Validate the address
Check that the address is fifty-six base32 characters plus the suffix, verify the embedded checksum and version byte, and reject anything matching the deprecated sixteen-character v2 format, which no longer resolves. Malformed or v2 addresses in older reporting are dead references and pursuing them wastes time. A good output is a validated address record with the validation method noted. Vanity prefixes are legitimate but indicate deliberate generation effort, which is itself a small operator signal worth recording.
Phase 3 — Confirm you are on the genuine service
Verify against PGP-signed mirror lists, the operator's own announcements on forums they control, and consistency of on-page artifacts such as PGP key fingerprints and cryptocurrency addresses across mirrors. Phishing clones of marketplaces and leak sites are pervasive and are built specifically to catch researchers and buyers. A good output is a documented authentication decision with the evidence. Stop and treat everything collected as suspect if the signature verification fails, because reporting from a clone is a false story.
Phase 4 — Collect through isolated infrastructure
Access through non-attributable, isolated infrastructure with scripting disabled, no persistent identity, and no route to organisational networks. Assume the operator logs and fingerprints visitors. Disable anything that could leak an identifier, and never reuse the collection environment for other work. A good output is a documented collection configuration that another operator could reproduce. Stop and rebuild the environment if any misconfiguration is discovered, rather than assuming a single exposure was harmless.
Phase 5 — Archive with integrity
Capture full page content, headers, screenshots and any linked artifacts, hash each capture and record the collection timestamp and operator. Hidden service content is ephemeral, and a claim you cannot substantiate later is worthless in both reporting and evidence. A good output is an archive that another analyst can verify was not altered after collection. Build this into the collection tooling rather than relying on analysts to remember, because it is always skipped under time pressure.
Phase 6 — Extract on-page identifiers
Pull every cryptocurrency address, PGP key and fingerprint, contact handle, email address, session and messaging identifier, and username, since these are the artifacts that actually pivot into other datasets. PGP key fingerprints and reused handles link operators across services and across years far more reliably than site content does. A good output is a structured identifier record per service with the page and date each came from. These are the leads that survive after the service disappears.
Phase 7 — Fingerprint the service technically
Record TLS certificates where presented, favicon hashes, HTTP headers, server banners, asset paths, template structure, error page text and any distinctive markup. These fingerprints enable correlation between hidden services and between a hidden service and clearnet hosts. A good output is a technical fingerprint record suitable for searching scan datasets. This is the material that supports the classic operator configuration failure discovery, and it costs almost nothing to collect at the time.
Phase 8 — Hunt for clearnet leakage
Search internet-wide scan data for hosts presenting the same certificate, favicon hash, distinctive content or misconfigured status page, which is the most common lawful de-anonymisation avenue and results from operators exposing the origin server directly. Record any candidate with the evidence and its strength rather than asserting a match. A good output is a candidate list with confidence ratings. Stop short of interacting with any candidate host, and route findings to law enforcement where the target is criminal infrastructure.
Phase 9 — Monitor on a schedule
Track uptime, content changes, new postings and mirror rotation on a defined cadence, because the timeline of a leak site is intelligence in itself: posting frequency indicates operational tempo, gaps indicate disruption, and rebranding indicates law enforcement pressure or internal splits. A good output is a time series per service with change detection and alerting on new victim postings. Automate the cadence, since manual monitoring degrades within weeks and the gaps are never where you expect.
Phase 10 — Handle victim data responsibly
Where postings name victims, notify affected parties through the appropriate route and treat published personal data as personal data with full obligations, not as open source material free to redistribute. Record the victim listing and date for tracking, but do not download the leaked corpus without a lawful basis and a specific purpose. A good output is a victim notification action with a record of what was and was not collected. Stop and take advice before ingesting any leaked dataset.
Phase 11 — Correlate across the ecosystem
Link services through shared PGP keys, cryptocurrency addresses, handles, template reuse, hosting fingerprints and cross-advertisement, then align that with malware family and infrastructure findings. Operators run multiple services and rebrand under pressure, so ecosystem-level correlation is what maintains continuity of tracking. A good output is a linked entity graph with the evidence for each edge recorded. Be explicit about which links are strong technical evidence and which are inference from behaviour.
Phase 12 — Report with collection provenance
State the address, the collection date, the archive hash, the authentication method used to confirm the genuine service, and the authority under which collection occurred. Withhold operational detail about your collection infrastructure. Where reporting names victims, apply an explicit decision about harm and notification. A good output is a product whose factual claims about a hidden service can be substantiated from your archive months later, when the service itself no longer exists.
The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.
Source register: what to collect from, and how
Sources are listed with their access model so you can plan around cost and licensing before you build a dependency on them. Open means no account required; registration means a free account or API key; licensed means paid or institutional access. Always confirm current terms — licensing changes, and a source that was free for research may not be free for commercial or evidential use.
| Source | Access | What it gives you | How it is used here |
|---|---|---|---|
| Ahmia | Open | Open search engine indexing Tor hidden services, with a published content policy and programmatic access to its index. | Discovery of hidden services by keyword and confirmation that an address is indexed and reachable. |
| Tor Metrics | Open | Official network statistics covering relays, bridges, users, onion service counts and historical consensus data. | Provides network-level context and authoritative statistics for reporting on hidden service ecosystem scale. |
| Tor Project documentation | Open | Specifications and guidance for onion service addressing, key derivation, client authorisation and operator security. | The authoritative reference for how v3 addresses encode keys and why they cannot be seized like domains. |
| Ransomware.live | Open | Aggregated ransomware leak site postings with victim names, dates, group attribution and onion addresses per operation. | Monitors leak sites for supplier and organisational exposure without direct collection from each site. |
| ThreatFox | Open | abuse.ch community indicator feed including onion command and control addresses linked to malware families. | Checks whether an onion address is already associated with a known malware family's infrastructure. |
| Censys | Registration | Internet-wide scan data covering certificates, favicons, headers and service fingerprints for clearnet hosts. | Hunts for clearnet hosts presenting the same fingerprints as a hidden service, indicating origin exposure. |
| Shodan | Registration | Host-centric scan data with banners, favicon hashes and historical observations across internet-facing services. | Second scan perspective for identifying misconfigured origin servers behind hidden services. |
| OnionScan concept and successors | Open | Open source tooling and methodology for identifying operational security failures in hidden service configuration. | Reference methodology for the classes of misconfiguration that expose hidden service operators. |
| Internet Archive and archive.today | Open | Public web archiving services preserving snapshots of clearnet pages, including mirrors of criminal service announcements. | Preserves clearnet-side references, forum posts and announcements that corroborate a hidden service's history. |
| Blockchair | Open | Multi-chain blockchain explorer with address search, filtering and export across major cryptocurrencies. | Pivots from cryptocurrency addresses published on a hidden service into transaction and cluster analysis. |
| Chainabuse | Open | Community reporting platform linking cryptocurrency addresses to scams, extortion and criminal services with narrative detail. | Corroborates that an address found on a hidden service is already reported in victim complaints. |
| Europol publications | Open | Law enforcement reporting on darkweb marketplace takedowns, coordinated operations and criminal service ecosystems. | Establishes whether a service has been seized or disrupted, which changes the meaning of continued activity. |
| CISA advisories | Open | Government advisories on ransomware operations including leak site infrastructure, indicators and mitigation guidance. | Authoritative reference linking an onion leak site to a named ransomware operation and its tradecraft. |
| Have I Been Pwned | Open | Verified breach catalogue with domain search, used to assess exposure without handling leaked corpora directly. | Assesses whether data claimed on a leak site has entered wider circulation, without ingesting the dataset. |
Prefer sources that publish a methodology and a revision history. A dataset that changes silently is a liability in any product that has to survive challenge.
Tooling
Tools commonly used against Onion / Hidden Service. None of these replace judgement, and each carries its own failure modes — know what a tool infers versus what it observes.
- Tor Browser — Reference client for accessing hidden services with a hardened, uniform configuration. Limitation: convenience features and misconfiguration can still leak identifying behaviour.
- Isolated collection virtual machine — Disposable environment with no persistent identity and no route to production. Limitation: only effective if genuinely rebuilt between operations rather than reused.
- Headless crawler with archiving — Automates scheduled collection with content hashing and change detection. Limitation: crawling patterns are detectable and may get the collector blocked or fingerprinted.
- PGP verification tooling — Validates signed mirror lists and operator announcements to confirm service authenticity. Limitation: only works where the operator actually publishes signatures.
- Favicon and content hashing — Produces fingerprints that can be searched against clearnet scan datasets. Limitation: common frameworks share favicons, producing large false positive sets.
- Scan dataset search — Finds clearnet hosts matching hidden service fingerprints, exposing misconfigured origins. Limitation: scan cadence means the exposure window may have closed before collection.
- Blockchain explorers — Trace cryptocurrency addresses extracted from hidden service pages. Limitation: address to entity clustering requires heuristics that must be stated explicitly.
- Evidential archive system — Stores captures with hashes, timestamps and operator attribution for later verification. Limitation: worthless unless the hashing happens at collection time, not afterwards.
AI skills and automation in detail
These are deterministic jobs with defined inputs and outputs, not open-ended prompting. Each is idempotent and cursor-based: interrupt one and it resumes where it stopped rather than duplicating work or losing progress.
- Enrichment Runner — Walks the indicator set through a chosen provider in time-boxed, cursor-based batches that resume rather than restart.
- Enrichment → Local — Materialises enrichment into the local store so dashboards render from your own database instead of a live third-party call.
- Correlate Infrastructure — Builds the cross-entity link graph: shared hosting, reused certificates, overlapping registrants, repeated selectors.
- Summarise (Copilot) — Produces a narrative summary beside the underlying records. It explains; it never creates indicators or assigns attribution.
- Generate Report — Assembles a sourced product from the current case or query, with provenance attached to each element.
A note on the boundary: the only skill that involves a language model is Summarise (Copilot), and it writes prose about records that already exist. Nothing else on this list involves generation of any kind. No indicator, relationship or attribution in the platform originates from a model. See the full skill list.
Tradecraft notes
The distinctions that separate a competent analyst from a fast one:
- The address is the stable identifier and everything else rotates. Content, mirrors, branding and even the criminal operation's name change, so build tracking on the address, the PGP key fingerprint and the cryptocurrency addresses rather than on the site name.
- Phishing clones of criminal services are pervasive and are built to catch researchers as well as buyers. Verifying against a signed mirror list before collecting is the difference between reporting on an operation and reporting on someone impersonating it.
- Leak site victim lists are operator-curated marketing, not a victim census. Organisations that paid, that were never listed, or that were quietly removed are invisible, so any scale claim derived from postings is a lower bound with a selection bias that should be stated.
- On-page artifacts outlive the page. PGP key fingerprints, cryptocurrency addresses and reused handles link operators across services and across years, whereas site content and design are replaced on every rebrand, so extract identifiers on the first visit.
- Clearnet leakage through a shared certificate, favicon or misconfigured status page is the classic lawful de-anonymisation route, and it results from operator error rather than from any weakness in Tor. Collect the fingerprints that make that search possible every time.
- Posting cadence is intelligence. A leak site's rhythm reveals operational tempo, and a sudden gap or a rebrand often precedes or follows law enforcement action, internal disputes or affiliate migration, which is visible only if you have been monitoring on a schedule.
- Archive with hashes at the moment of collection. Hidden service content cannot be re-verified later, so an unhashed screenshot is an assertion rather than evidence, and this is the step that is always skipped when the analyst is in a hurry.
- Plan for exposure to abusive content before it happens. Analyst welfare, rotation and psychological support are operational requirements in this area, and teams that treat them as optional lose experienced people and make worse decisions under stress.
Measuring whether it is working
Capability claims should be falsifiable. These are the measures that show whether work on Onion / Hidden Service is producing anything, and they are worth baselining before you change process or tooling.
- Time from a leak site posting naming the organisation or a supplier to internal notification and assessment, measured in hours. This is the metric that determines whether monitoring provides warning.
- Proportion of collected captures with a hash and timestamp recorded at collection time, which determines whether the archive can support any later claim.
- Percentage of monitored services where authenticity was verified against signed mirror lists before collection, guarding against reporting from clones.
- Number of pivots from on-page identifiers into other datasets per monitored service, which measures whether collection is generating leads rather than accumulating screenshots.
- Uptime and coverage of scheduled monitoring against the target service list, since gaps in collection are invisible in the output but destroy timeline analysis.
- Count of collection operations conducted outside written authorisation. The target is zero, and any occurrence is a governance incident rather than a performance issue.
- Analyst rotation and welfare check completion for staff exposed to abusive content, tracked as an operational rather than an administrative measure.
Beware of measuring volume alone. Indicator counts and report counts rise easily and say little; time-to-attribution, proportion of findings that survive review, and how often a product changed a decision say a great deal.
Common pitfalls
- Onion addresses are not resolvable to an IP by design, and any tool claiming to geolocate a hidden service directly should be distrusted.
- Phishing clones of marketplaces and leak sites are pervasive, so unverified content may be an entirely fabricated copy of the real service.
- Content on criminal leak sites is adversary-authored marketing and may include invented, recycled or exaggerated victim claims.
- Accessing certain categories of hidden service content is a criminal offence irrespective of investigative intent, and requires proper authority.
- Downloading files from hidden services risks both malware and unlawful material entering your environment and your evidence store.
- Services rotate addresses frequently, so a dead address does not indicate the operator is gone, only that this identity was retired.
Legal and ethical considerations
Dark web collection needs an explicit policy and, in many organisations, prior authorisation. Passive viewing of public hidden services is generally lawful, but purchasing, downloading certain material, engaging operators or attempting to access non-public areas can constitute criminal offences and must not proceed without legal sign-off and, where relevant, law enforcement involvement. Victim data appearing on leak sites remains personal data with data protection obligations. Preserve collection with hashes, timestamps and operator notes so the record is defensible.
Data integrity: no fabrication, no drift, no hallucination
Intelligence that cannot be traced back to a source is not intelligence, it is assertion. Everything in this entry — and everything in the platform behind it — is built on a small number of non-negotiable rules.
Provenance on every record
Every indicator carries the source that supplied it, a first-seen and last-seen timestamp, and a sighting count. Where several feeds report the same artifact, each contribution is recorded separately rather than collapsed, so you can see whether a finding rests on one source or twelve. Source attribution travels with the data into every export, so a recipient can audit a claim without asking you for the working.
Nothing is invented to fill a gap
If the platform has no data for Onion / Hidden Service, it says so. Empty is displayed as empty — never padded with plausible-looking placeholder values, sample records or illustrative examples that a reader might mistake for observations. A dashboard with no rows is a true statement about collection coverage, and it is treated as a gap to close, not a blemish to hide.
Scoring is deterministic and reproducible
Threat scores, reputation grades and risk tiers are computed from stated inputs with fixed weights, not estimated. The same inputs always produce the same output, and the formula is visible rather than a black box. Aggregates are cached with an explicit time-to-live so a figure on screen is never silently stale — and when a heavy query exceeds its time budget the platform serves the last known-good value and labels it, rather than inventing a fresh number or hanging.
Where AI is used, and where it is not
Language models summarise and explain. They do not create indicators, assign attribution or manufacture relationships. No IP address, wallet, hash or identity in the platform originates from a model — every one is ingested from a named feed, resolved from a reference dataset, or entered by an analyst with a source recorded. Copilot output is presented as narrative alongside the underlying records, never in place of them, so a reader can always check the summary against the evidence.
Guarding against drift
Enrichment is additive and timestamped rather than overwriting. Reference data — sanctions lists, allocations, taxonomies — is re-synchronised from the authority on a schedule instead of being edited in place, so local copies cannot quietly diverge from the source of truth. Attribution is recorded with a confidence level and the reporting it rests on, and inferred relationships are labelled as inferred. When a source retracts or corrects, the correction propagates rather than leaving a stale assertion behind.
What this means for you
You can put a finding from this platform in front of a regulator, a court, a board or a partner agency and show where each element came from. That is the standard the tooling is built to — because in this work, being confidently wrong is more damaging than being usefully uncertain.
By the numbers
The taxonomy this entry belongs to is not a marketing list — it is the actual structure of the platform: 52 mission domains, 52 intelligence disciplines and 65 data points, each with a live dashboard behind it. Supporting that: 18 indicator types, 14 playbooks, 16 AI skills, 18 export formats and a 30-step automated pipeline.
This particular entry connects directly to 8 intelligence disciplines, 5 mission domains, 6 closely related entries — every one of them a tag you can follow, and a dashboard you can open.
Questions analysts actually ask
Can a v3 onion address be seized like a domain?
Not in the same way. A v3 address is derived directly from the service's ed25519 public key, so there is no registrar to compel, no WHOIS record and no DNS entry to redirect. Control of the address means possession of the private key. Law enforcement seizure of hidden services therefore happens by seizing the server and its key material, or by compromising the operator, after which the seized key can be used to serve a seizure notice at the same address. That is why seizure banners on onion sites are evidence of key compromise rather than of a registry action.
Is it legal to browse dark web marketplaces?
Passive viewing of publicly accessible pages is generally lawful in most jurisdictions, but the boundaries are narrow and jurisdiction-specific. Purchasing, downloading stolen data, engaging operators, accessing non-public areas or possessing certain categories of material can constitute serious criminal offences regardless of investigative intent, and in some jurisdictions merely viewing specific content is an offence. Operate under a written policy with named authorisation, define what is prohibited, and establish an escalation route in advance. Encountering child sexual abuse material requires immediate cessation and reporting through the defined channel, never collection.
How do I know I am on the real site and not a clone?
Verify against PGP-signed mirror lists published by the operator, check the signature rather than just the presence of a key block, and compare on-page artifacts such as cryptocurrency addresses and key fingerprints across mirrors. Cross-reference the address against forum announcements the operator controls and against multiple independent third-party trackers. Clones exist specifically to intercept payments and to catch researchers, and they mirror content faithfully. If verification fails, treat everything collected as unreliable, because a report based on a clone attributes another actor's content to the operation you were tracking.
How are hidden service operators actually identified?
Almost always through operator error rather than through any weakness in Tor. The recurring failures are exposing the origin server so it presents the same certificate, favicon or content on the clearnet, reusing handles, email addresses or PGP keys that exist elsewhere, cashing out cryptocurrency through regulated services subject to lawful process, leaking metadata in uploaded files, and operational security lapses in communications. Investigation therefore concentrates on collecting and correlating those artifacts, plus financial tracing toward off-ramps where identity data exists and can be obtained through proper legal process.
Should we monitor leak sites ourselves or use an aggregator?
Both, with different purposes. Aggregators give broad, low-cost coverage across many operations and are sufficient for detecting that your organisation or a supplier has been named. Direct monitoring gives you the primary artifact, the full posting content, the timing precision and the on-page identifiers that aggregators do not extract, plus independence from an aggregator's coverage gaps. For most organisations the correct answer is an aggregator for breadth plus direct monitoring of the small number of operations that specifically threaten your sector, conducted under proper policy and isolated infrastructure.
What do we do when our data appears on a leak site?
Treat the posting as the trigger for incident response and regulatory obligations, not as an intelligence curiosity. Preserve the posting with a hash and timestamp, confirm the claim's credibility, and start the notification clock, because many regimes measure from awareness. Do not download the published corpus without a lawful basis and a specific purpose, since that creates its own processing problem. Assess what data classes are involved and notify affected individuals where required. Route any question of payment to counsel and compliance immediately, because sanctions screening of the operator is a legal prerequisite.
Why do old onion addresses in reports no longer work?
Most commonly because they are version 2 addresses. The sixteen-character v2 scheme was deprecated and removed from the network, so every v2 address in older reporting is permanently dead and cannot be resolved by any client. Beyond that, hidden services move constantly: operators rotate addresses after law enforcement attention, rebrand, or simply lose their key material. Record the address with the date it was observed live, and when working from historical reporting, check the format first, because pursuing v2 addresses is a recurring and entirely avoidable waste of collection effort.
Standards, frameworks and further reading
Work that references a recognised framework is easier to defend, easier to hand over, and easier for a partner to consume:
- Tor rendezvous specification version 3 defines onion service addressing, key derivation, the distributed hash table and client authorisation.
- Tor Project research safety guidelines set the ethical baseline for measurement research that could affect user anonymity.
- Budapest Convention on Cybercrime provides the international framework for cross-border evidence gathering and cooperation in cybercrime investigations.
- ISO/IEC 27037 governs preservation of digital evidence, including the timestamped and hashed archiving of collected web content.
- Traffic Light Protocol governs onward sharing of collection findings within trust communities, particularly where collection methods are sensitive.
- GDPR and equivalent regimes apply in full to victim personal data published on leak sites, which remains personal data regardless of how it was disclosed.
- OFAC and equivalent sanctions regimes create strict liability exposure for payments or transactions involving designated criminal operations.
- National mandatory reporting frameworks for child sexual abuse material define the escalation route that overrides all collection activity.
References
Primary sources and authoritative references for this entry. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.
- Tor Project — The Tor Project. Specifications and documentation for onion services, addressing and operator security.
- Tor Metrics — The Tor Project. Authoritative network statistics including onion service counts and relay data.
- Ahmia — Ahmia. Open search engine indexing Tor hidden services with a published content policy.
- Ransomware.live — Ransomware.live. Aggregated ransomware leak site postings with victim listings and group attribution.
- Internet Organised Crime Threat Assessment — Europol. Law enforcement assessment of darkweb criminal ecosystems and disruption operations.
- Stop Ransomware — Cybersecurity and Infrastructure Security Agency. Government guidance and advisories covering ransomware operations and their leak infrastructure.
- Chainabuse — TRM Labs and partners. Community reports linking cryptocurrency addresses to criminal services and extortion.
- Convention on Cybercrime — Council of Europe. International treaty framework for cybercrime investigation and cross-border cooperation.
Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.
Put it into practice
The Quantus Intel threat intelligence platform operationalises this entry: monitored hidden service tracking with content archiving, extracted identifiers and automatic pivots to crypto and infrastructure. Explore the platform, or browse the rest of the library by following any tag above.