September 14, 2026

CryptoScamDB: Intelligence Source Guide

0

An open-source database of cryptocurrency scam domains and addresses, descended from EtherScamDB and built during the ICO and wallet-phishing era. Its value today is largely historical, and reading it as a current blocklist is a mistake worth understanding.

cryptoscamdb-intelligence-source-guide

An open-source database of cryptocurrency scam domains and addresses, descended from EtherScamDB and built during the ICO and wallet-phishing era. Its value today is largely historical, and reading it as a current blocklist is a mistake worth understanding.

At a glance

Source CryptoScamDB
Category Cryptocurrency & Blockchain › Crypto Scam & Phishing Databases
Homepage https://cryptoscamdb.org/
Machine interface https://api.cryptoscamdb.org/v1/addresses
Format JSON
Access Open — no account required
Disciplines Cryptocurrency Intelligence
Mission domains Fraud & Identity

Known scam/phishing crypto addresses & domains. — as catalogued in the platform’s own source registry.

CryptoScamDB is a public, open-source registry of scam entries, each associating a fraudulent website, project or campaign with the domains it used and the cryptocurrency addresses it collected funds at. It grew out of EtherScamDB, a project started to catalogue the wave of Ethereum wallet phishing and fraudulent token sales that accompanied the 2017 to 2018 boom, and was later broadened and rebranded. The dataset is structured rather than narrative: an entry carries a name, a category and subcategory such as phishing or fake token sale, one or more domains including subdomains and lookalikes, associated addresses with the chain they belong to, a status indicating whether the site was still live when last checked, and in some cases a description and the reporter. Alongside the scam list the project maintains a verified list of legitimate domains, which is the less celebrated and arguably more useful half, because a whitelist is what stops a lookalike-detection system from flagging the real site. The data lives in a public repository and is served by a JSON API with endpoints for addresses, domains and whole entries, plus check endpoints designed for a browser extension or wallet to call before a user interacts with a site. That design origin matters: this was built to protect wallet users in real time, not to support investigations, and the schema reflects it.

Two jobs, and they are not the ones the project was built for. The first is historical research. The corpus is one of the better surviving structured records of the phishing and fraudulent-offering ecosystem of the late 2010s, with domains and addresses linked into named campaigns, and that period is otherwise documented mainly in blog posts and dead links. If you are studying how crypto fraud infrastructure evolved, how lookalike domain strategies developed, or how the same operators recur across campaigns, this is a primary source with structure. The second is retrospective enrichment. When an old address or domain surfaces in a case – in a wallet's transaction history, in an archived email, in a seized device – a match here identifies the campaign it belonged to and dates it, which is context that current-generation blocklists do not carry because they discard what is no longer live. What it is not is a current threat feed. Freshness is the whole value of a phishing blocklist and it is precisely what this source cannot be assumed to have. Use it the way you would use an archived intelligence report: valuable about its period, silent about now, and dangerous only if you mistake the two.

Who publishes it, and why that matters

This is a community project with roots in the wallet-software world, built and maintained largely by a small number of security-minded volunteers associated with the Ethereum tooling ecosystem, and funded by nobody in particular. That is a familiar shape and it comes with a familiar trajectory: intense, high-quality curation while the founding maintainers were engaged, followed by a decline in submissions and review as their attention moved on and as commercial and better-resourced alternatives absorbed the function. Contributions were reviewed by hand, which produced good precision and does not scale, and there is no sustaining organisation behind it in the way TRM stands behind Chainabuse or a wallet vendor stands behind its own blocklist. The practical instruction is therefore blunt: before you use this source for anything operational, check when the data was last updated. That check takes a minute in the public repository and it determines whether you are looking at a live feed or an archive. Do not build a pipeline that assumes freshness, do not present a hit as current risk without dating it, and do not assume the API will be available indefinitely – unfunded community infrastructure in this niche has a poor survival record, and the correct hedge is to keep your own copy of anything you rely on.

Provenance is the first question to ask of any dataset and the one most often skipped. Who collects it, what their incentive is, whether they publish a methodology, and whether they correct the record when they get something wrong all bear directly on how much weight a finding drawn from it can carry.

What a record actually contains

The fields you will be working with, what each one means, and whether it is something you can pivot on. Read the meanings carefully — more analysis is wrecked by misreading a field than by failing to find one, and a field that looks like an observation is often an inference.

Field Type What it means Pivot value
name string The campaign or entity label the entry is filed under, usually the brand being impersonated or the fraudulent project's own name. It is a curatorial choice rather than a formal identifier, so the same operation can appear under variant names. Other entries impersonating the same brand, which is how you see a campaign rather than a site.
category / subcategory enum The classification, distinguishing phishing from fraudulent token sales, fake wallets, fake exchanges and similar. Assigned by a reviewing human at submission time, so it is more consistent than victim-chosen categories and reflects the taxonomy of its era rather than of today. The peer group of entries in the same class.
url / domain string The fraudulent site, frequently a lookalike of a well-known wallet or exchange using character substitution or an alternative top-level domain. Entries often list many domains for one campaign, which is the useful part. Registration and passive DNS records, hosting infrastructure, and the current occupant of the domain, which may be entirely unrelated.
addresses array Cryptocurrency addresses associated with the entry, mostly Ethereum-family. These are collection addresses observed at the time, which means they are the operation's funnel rather than its treasury, and they may have been in use for only days. The chain record for the address, which shows the actual take, the timing, and where funds moved next.
coin / chain enum Which network an address belongs to. Essential for hexadecimal addresses, which are syntactically valid on every Ethereum-compatible chain, and easy to lose when the field is sparsely populated in older entries. The relevant explorer for that chain.
status enum Whether the site was reachable when last checked. This is the field most likely to be stale, and a status of active on an entry from years ago means only that it was active at some point in the past. Archived captures of the site, which show what it actually looked like when it was live.
description string A short human note on what the entry is, present inconsistently. Where it exists it usually explains the impersonation target and the mechanism, which is the context that makes the domain list intelligible. none
reporter string Who submitted the entry, where recorded. A handful of prolific researchers account for a large share of the corpus, which is worth knowing because it means the dataset reflects their focus and their coverage rather than a systematic sweep. Other entries by the same reporter, which cluster by campaign type and period.
subdomain and path variants array Entries frequently enumerate multiple hostnames and paths for one operation. This is the closest the corpus comes to infrastructure analysis and it is genuinely useful, because domain families reveal registration patterns. Registrar and name server patterns across the family, which sometimes identify a bulk registration.
verified / whitelist entries array The separately maintained list of legitimate domains for wallets, exchanges and projects. Its purpose is to prevent lookalike detection from flagging the genuine article, and it is the part of the dataset most likely to be quietly wrong as legitimate services change domains. none
date added / last updated timestamp When the entry entered or was last touched. This is the single most important field in the dataset for a modern consumer, because it is the only thing that tells you whether you are reading intelligence or history. none

Coverage — and what is not in it

Coverage is heavily weighted to the Ethereum ecosystem and to the period roughly spanning 2017 to 2019, which is when the project's contributors were most active and when the threat it was built for was at its peak. Within that window it is a reasonably thorough record of wallet phishing, fake token sales, impersonation sites for major wallets and exchanges, and the domain families supporting them. Bitcoin and other chains appear but are secondary. Coverage of anything after that window should be assumed thin unless you have verified otherwise from the repository's own history, and coverage of the current threat picture – wallet drainers as a service, malicious signature requests, fraudulent decentralised applications, social engineering through compromised project accounts – is essentially absent, because that landscape emerged after the project's active period. Geographically the dataset has no meaningful dimension; phishing infrastructure is registered wherever is cheap and hosted wherever is tolerant, and the corpus reflects the researchers' visibility rather than any distribution of activity. Update rhythm was contribution-driven and human-reviewed, which produced good precision and slow throughput even at its most active. Treat the coverage statement here as a description of what the corpus contains, not as a claim about what is currently being added, and check the repository before assuming anything about the latter.

Known blind spots

Absence of evidence here is not evidence of absence. These are the conditions under which CryptoScamDB will not show you something that is nevertheless real:

  • It does not cover the current threat landscape. Drainer-as-a-service kits, malicious token approvals, signature-phishing and compromised-project social accounts postdate the project's active curation and are largely absent, so a modern phishing campaign will pass this check cleanly.
  • Maintenance has been intermittent and may have stopped. Verify the last update in the public repository before drawing any conclusion from a negative result, because an unmaintained blocklist returns clean answers indefinitely and gives no indication that it has stopped learning.
  • Domains outlive campaigns and change hands. An expired scam domain is routinely re-registered by an unrelated party, so an old entry can flag a site that is now entirely legitimate, and there is no automatic process reversing that.
  • The addresses are collection points, not infrastructure. Phishing operations generate throwaway receiving addresses, so a matched address usually identifies the funnel for one campaign rather than anything about the operator's holdings or ongoing activity.
  • Ethereum-centric. Bitcoin, Tron, Solana and the chains where a great deal of current fraud settles are thinly represented, and the account-model bias also means chain qualification on addresses is inconsistent.
  • Curation reflects a handful of contributors. The dataset covers what a small number of researchers were looking at, which was real and important and was not systematic, so gaps are not random and cannot be treated as such.
  • The whitelist ages badly in a specific and dangerous way. Legitimate projects change domains, get acquired and let old domains lapse, and a stale verified entry can vouch for a domain that has passed into someone else's hands.
  • No victim narratives, loss figures or reporting linkage. The corpus records infrastructure, not harm, so it cannot establish scale, intent or victim count the way a reporting platform can.
  • Availability is not assured. This is unfunded community infrastructure with no sustaining organisation, and the API being reachable today is not a basis for a dependency tomorrow.

Write the blind spot into the product. A statement that something “was not observed in CryptoScamDB” is defensible; a statement that it “did not happen” is not, and the difference is what survives cross-examination.

Access, licensing and what you may do with it

Access model: Open — no account required

The data is published openly in a public repository and served through a JSON API with endpoints for addresses, domains and full entries, plus check endpoints designed for a wallet or browser extension to call at the moment a user navigates or signs. In practice the repository is the more reliable access route and the one you should build on, for three reasons: it is the actual source of truth rather than a service in front of it, it carries revision history so you can see when entries were added and when curation slowed, and it does not depend on a hosted service that may or may not still be running. Clone it, keep a local copy, and refresh it on a schedule that reflects the actual change rate rather than a habitual hourly poll. If you use the API, handle failures as expected conditions rather than as exceptions, and never let a check-endpoint timeout resolve to a clean result in your logic, because that turns an outage into a silent approval. Before designing anything, look at the commit history and establish for yourself when the corpus was last meaningfully updated – that single check determines the entire correct use of this source.

Licence

The project is open source and the code and data are published publicly for reuse, which is the whole point of a community blocklist – it exists to be consumed by wallets, extensions and security tools. Confirm the specific licence text in the repository rather than assuming, because code and data in projects of this shape are sometimes covered by different terms and the data question is the one that matters to you. Attribution is both good practice and practically useful, since a downstream consumer of your product will want to know that a hit came from a community list of a particular vintage rather than from an authoritative determination. A separate consideration applies regardless of licence: the entries name domains and, by implication, the people who operated them, and republishing an assertion that a named project was a scam carries the ordinary risks of publishing an allegation. The licence permits redistribution; it does not resolve whether a particular entry is accurate, and an entry that was correct in 2018 about a domain that changed hands in 2021 is a defamation problem rather than a licensing one.

Rate limits and fair use

Nothing is published, and the correct assumption is that the hosted API is small, unfunded infrastructure that will not tolerate being treated as a commercial endpoint. If you use it, one connection, real delays between requests, an identifying user agent with a contact address, and backoff that actually backs off. But the better answer is to stop using it at volume entirely: because the data changes slowly and is published in full in the repository, there is no legitimate reason to make repeated network calls for lookups you can answer from a local copy in microseconds. Pull the repository on a schedule measured in days, index it locally, and query locally. This is faster, it removes the availability risk, it removes any question of whether you are abusing a volunteer's server, and it gives you a dated snapshot you can cite – which matters for a source whose freshness is the main thing a reviewer will question. The only case for calling the hosted API at all is a real-time wallet-protection use case, which is what the check endpoints were designed for and is probably not what you are building.

Licensing changes, and it changes without warning. A dataset that was free for research this year may not be free for commercial or evidential use next year. Confirm the current terms before you build a dependency on it, and record the terms you relied on alongside the data — the licence in force at the time of collection is part of the provenance.

Collecting it

How CryptoScamDB is actually pulled, in the order you would set it up. Prefer the bulk or export interface over per-item lookups wherever one exists: it is kinder to the publisher, faster for you, and gives a reproducible snapshot rather than a series of point-in-time answers you cannot reconstruct later.

Method Format Cadence Notes
Repository clone and periodic refresh bulk Weekly, or on notification of a change The primary route. Take the whole dataset from the public repository, keep it locally, and refresh on a slow schedule. This is the only collection method that gives you revision history, which is what tells you whether the corpus is alive.
Local address and domain index JSON Rebuilt on each refresh Build lookup indexes on addresses and on domains, including subdomain handling, and query them locally. Sub-millisecond checks, no network dependency, no rate limit, and no partial-outage failure mode.
Hosted API lookup JSON On demand only, if at all The address, domain and entry endpoints are convenient for occasional manual checks and for the real-time wallet use case they were designed for. Treat availability as best-effort and never let a failed call resolve to a negative result.
Revision-history mining bulk Once, for research The commit history dates every entry's addition and modification, which converts a flat list into a timeline of when campaigns were discovered. For historical research this is more valuable than the current state of the file.
Whitelist extraction for false-positive control JSON Weekly Take the verified-domain list separately and use it to suppress lookalike-detection false positives, while re-verifying independently that each whitelisted domain is still controlled by the party it names.

Ingesting it into the platform

Every step below is idempotent and cursor-based: interrupt one and it resumes from where it stopped rather than duplicating rows or losing progress. Collection is recorded per source, so a feed that quietly stops publishing shows up as a stale timestamp instead of silently thinning your coverage.

  1. Snapshot with a dated version stamp — collect.php records the repository revision and retrieval date with every import, so every downstream hit can state which vintage of the corpus produced it. For a source whose principal weakness is age, an undated import is unusable.
  2. Mark the whole feed as historical unless proven current — ingest.php applies a source-level qualifier derived from the newest entry in the snapshot. If the most recent addition is years old, every hit inherits that context automatically rather than depending on an analyst remembering.
  3. Split addresses and domains into separate indicator types — resolve-everything.php registers addresses as crypto address entities and domains as domain entities, linked through the campaign entry rather than merged. The two age at completely different rates and must not share a confidence value.
  4. Re-resolve domains before presenting them — resolve-domains.php and passive-dns.php check current registration and hosting for every matched domain, because the most common failure here is an expired scam domain now operated by an unrelated party. A hit on a re-registered domain is presented as a historical association, never as a current one.
  5. Verify addresses against the chain rather than trusting the entry — enrich.php pulls the actual transaction history for matched addresses, which establishes when the address was active, how much it collected and whether it has moved since. That converts a list membership into a dated fact.
  6. Cross-reference against actively maintained lists — correlate.php compares hits against current phishing and drainer lists. Presence in both is corroboration of an ongoing operation; presence only here is a strong signal that you are looking at history.
  7. Route to phishing and fraud views with the vintage attached — Matches surface in phishing-center.php and blockchain.php carrying their snapshot date, so an analyst sees the age at the point of use rather than in a footnote. Routing is on the entry's own category, not on any inferred classification.
  8. Export with source characterisation intact — export.php labels these indicators as community-curated and dated in STIX 2.1 or MISP output, so a recipient consuming the feed knows they have received an archive record and not a live detection.

Registered sources and their last-collected state are listed in sources.php, and the scheduled chain that keeps them current is in automation.php.

How it is wrong, and how to tell

Every dataset is wrong in characteristic ways. Knowing which ways is the difference between using a source and being used by one, and it is the part of source evaluation most often skipped because it is the part that takes work.

Precision was good and recall was always narrow, and the gap between those two facts is the whole story of this source. Entries were reviewed by hand before publication, so the false-positive rate at the time of curation was low by the standards of community blocklists – when this dataset said a domain was a phishing site, it usually was. What it never had was breadth: a small group of contributors covering a fast-moving landscape catches what they see, and the corpus is a record of their visibility. That was an acceptable trade when the data was fresh, because a precise partial blocklist is useful. It is a poor trade now, because the precision applies to a snapshot of the past while the recall gap applies to the present, and the combination produces a source that is confidently right about things that no longer matter and silent about things that do. Judge it accordingly: high confidence in any individual historical entry, low confidence that a clean result means anything at all, and no confidence whatsoever that a domain flagged years ago is currently hostile. The verified whitelist deserves separate and harsher treatment, because a stale entry there fails in the direction that causes harm.

Characteristic false positives

  • Re-registered domains. An expired phishing domain picked up by an unrelated buyer keeps its entry forever, so the list flags a site that may now be a small business, a parked page or a legitimate project.
  • Clean results from a stopped feed. An unmaintained blocklist answers negative to everything with complete confidence, and there is nothing in a negative response to indicate that the corpus has not learned anything new in years.
  • Chain-ambiguous addresses. Hexadecimal addresses are valid across every Ethereum-compatible network and older entries do not always qualify which chain they mean, so a match can attach a campaign to activity by a different party entirely.
  • Stale whitelist entries vouching for domains that changed hands. This is the most damaging failure mode in the dataset, because it fails safe in the wrong direction and actively suppresses a warning.
  • Subdomain and path matching errors. Entries sometimes list a specific path or subdomain on a shared hosting or platform domain, and a matcher that generalises to the parent domain will flag an entire hosting provider.
  • Campaign name collisions. Entries are labelled by the brand being impersonated, so searching by name returns both the impersonators and, if your matching is loose, the legitimate brand itself.
  • Addresses read as attribution. A collection address for a phishing campaign was controlled by the campaign for a matter of days; treating it as an enduring identifier for an actor produces links across years that have no basis.

None of these make the source unusable. They make it a source that requires corroboration before an assertion built on it goes into a product, which is true of every source and admitted by few.

Ageing

This source is a case study in how blocklists die, and the failure is instructive because it is silent. Individual entries do not become false: a domain that ran a phishing site in 2018 did run one. What decays is every inference you might draw. The domain has probably lapsed and may belong to someone else. The addresses were emptied years ago and identify nothing ongoing. The status field, if it says active, is reporting an observation nobody has repeated. And the absence of an entry, which is the answer you get most of the time, has degraded from a weak signal to no signal at all. A stale record here looks completely healthy – well-formed, categorised, plausible – which is precisely why the mitigation has to be structural rather than a matter of analyst care. Attach the snapshot date to every hit, derive a feed-level currency qualifier from the newest entry in the corpus, and re-resolve any domain before it is shown to a human. Do those three things and the source becomes a useful historical index. Skip them and it becomes a machine for producing confident, dated-looking, wrong answers.

What this source feeds

A source is only worth what it lets you conclude. These are the disciplines that collect through it, the mission domains it serves and the data points it yields — every one is a tag, so you can follow any thread from here into the rest of the library.

Collected by these intelligence disciplines

Serves these mission domains

Yields these data points

How each sector uses CryptoScamDB

The same dataset is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The records are shared — the constraints, thresholds and outputs are not.

🎖 Military and defence

Limited direct relevance, with one genuine use: retrospective analysis of compromise. When an incident review turns up an old phishing domain or a crypto address in an archived mailbox, on a recovered device or in historical proxy logs, a match here dates the campaign and names what was being impersonated, which is context that current-generation feeds have discarded. That helps establish when an exposure actually began, which is frequently the hardest question in a retrospective. Do not use it for current detection – the coverage gap against modern techniques is total – and do not treat a clean result as any kind of assurance. Where crypto fraud touches personnel welfare and financial vulnerability, current reporting platforms are the appropriate source.

🕵 National intelligence

For CRYPTINT and OSINT this is an archive with structure, and archives with structure are rarer than they should be. Its use is historical reconstruction: identifying the campaign an old address or domain belonged to, dating an operator's earliest observed activity, and tracing lineage where the same techniques or naming conventions recur years later. The revision history is the underused asset, because it dates discovery rather than merely listing entries, and discovery dates support timeline work that the current file cannot. Characterise it honestly in any product as a community-curated dataset with a defined active period, and never cite it in support of a present-tense claim about infrastructure without independent current verification.

👮 Law enforcement

Useful in cold and historical matters, of little use in live ones. In a case involving conduct from the relevant period, a match ties an address or domain to a named campaign and gives you a starting frame for what the victim experienced, which can be corroborated against archived captures of the site and against the chain record. As with any community dataset it is lead information and not evidence: the entry is an unverified curatorial judgement by a volunteer, and the evidential version is the archived site, the chain data and the victim statement. For current investigations, use actively maintained sources and victim reporting platforms instead, because a negative result here carries no exculpatory weight whatsoever.

🔍 Private investigation and corporate security

In due diligence work this has a narrow but real use: checking whether a counterparty, project or domain has a history in the fraudulent-offering era. Many people currently promoting crypto ventures were promoting different ones in 2018, and a documented association with a catalogued scam from that period is a material finding for a background report. The technique is to search names, domains and known addresses across the corpus and then verify each hit independently through archived captures and chain data before it goes anywhere near a client deliverable. What you must not do is present a hit as current risk or a clean result as a clearance, both of which misrepresent a dormant dataset as an active one.

📰 Journalism and OSINT media

The corpus is a usable primary source for reporting on the history of crypto fraud, and for the specific and recurring story of people who ran or promoted fraudulent projects in one cycle and reappeared in the next. Entries give you names, domains and addresses that can be corroborated against archived versions of the sites and against the chain, which is the standard of verification such a story requires. Two cautions. Entries are volunteer judgements, not adjudications, so a name in this database is a lead to investigate rather than a fact to publish. And domains change hands, so the current occupant of a domain listed here may have nothing to do with what it once hosted, which is an easy and defamatory mistake to make.

🌍 NGO, humanitarian and human rights

For consumer protection and digital safety work the material value is educational and archival rather than operational. The corpus documents what a generation of crypto phishing looked like, which supports awareness material grounded in real examples rather than in invented ones, and it supports advocacy arguments about how long fraudulent infrastructure persisted and how little was done about it. For protecting people today, use current sources: an unmaintained blocklist offers no protection and offering it as protection creates false confidence, which in consumer safety work is worse than offering nothing. Victim referral remains the national fraud reporting mechanism first.

🎓 University and research

This is a good research dataset with a clearly bounded period, structured campaign-to-infrastructure linkage and a public revision history, which together support longitudinal work on the evolution of crypto phishing infrastructure, domain strategies, campaign lifetimes and the economics of collection addresses. It is also a good case study in the sustainability of volunteer security infrastructure, which is a research question in its own right. The methodological requirements are explicit period bounding, acknowledgement that the sample reflects a small set of contributors rather than systematic collection, and independent verification through web archives and chain data rather than treating entries as ground truth. Any study that uses this corpus to characterise the current threat landscape is measuring the past and calling it the present.

Playbook: working CryptoScamDB end to end

A repeatable sequence from first pull to finished product. Each phase states what you are trying to establish, not merely what to click — the objective is a defensible chain of reasoning, not a completed checklist.

Phase 1 — Establish the vintage before anything else

Open the public repository and find the date of the most recent meaningful update. This single fact determines whether the source is a feed or an archive, and it changes every subsequent step. An analyst who skips it will make confident statements about current risk on the basis of a dataset that stopped learning years ago.

Phase 2 — Decide explicitly that you are doing historical work

Frame the question as what was this domain or address associated with, and when. If your question is whether something is dangerous today, this is the wrong source and you should go to actively maintained phishing and drainer lists instead. Being clear about this at the outset prevents a whole class of error.

Phase 3 — Take a local copy and query it locally

Clone the repository, index addresses and domains, and never make a network call for a lookup again. Local querying is faster, removes availability risk, and gives you a citable dated snapshot, which is exactly what a source of this kind needs to be defensible.

Phase 4 — Search names, domains and addresses in parallel

Campaigns appear under a brand name, a set of domains and a set of addresses, and a search on only one dimension misses entries where the linkage was recorded differently. Run all three and merge on the entry identifier rather than assuming any single field is complete.

Phase 5 — Date the discovery, not just the entry

Use the repository revision history to establish when an entry was added and when it was last changed. That gives you a discovery date for a campaign, which is far more analytically useful than the flat current state of the file and is what supports timeline work.

Phase 6 — Re-resolve every domain before you rely on it

Check current registration, name servers and hosting, and check whether the domain has changed hands since the entry. A domain that lapsed and was re-registered is now somebody else's property, and presenting it as a scam site is both wrong and actionable against you.

Phase 7 — Pull archived captures of the site

Web archives frequently hold captures of a listed phishing page, which is the actual evidence of what the site did. That capture is worth far more than a database entry asserting it, and in any context where the finding will be tested it is what you should be citing.

Phase 8 — Verify addresses against the chain

Look up each matched address on the relevant explorer and establish its active window, its total receipts and where funds moved. This tells you whether the campaign was significant or trivial and whether the address has any current relevance, neither of which the entry can tell you.

Phase 9 — Cross-check against a currently maintained source

Compare hits against present-day phishing domain lists and drainer address lists. Overlap indicates something that has persisted or recurred, which is genuinely interesting. Absence from current lists confirms that you are working with history and should say so.

Phase 10 — Use the whitelist, but re-verify it

The verified-domain list is useful for suppressing lookalike false positives, and it fails dangerously when a legitimate project has moved on and a domain has lapsed. Check independently that each whitelisted domain is still controlled by the named party before letting it suppress a warning.

Phase 11 — Record what a negative result actually means

Write down, in the case file, that no entry was found in a dataset with a stated last-update date. That sentence prevents the finding from being read later as a clearance, which is how a dormant blocklist causes real harm in due diligence work.

Phase 12 — Preserve your snapshot with the finding

Keep the exact copy of the dataset that produced a hit, with its revision identifier, alongside the case record. Community projects disappear, and a finding you cannot reproduce because the source went offline is a finding you will eventually have to withdraw.

The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.

What to pair it with

No single source carries a finding. These are the datasets that corroborate, extend or contradict this one — and a source that contradicts is worth more than one that agrees, because it is the only thing that will tell you when you are wrong.

Source Relationship What it adds
MetaMask eth-phishing-detect supersedes An actively maintained crypto-phishing domain blocklist with far broader current coverage. Where this project stopped, that one continued, and for present-day domain checking it is the right source.
ScamSniffer scam database supersedes Current-generation address and domain lists focused on wallet drainers, covering the threat class that emerged after this project's active period.
Chainabuse extends Victim-reported abuse across many chains, supplying the harm and typology layer that an infrastructure list does not contain, and covering the present rather than the past.
Internet Archive Wayback Machine (archived) corroborates Archived captures of listed sites, which is the actual evidence of what a phishing page did and the only way to verify a historical entry properly.
Etherscan prerequisite Chain records for Ethereum-family addresses, establishing when a collection address was active, what it received and where funds went next.
urlscan.io extends Historical and current scans of URLs with screenshots, DNS and request chains, useful both for verifying old entries and for checking whether a listed domain has changed character.
ICANN extends Domain registration lookup infrastructure, which is how you establish whether a listed domain lapsed and was re-registered by an unrelated party.
Office of Foreign Assets Control corroborates Formal designations of addresses, which occasionally intersect with historical scam infrastructure and carry legal force where community lists do not.

Legal, ethical and operational constraints

The dataset is published for reuse and the ordinary open-source conditions apply, so consumption and redistribution are not the problem. The problem is assertion. Every entry is an accusation that a named project, site or operator committed fraud, made by a volunteer, reviewed informally and now years old. Repeating that accusation in a published report, a due diligence product or an article is publishing an allegation on somebody else's authority, and if the entry was wrong, or was right about a domain that has since changed hands, the exposure is yours and not the project's. In most jurisdictions the defence for such a publication rests on your own verification, which means archived captures, chain records and current registration checks rather than a database hit. Data protection obligations are lighter here than for victim-report sources because the corpus contains infrastructure rather than personal narratives, but where an entry names an individual you are processing personal data and accuracy obligations attach – and an inaccurate accusation retained indefinitely is the specific harm those obligations address. Finally, if you redistribute this data in a product, characterise it as a dated community list, because a customer who acts on it as current risk will have been misled by your framing rather than by the source.

Operational security

Querying the hosted API discloses to whoever operates it which domains and addresses you are checking, which is a modest exposure for a small volunteer project and not a modest one if the infrastructure has changed hands or been compromised – an unmaintained service is a plausible target precisely because nobody is watching it. The clean answer is the one that is also better engineering: clone the repository, query locally, and make no network calls at all. That eliminates disclosure entirely and it is not a compromise, because local lookups are faster and more reliable than remote ones. If you do use the hosted endpoints, do not do so from infrastructure that identifies your organisation, and be aware that a check-endpoint call pattern reveals your target set in real time. There is a further consideration for anyone using the check endpoints in a user-facing tool: doing so sends your users' browsing or transaction targets to a third party, which is a privacy disclosure you are making on their behalf and should not make silently. Bundling the list locally is the correct design for that case too.

Two rules that hold regardless of jurisdiction. Collection that is lawful is not automatically proportionate, and a dataset assembled for one purpose does not carry consent for another. Where the records concern identifiable people, the question is not only whether you may hold the data but whether holding it serves the purpose you are accountable for.

Is it earning its place?

Sources accumulate. Feeds get added during an incident and are never reviewed again, and a decade later the pipeline is carrying dead weight that nobody dares remove. These are the measures that show whether CryptoScamDB is contributing anything, and they are worth baselining now so the answer is available later.

  • Age of the newest entry in your current snapshot, monitored continuously, because it is the single number that determines whether the source should still be consulted at all.
  • Proportion of hits that survive independent verification through archived captures and chain records, which is the real precision figure rather than the historical one.
  • Proportion of matched domains that have changed registration since the entry was created, tracked as the leading indicator of the source's dominant false-positive mode.
  • Overlap rate between hits here and hits in actively maintained lists, which distinguishes persistent operations from purely historical ones.
  • Number of case findings where this source supplied campaign context unavailable from any current feed, which is the honest measure of its remaining value.
  • Count of whitelist entries that fail independent control verification, audited on each refresh, since these fail in the direction that suppresses warnings.
  • Frequency with which a negative result from this source appears in a product without an accompanying statement of the corpus date, tracked as a discipline failure.

Beware of volume. Indicator counts rise easily and say almost nothing. Unique contribution — findings this source produced that no other source in your stack would have — is the measure that matters, and it is usually far lower than anyone expects.

Tradecraft notes

The distinctions that separate a competent analyst from a fast one:

  • Check the last update date before you check anything else. For a blocklist, currency is not a quality attribute, it is the entire product, and a dormant list is a different kind of object from a live one.
  • A clean result from an unmaintained feed is not information. Write that sentence into the case file explicitly, because six months later somebody will read the absence of a hit as a clearance.
  • Domains change hands and entries do not. Always re-resolve before presenting a domain, and treat any match on a lapsed and re-registered domain as historical association only.
  • The revision history is the better dataset. Discovery dates support timelines; a flat current file supports only lookups, and the flat file is what most consumers use.
  • Collection addresses have short lives. Matching one identifies a campaign's funnel during a specific window and says nothing about the operator's holdings, ongoing activity or identity.
  • The whitelist is the dangerous half. A stale verified entry suppresses a warning that should have fired, which is a failure that produces harm rather than noise.
  • Precision from the past does not transfer to the present. This dataset was carefully curated, which makes individual old entries trustworthy and makes the corpus as a whole a poor guide to current risk.
  • Keep your own snapshot with every finding. Unfunded community projects go offline, and a citation to a dataset that no longer exists is a finding you cannot defend.
  • Use it to explain, not to detect. The right question is what was this, and when – a question archives answer well and live feeds answer badly.

Questions analysts actually ask

Is this data still being updated?

Check the public repository yourself before relying on it – the commit history answers the question definitively and takes a minute. Community maintenance has been intermittent and the project's most active curation belongs to an earlier period. Design your use around what you find rather than around any assumption, including the ones in this guide.

Can I use it as a phishing blocklist for users?

Not on its own, and not as a primary control. Its coverage of current techniques – drainer kits, malicious approvals, signature phishing – is essentially absent, so it will pass modern attacks cleanly while creating the impression of protection. Use an actively maintained list for that job and keep this one for retrospective questions.

How does it differ from eth-phishing-detect?

Both grew out of the same ecosystem and both list crypto phishing domains, but eth-phishing-detect is maintained continuously by a wallet vendor with a direct operational stake, includes lookalike-detection machinery and a whitelist, and is consumed in production by a large user base. CryptoScamDB adds campaign-level structure and associated addresses, which eth-phishing-detect does not, and is the better of the two for historical work.

The addresses look empty. Did the entry get it wrong?

Usually not. Phishing operations sweep collection addresses quickly, so an address with a small historical inflow and a zero balance is exactly what a working campaign looks like years later. Check the address history rather than the balance, and note the active window – that window is often the most useful thing the entry gives you.

A domain in the list is now a legitimate business. What happened?

The scam domain lapsed and somebody else registered it, which is routine. The entry is a historical record and nothing in the dataset reverses it automatically. Never present such a domain as currently malicious, and build the re-resolution check into your pipeline rather than relying on analysts to notice.

Should I query the API or clone the repository?

Clone the repository. It is the source of truth, it carries revision history, it removes any availability dependency, and local lookups are orders of magnitude faster. The hosted API exists for real-time wallet protection, which is a use case with different requirements from yours.

Can I cite this in a due diligence report?

Yes, if you cite it accurately: as a community-curated dataset, with the date of the snapshot you used, and with each relevant hit independently verified through archived captures and chain data. Citing it as a current risk determination, or citing an absence of hits as a clearance, misrepresents the source and puts the exposure on you.

Does it cover Bitcoin?

Partly and thinly. The project grew out of Ethereum wallet phishing and its centre of gravity stayed there, so Bitcoin entries exist but should not be treated as any kind of systematic coverage. For Bitcoin-centric questions, victim reporting platforms and chain analysis are more productive routes.

What is the single most useful thing in this dataset today?

The linkage between a campaign, its family of domains and its collection addresses, dated by the revision history. Almost nothing else preserves that structure for the late-2010s phishing era, and it is what lets you recognise that a name resurfacing today has a documented past.

Standards, formats and interoperability

What this source speaks natively, and what it has to be translated into before a partner can consume it. Work that arrives in a recognised format is easier to defend, easier to hand over and easier to automate against:

  • Domain name syntax and registration records, including registrar and name server data, which is what you must re-check before treating any listed domain as currently relevant.
  • Ethereum account address format, a hexadecimal identifier valid across every Ethereum-compatible network, which is why chain qualification matters and why its absence in older entries is a real problem.
  • Bitcoin address encodings, present in the corpus but secondary to its Ethereum focus.
  • Common phishing taxonomies distinguishing impersonation, credential harvesting and fraudulent offerings, which the entry categories approximate at the granularity of their era.
  • STIX 2.1 and MISP for exporting domains and addresses as indicators with explicit dating and source characterisation, which this source requires more than most.
  • Web archive capture formats, the practical evidence layer that turns a database entry into something verifiable.
  • Open-source repository revision history as a dating mechanism, which for this dataset is a first-class analytical feature rather than a development artefact.

References

Primary documentation and authoritative references for this source. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.

  1. CryptoScamDB — CryptoScamDB. The project site, with search over the corpus and links to the underlying data. The quickest way to see the shape of an entry before deciding how to consume the dataset.
  2. CryptoScamDB repositories — CryptoScamDB. The public repositories holding the dataset, the API implementation and the revision history. This is the authoritative source and the place to establish when curation actually stopped.
  3. MetaMask eth-phishing-detect — MetaMask. The actively maintained crypto phishing domain list that took over the live-protection role. The correct source for present-day domain checking.
  4. ScamSniffer scam database — ScamSniffer. Current-generation drainer and phishing address lists, covering the threat class that postdates this project's active period.
  5. Chainabuse — TRM Labs. Victim-reported abuse across many chains, which supplies the harm dimension an infrastructure list lacks and is current where this source is not.
  6. Internet Archive Wayback Machine — Internet Archive. Archived captures of listed sites. For historical verification this is the evidence, and a database entry is merely the pointer to it. (archived copy — the publisher moved or withdrew the original)
  7. Etherscan — Etherscan. Ethereum chain explorer for verifying the activity, timing and onward movement of listed collection addresses.
  8. urlscan.io — urlscan.io. URL scanning with historical results, screenshots and request chains, useful for both verifying old entries and detecting that a listed domain has changed hands.
  9. ICANN — Internet Corporation for Assigned Names and Numbers. Registration data lookup infrastructure and policy, the route to establishing whether a listed domain lapsed and who holds it now.
  10. Ethereum — Ethereum Foundation. Reference material on accounts, token approvals and transaction signing, which is the background needed to understand what the phishing techniques in this corpus actually did.
  11. Office of Foreign Assets Control — US Department of the Treasury. Designated addresses, the one attribution layer with legal force, occasionally intersecting with historical fraud infrastructure.
  12. Internet Crime Complaint Center — US Federal Bureau of Investigation. The victim reporting pathway. Historical infrastructure research is not victim assistance, and anyone with a current loss should be routed here first.

Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.

Put it into practice

The Quantus Intel threat intelligence platform operationalises this source: it imports the corpus as a dated snapshot with a feed-level currency qualifier derived from the newest entry, re-resolves every matched domain against current registration before it reaches an analyst, verifies listed addresses against the chain, and exports hits labelled as dated community records rather than live detections.. Browse the full source catalogue, or follow any tag above into the rest of the library.

Leave a Reply