Social Profile: Data Point Intelligence Guide
A social profile is a persona’s self-published intelligence report. The metadata around it is usually more truthful than anything written in the bio.
A social profile is a persona's self-published intelligence report. The metadata around it is usually more truthful than anything written in the bio.
Understanding the Social Profile as an intelligence artifact
A social profile is a platform-hosted account representing a persona: a stable internal account identifier, a mutable handle, a display name, biography, avatar, follower and following graph, and a content timeline. Its structure is dictated by the platform, and each platform exposes a different subset publicly. A profile represents a curated self-presentation rather than an identity, so the analytic value lies less in the claims it makes than in the observable metadata surrounding those claims: timing, network, media and cross-references.
Analytically, separate the immutable numeric or opaque account identifier from the display handle, because the identifier survives renaming and lets you track a persona through rebranding. Distinguish personal accounts from business pages, bot-operated accounts and coordinated inauthentic clusters, which have different creation patterns, posting rhythms and network structures. Archived captures matter because deletion is the norm once an actor senses attention.
Why it matters
Profiles bridge personas and real-world context. Posting time distributions reveal an operator's working hours and therefore likely timezone; media reveals devices, locations and associates; the follower graph reveals affiliation and coordination. For threat intelligence, profiles are where recruitment, fraud lures, extremist mobilisation and stolen data advertising actually occur. For due diligence, they corroborate or contradict claimed employment and history. And because platforms retain far more than they display, the profile identifies where to serve legal process.
What analysts actually look for
These are the concrete, observable signals that carry weight in this area of work:
- Immutable account identifier and creation date, which persist through handle changes and expose accounts that predate or postdate a claimed history.
- Posting time-of-day distribution across a long window, which estimates the operator's working hours and probable timezone.
- Follower and following graph structure, revealing affiliation, coordination clusters and accounts created in the same batch.
- Reused avatar images whose hashes or reverse search results link the persona to other accounts and to stock or stolen photography.
- Cross-platform links published in the bio, which are self-asserted but frequently accurate and provide immediate pivots.
- Language, dialect, transliteration and idiom in posts, indicating native language and region independent of stated location.
- Engagement patterns inconsistent with follower counts, indicating purchased audience or automated amplification.
- Deleted or edited content recoverable from archives, which often contains the operational detail later judged too revealing.
Where the data comes from
Authoritative and openly available collection points. Always confirm licensing and terms before operational or commercial use:
- Platform-native public search and profile pages — Authoritative current profile data, follower counts, timeline content and account identifiers where exposed.
- Archive.today and Wayback Machine — Historic captures of profiles and posts, preserving content removed after the account drew attention.
- Reverse image search services such as Google Lens, Yandex and TinEye — Whether an avatar is stock, stolen or reused across other accounts and websites.
- Have I Been Pwned and breach index services — Whether an associated address appears in platform breaches, corroborating account ownership timelines.
- OSINT frameworks such as Maigret and WhatsMyName — Presence of the same handle across other platforms, giving further profiles to assess.
- Platform transparency and law enforcement request portals — Correct legal process route and data retention detail for the platform in question.
- Public academic and research datasets on coordinated inauthentic behaviour — Documented takedown datasets useful for recognising known influence operation patterns.
A working method
A repeatable sequence beats ad-hoc searching. This is a practical starting workflow:
- Define basis and scope — Record the lawful basis, the investigative question and the boundaries of collection before viewing or capturing any personal content.
- Capture immediately — Preserve the profile and relevant posts with timestamps, URLs and hashes at first sighting, because deletion frequently follows analyst attention.
- Anchor on the stable identifier — Record the internal account identifier and creation date so the persona remains trackable if the handle or display name changes.
- Analyse metadata over claims — Build posting time histograms, language indicators and network structure rather than relying on stated location, employer or age.
- Test the imagery — Reverse search avatars and posted media to detect stolen photographs, stock imagery and reuse across other personas.
- Map the network — Identify accounts with mutual, early or batch-created relationships, which distinguishes an organic account from a coordinated cluster.
- Escalate through proper channels — Where subscriber or private data is required, use platform legal process rather than attempting access, and document the request and response.
How this connects across the intelligence taxonomy
Intelligence work does not respect neat boundaries. The mission domain you are working, the disciplines you practise, and the data points you pivot on are one connected system. These are the direct relationships for this entry — every link is also a tag, so you can follow any thread across the whole library.
Collected by these disciplines
- Social Media Intelligence — Intelligence from Social Platforms and Networks
- Human Intelligence — Information from People, Ethically Obtained
- Criminal Intelligence — Intelligence Supporting Criminal Investigation
- Legal Intelligence — Law, Litigation, and Regulatory Intelligence
- Geospatial Intelligence — Intelligence Derived from Place
- Threat Actor Intelligence — Tracking Adversary Groups Over Time
- Disinformation Intelligence — Detecting and Analyzing Information Manipulation
- Identity Intelligence — Resolving and Verifying Who Someone Is
- News Intelligence — Media Reporting as an Intelligence Source
- Financial Intelligence — Following Value Through the Financial System
Investigated in these domains
- Human Trafficking
- Wildlife Trafficking
- Child Protection
- Gangs & Street Crime
- Counter-Terrorism
- Extremism & Radicalization
- Transnational Repression
- Election Security & PSYOP
- Disinformation / IO
Pivots to these data points
- Person / Name — A named individual — the subject of identity resolution and profiling.
- Email Address — Electronic mail address tied to an individual or organization.
- Username / Handle — Screen name or handle used across online platforms and services.
- Phone Number — Telephone number for voice, SMS, or messaging identification.
- Physical Address — A physical or mailing address tied to a person, company, or registered entity.
- Device / Advertising ID — A mobile advertising or device identifier used in adtech data to track and locate devices.
Inside the platform: where Social Profile lives
The Quantus platform is 204 pages behind a 147-item sidebar organised into six working groups: Command (24 items), Dashboards (15), Threat Theaters (14), Intelligence Domains (15), Investigate (34), and Administration (45). This entry is not a page in isolation — it is a thread running through several of them.
The modules that matter most here:
datapoint.php?dp=dp_profile— Data point hubhuman-trafficking.php— Human Trafficking dashboarddomain.php?d=wildlife— Wildlife Trafficking dashboarddomain.php?d=cp— Child Protection dashboarddomain.php?d=gangs— Gangs & Street Crime dashboardsearch.php— Advanced search, filter and pivotcorrelate.php— Correlation graphcases.php— Case management
Each dashboard is local-first: it renders from the platform’s own database rather than depending on a live third-party call, so it still works when an upstream API is unreachable or rate-limited. Heavy aggregates are cached with a hard query time cap and degrade to the last good value instead of hanging the page.
Automation, playbooks and AI skills
Analysis that only happens when someone remembers to run it is not a capability. The platform ships a 30-step automation pipeline (cron.php) that collects, ingests, resolves, enriches, correlates and scores on a schedule — 25 seeders, 11 resolvers and 7 enrichment runners, all idempotent and cursor-based so a run can be interrupted and resumed without duplicating or losing work.
AI skills that apply
The 16 one-click operations in ai-skills.php are deterministic jobs, not free-text generation. The ones that matter here:
- Enrichment Runner
- Enrichment → Local
- Correlate Infrastructure
- Summarise (Copilot)
- Generate Report
Alerting closes the loop: rules in alerts.php fire on new indicators matching a saved query, so a first sighting in this area raises a notification rather than waiting to be noticed at the next review.
Feeds, data sources and the API
The collection layer runs a feed registry of free, machine-readable sources — bulk blocklists and trackers (Maltrail, IPsum, FireHOL, the full abuse.ch corpora, phishing databases, Emerging Threats, Spamhaus, DigitalSide, ThreatView), authoritative government feeds (CISA KEV, OFAC, UN and EU sanctions lists), and reference datasets (RIR allocations, ip-to-ASN and geolocation tables, MITRE ATT&CK, EPSS). collect.php pulls them server-side on a schedule; feeds.php and source-catalog.php show what is registered, what it covers and when it last ran.
Anything the platform holds is reachable programmatically. The REST API in api.php exposes 11 endpoints — status, stats, search, lookup, recent, export, bulk_check, top_threats, by_category, categories, check — and export.php streams 18 formats in bounded chunks, so a million-row export neither exhausts memory nor times out:
STIX 2.1, MISP, OpenIOC 1.1, CEF (ArcSight), LEEF 2.0 (QRadar), Zeek/Bro intel, Snort/Suricata rules, Palo Alto EDL, BIND RPZ, hosts blackhole, iptables, CSV, JSON, NDJSON/JSONL, XML.
That covers the CTI standards (STIX 2.1, MISP, OpenIOC), SIEM ingestion (CEF, LEEF, Zeek), detection engines (Snort/Suricata), and direct enforcement (Palo Alto EDL, BIND RPZ, hosts, iptables) — so intelligence developed here can be actioned in the tools you already run, without a manual reformatting step. A TAXII 2.1 server and a MISP/RSS feed are also served for pull-based sharing.
Use cases
Three ways this entry earns its keep in day-to-day work:
- Triage under time pressure. An artifact or report lands and you need a defensible read in minutes, not days. Define basis and scope is the first move; the platform pre-computes the enrichment so the analyst spends the time on judgement rather than lookups.
- Building the picture. A single indicator is rarely the story. Anchor on the stable identifier turns one artifact into a network — shared infrastructure, repeated selectors, the same operator behind different names — via the correlation graph and the cross-entity link engine.
- Producing something actionable. Analysis that ends in a document nobody can use is wasted. Escalate through proper channels feeds the case file, the detection rule, the block list or the referral — with sourcing attached so the recipient can verify it.
Case management (cases.php), watchlists, saved searches and scheduled reports mean the work persists between sessions and survives an analyst leaving the team.
How each sector uses Social Profile
The same entry is worked very differently depending on who you are, what authority you hold, and what you are ultimately producing. A military analyst is supporting a commander’s decision; a journalist is meeting a publication standard; an NGO caseworker is protecting a person. The underlying artifacts are shared — the constraints, outputs and thresholds are not.
🎖 Military and defence
Profiles support force protection, information operations awareness and personnel security. Detecting recruitment approaches to service members, impersonation of unit or command accounts, and coordinated inauthentic amplification targeting a deployment are all legitimate defensive uses. Profile metadata, particularly posting time distributions and account creation clustering, characterises operator tempo and coordination for J2 reporting. Constraints are strict: collection on own nationals is restricted, covert engagement requires specific authorisation under national policy, and platform terms breaches can undermine both admissibility and access. Report observable metadata and network structure, not inferred personal attributes about identifiable individuals.
🕵 National intelligence
Profiles are where recruitment, procurement, radicalisation and illicit trading actually occur, which makes them collection rich and handling heavy. Requirements driven use covers persona tracking through rebranding, coordination analysis across account clusters, and identifying where legal process should be served because platforms retain far more than they display. Minimisation is central: profile content readily reveals political opinion, religion, health and sexuality, which attract stricter handling in most regimes and are frequently inferred inadvertently from group membership. Preserve captures with hashes and record the immutable account identifier, which survives renaming and is what a platform request must reference.
👮 Law enforcement
For law enforcement the profile is both evidence and a pointer to evidence. Public content should be captured with hashes, timestamps and the platform account identifier before any action that might prompt deletion. Non public content, message contents, login records, registration data and device information are held by the platform and require production orders, warrants or mutual assistance depending on category. Undercover engagement with an account is a regulated activity requiring specific authorisation in most jurisdictions. Chain of custody depends on documented capture method and tooling, since screenshots without provenance are readily challenged.
🔍 Private investigation and corporate security
Corporate security uses profiles for executive impersonation detection, insider risk indicators, brand abuse, pre litigation evidence and open source verification of claimed employment. The permitted ground is observation of publicly visible content for a defined purpose with a recorded balancing test. Not permitted: fake accounts to view restricted content, connection requests to gain access, engaging the subject, or sustained monitoring of a private individual amounting to harassment. Preserve with hashing since content is deleted, and disclose in the client report exactly what was visible publicly and what was not accessible.
📰 Journalism and OSINT media
Profiles are primary source material and among the most volatile. Standard practice is to archive before analysing, using both a third party archiving service and a hashed local capture, because subjects delete or lock accounts as soon as they sense attention. Verification requires metadata rather than claims: creation dates, posting patterns, network composition and media provenance. Consider whether the account is a real person, an impersonation or an inauthentic asset before attributing anything to a named individual. Seek comment before publication and weigh the safety consequences of naming, particularly for pseudonymous accounts in repressive contexts.
🌍 NGO, humanitarian and human rights
Civil society organisations document harassment campaigns, incitement and coordinated attacks on defenders, and support individuals whose profiles are targeted. Victim centred practice means the affected person controls what is documented and shared, consent is informed, and findings return to them first. Do no harm governs publication, since exposing a coordinated network can trigger escalation against the people it targets. Preservation follows Berkeley Protocol practice with hashes, capture times and tool versions so material supports later accountability. Duty of care includes rotating staff exposed to violent or abusive content and providing psychological support.
🎓 University and research
Profile research covers disinformation, coordination detection, platform governance and online harm, and it is among the most ethically scrutinised areas of internet research. Institutional review normally applies even to public content, and increasingly expects justification for identifiability, a data management plan and a deletion schedule. Method must state the collection interface, the sampling frame, the collection window and how deleted content was handled, since attrition biases results substantially. Platform terms constrain automated collection and researcher access programmes vary. Publish aggregate findings, code and where possible identifier lists rather than content, respecting platform redistribution rules.
Playbook: working Social Profile end to end
A repeatable sequence, from the moment the requirement lands to the moment a product is delivered and the case is closed out. Each phase states what you are trying to establish, not merely what to click — the point is a defensible chain of reasoning, not a checklist.
Phase 1 — Record purpose, basis and prohibitions
Document the investigative question, lawful basis and proportionality, and record explicitly what is prohibited: creating accounts to view restricted content, sending connection requests, engaging the subject and any covert persona use without authorisation. A good output is an authorisation naming the accounts and platforms in scope. Stop if the request amounts to open ended monitoring of a private individual.
Phase 2 — Preserve before you analyse
Capture the profile, timeline, media and any relevant threads immediately, using a capture tool that records hashes, retrieval times and page provenance, plus an independent third party archive. Content disappears within hours once a subject senses attention. A good output is a hashed evidence bundle with capture times and tooling recorded. Stop when everything you may need to cite is preserved, before any further action.
Phase 3 — Capture the immutable account identifier
Record the platform assigned numeric or opaque identifier beneath the visible handle, because it survives renaming and rebranding and is what a preservation or production request must reference. Handles change; identifiers do not. A good output is a handle to identifier mapping with the date. Stop when the identifier is captured for every account you intend to rely on.
Phase 4 — Establish the account timeline
Record creation date, earliest available content, handle changes where the platform exposes them, and any dormancy periods followed by reactivation. Creation clustering across accounts is one of the strongest coordination indicators available. A good output is a timeline per account with the evidence for each date. Stop when the timeline covers the period relevant to your question.
Phase 5 — Classify the account type
Distinguish a personal account, a business page, an automated or scheduled account, a commercially operated persona and a coordinated inauthentic asset, using creation pattern, posting rhythm, content originality, network composition and profile completeness. A good output is a classification with the specific indicators supporting it. Stop when the type is assessed, because attributing statements to a person requires knowing a person operates the account.
Phase 6 — Analyse posting metadata
Derive posting time distributions across a large sample to estimate operator working hours and likely timezone band, noting the timezone convention of the data source and controlling for scheduling tools. Combine with posting frequency and burst patterns. A good output is a distribution with sample size and caveats. Stop when the result is presented as a distribution rather than a location claim.
Phase 7 — Examine media provenance
Check images and video for reuse through reverse search, examine any retained metadata, and look for internal evidence of location, device and time. Platforms strip most embedded metadata, so treat its absence as normal rather than as concealment. A good output is a media provenance note per significant item. Stop when reuse is established or the item is assessed as original with the basis stated.
Phase 8 — Map the network
Record followers, following, interactions and mentions at a defined point in time, focusing on structural features such as reciprocal clusters, shared early followers and synchronised amplification rather than on individual connections. A good output is a dated network snapshot with the extraction method recorded. Stop when the structural question is answered; enumerating an individual's entire social graph is rarely proportionate.
Phase 9 — Test for impersonation and inauthenticity
Compare against the genuine account where one exists, checking creation date, follower composition, content originality and profile image reuse. Impersonation is common and folding an impersonator into a cluster attributes someone else's statements to your subject. A good output is an explicit authenticity assessment. Stop when you can state whether the account is genuine, impersonating or inauthentic with the evidence.
Phase 10 — Cross platform linkage with corroboration
Where the question requires linking profiles across services, rely on content correspondence, reused imagery hashes, self declared links and documentary identifiers rather than handle similarity alone. A good output is a linkage set with at least two independent corroborations per link. Stop when linkage is evidenced; a matching name is a lead, not a link.
Phase 11 — Minimise special category exposure
Review what the collection has revealed about political opinion, religion, health, sexuality or trade union membership, all of which are readily inferred from platform membership and content. Remove anything not necessary for the stated purpose and justify what remains. A good output is a minimised record with a special category note. Stop when unnecessary inferences are stripped before circulation.
Phase 12 — Report, restrict and review
Present findings separating observation from assessment, with capture references for every cited item, restrict access to the case team, apply retention and schedule deletion. Where platform enforcement is sought, package the evidence in the platform's reporting format. A good output is a graded product with a controlled evidence bundle and a deletion date. Stop when retention is scheduled and the evidence bundle is verifiable.
The platform ships this as a step-checked workflow in playbooks.php, so progress is recorded against a case rather than held in someone’s head.
Source register: what to collect from, and how
Sources are listed with their access model so you can plan around cost and licensing before you build a dependency on them. Open means no account required; registration means a free account or API key; licensed means paid or institutional access. Always confirm current terms — licensing changes, and a source that was free for research may not be free for commercial or evidential use.
| Source | Access | What it gives you | How it is used here |
|---|---|---|---|
| Platform transparency and researcher access programmes | Registration | Documented interfaces and datasets that platforms make available to researchers and to legal process, with terms and eligibility criteria. | The lawful route to structured profile and content data at scale, and the reference for what a platform will disclose. |
| Internet Archive Wayback Machine (archived) | Open | Historic captures of profile pages and posts including content later deleted, edited or made private by the account holder. | Recovers earlier profile states, biography text and linked sites which frequently carry the linkage or contradiction evidence. |
| Conifer by Rhizome | Open | On demand archiving service producing timestamped, replayable captures of dynamic pages including much social content. | Third party verifiable preservation at the moment of capture, before an account is locked, renamed or deleted. |
| Hunchly | Licensed | Investigation capture tool recording every page visited with hashes, timestamps and full page provenance in a case database. | Chain of custody grade capture of profile evidence with automatic hashing and retrieval time recording. |
| Reverse image search services | Open | Services matching an image against indexed web content to identify prior appearances, sources and reuse across sites. | Detects stolen or stock profile images, which is one of the fastest indicators of an inauthentic or impersonating account. |
| InVID and WeVerify verification toolkit | Open | Browser based toolkit for video and image verification including keyframe extraction, reverse search and metadata inspection. | Assesses provenance of media posted by a profile, including whether footage predates the claimed event. |
| Digital Forensic Research Lab publications | Open | Published case studies and methodology on detecting coordinated inauthentic behaviour and information operations across platforms. | Reference methodology for coordination indicators and for calibrating what evidence supports an inauthenticity claim. |
| Stanford Internet Observatory research outputs | Open | Peer reviewed and technical analyses of platform manipulation, network coordination and takedown datasets provided by platforms. | Benchmarks for network analysis method and access to documented takedown case material for comparison. |
| EU Digital Services Act transparency database | Open | Repository of statements of reasons for content moderation decisions submitted by very large online platforms operating in the EU. | Documentary evidence of platform enforcement actions and their stated grounds, usable for accountability reporting. |
| Bellingcat online investigation resources | Open | Published methodology and case studies covering profile verification, geolocation from imagery and cross platform linkage. | Practical technique reference with worked examples showing how verification claims are validated and where they fail. |
| Berkeley Protocol on Digital Open Source Investigations | Open | Methodological standard covering collection, verification, preservation and analysis of online material for accountability proceedings. | Governing framework for capture provenance, hashing and analyst documentation of profile evidence. |
| Have I Been Pwned | Open | Breach exposure service which can confirm that an address associated with a persona was registered with particular services and when. | Corroborates service registration timeline for a persona, supporting or contradicting a claimed account history. |
| CrowdTangle successor and content library interfaces | Registration | Platform provided research interfaces for querying public content, page and account level data within defined eligibility rules. | Structured access to public post and page data for coordination analysis without breaching platform terms. |
| Global Investigative Journalism Network resources | Open | Methodological guides and tool directories for social media research, verification and cross border investigation. | Practical guidance on verification workflow and on the ethics of identification in published reporting. |
| Association of Internet Researchers ethics guidelines | Open | The reference ethics framework for research involving online communities, public content and identifiable users. | Standard against which institutional review assesses profile based research design and identifiability decisions. |
Prefer sources that publish a methodology and a revision history. A dataset that changes silently is a liability in any product that has to survive challenge.
Tooling
Tools commonly used against Social Profile. None of these replace judgement, and each carries its own failure modes — know what a tool infers versus what it observes.
- Hunchly — Captures every page visited with hashes and timestamps automatically during an investigation. Limitation: fidelity degrades on infinite scroll and heavily scripted interfaces.
- Conifer and Wayback capture — Produces third party verifiable snapshots of profiles and posts. Limitation: several major platforms block archiving, leaving gaps exactly where evidence matters.
- Reverse image search services — Identify reuse of profile and post imagery across the web. Limitation: coverage differs sharply between services, so a single negative result proves nothing.
- InVID and WeVerify toolkit — Extracts keyframes and metadata for verifying video and image provenance. Limitation: platforms strip most embedded metadata on upload.
- Gephi and graph analysis platforms — Visualise and cluster follower and interaction networks to expose coordination. Limitation: layout algorithms create visually compelling structures that may not be statistically meaningful.
- Platform research APIs — Provide structured access to public content within terms and eligibility rules. Limitation: access programmes change frequently and coverage is partial and platform controlled.
- Stylometry and text similarity tooling — Compare writing across accounts as corroboration for common authorship. Limitation: unreliable on short posts and easily confounded by translation and templates.
- Timezone and activity analysis scripts — Derive posting hour distributions to estimate operator working patterns. Limitation: scheduling tools, automation and shared operation all distort the result.
- Case management with capture provenance — Stores profile evidence with hashes, capture times and access controls. Limitation: only effective where analysts capture at the moment of viewing rather than later.
AI skills and automation in detail
These are deterministic jobs with defined inputs and outputs, not open-ended prompting. Each is idempotent and cursor-based: interrupt one and it resumes where it stopped rather than duplicating work or losing progress.
- Enrichment Runner — Walks the indicator set through a chosen provider in time-boxed, cursor-based batches that resume rather than restart.
- Enrichment → Local — Materialises enrichment into the local store so dashboards render from your own database instead of a live third-party call.
- Correlate Infrastructure — Builds the cross-entity link graph: shared hosting, reused certificates, overlapping registrants, repeated selectors.
- Summarise (Copilot) — Produces a narrative summary beside the underlying records. It explains; it never creates indicators or assigns attribution.
- Generate Report — Assembles a sourced product from the current case or query, with provenance attached to each element.
A note on the boundary: the only skill that involves a language model is Summarise (Copilot), and it writes prose about records that already exist. Nothing else on this list involves generation of any kind. No indicator, relationship or attribution in the platform originates from a model. See the full skill list.
Tradecraft notes
The distinctions that separate a competent analyst from a fast one:
- Archive first, always. The single most common failure in profile work is analysing content that is deleted before it was preserved, and no amount of later effort recovers a locked or removed account.
- Record the immutable account identifier at first contact. Handles are renamed and reassigned, and a preservation request naming a display name that has since changed will return nothing.
- Creation date clustering is the strongest coordination indicator available from public data. Networks assembled in batches show tight creation windows that organic communities never produce.
- Absence of embedded media metadata is normal, not suspicious. Platforms strip it on upload, so treating its absence as evidence of concealment reveals unfamiliarity with how the platforms work.
- Follower counts are the least informative field on a profile and the most cited. Composition matters instead: who followed first, how reciprocal the graph is, and whether the followers themselves are recently created.
- Distinguish the account from the person throughout. Accounts are shared, sold, compromised, ghostwritten and operated by teams, so attributing a statement to a named individual requires evidence that the individual operated the account at that moment.
- Posting time analysis needs a large sample and an explicit timezone convention. Platforms display times in the viewer's timezone, so an analysis that does not state the source convention is uninterpretable and frequently wrong.
- Assume the subject will see your interest eventually. Views, follows, archive requests and even some scanning services leave traces, so plan the collection order so that the most volatile evidence is captured before anything that could tip them off.
Measuring whether it is working
Capability claims should be falsifiable. These are the measures that show whether work on Social Profile is producing anything, and they are worth baselining before you change process or tooling.
- Proportion of cited profile evidence preserved with hashes and capture timestamps before any subject facing action, sampled from closed cases.
- Share of investigations capturing the platform immutable account identifier, measuring readiness for preservation and production requests.
- Rate of authenticity assessments performed before attributing content to a named individual, tracked as a quality control.
- Evidence loss rate: proportion of cases where content was deleted before capture, tracked downward as preservation discipline improves.
- Number of cross platform linkages supported by two or more independent corroborations rather than handle similarity.
- Retention compliance for profile evidence bundles, measured as the proportion deleted or reviewed on schedule.
- Analyst exposure management: proportion of staff working with abusive or violent content who received scheduled rotation and support.
Beware of measuring volume alone. Indicator counts and report counts rise easily and say little; time-to-attribution, proportion of findings that survive review, and how often a product changed a decision say a great deal.
Common pitfalls
- Everything on a profile is self-asserted, and stated location, employer, age and gender are among the least reliable fields available.
- Avatars are routinely stolen from unrelated real people, so an image match can misidentify a victim as the operator.
- Impersonation and parody accounts closely mimic genuine profiles, and follower counts do not distinguish them.
- Automated collection often breaches platform terms of service and can invalidate evidence or trigger account and IP bans.
- Timezone inference from posting times fails for shift workers, schedulers and multi-operator accounts run from several countries.
- Viewing profiles from an attributable account can notify the subject and burn the investigation, particularly on professional networks.
Legal and ethical considerations
Profile content is personal data and can readily reveal special category information about political opinion, religion, health or sexuality, which attracts stricter conditions in most privacy regimes. Collect only what the investigative question requires, avoid covert engagement or fabricated accounts unless specifically authorised under a controlled legend policy, and respect platform terms since breaches can undermine evidential admissibility. Capture with hashes and timestamps for integrity, restrict access to the case team, and apply defined retention with periodic review.
Data integrity: no fabrication, no drift, no hallucination
Intelligence that cannot be traced back to a source is not intelligence, it is assertion. Everything in this entry — and everything in the platform behind it — is built on a small number of non-negotiable rules.
Provenance on every record
Every indicator carries the source that supplied it, a first-seen and last-seen timestamp, and a sighting count. Where several feeds report the same artifact, each contribution is recorded separately rather than collapsed, so you can see whether a finding rests on one source or twelve. Source attribution travels with the data into every export, so a recipient can audit a claim without asking you for the working.
Nothing is invented to fill a gap
If the platform has no data for Social Profile, it says so. Empty is displayed as empty — never padded with plausible-looking placeholder values, sample records or illustrative examples that a reader might mistake for observations. A dashboard with no rows is a true statement about collection coverage, and it is treated as a gap to close, not a blemish to hide.
Scoring is deterministic and reproducible
Threat scores, reputation grades and risk tiers are computed from stated inputs with fixed weights, not estimated. The same inputs always produce the same output, and the formula is visible rather than a black box. Aggregates are cached with an explicit time-to-live so a figure on screen is never silently stale — and when a heavy query exceeds its time budget the platform serves the last known-good value and labels it, rather than inventing a fresh number or hanging.
Where AI is used, and where it is not
Language models summarise and explain. They do not create indicators, assign attribution or manufacture relationships. No IP address, wallet, hash or identity in the platform originates from a model — every one is ingested from a named feed, resolved from a reference dataset, or entered by an analyst with a source recorded. Copilot output is presented as narrative alongside the underlying records, never in place of them, so a reader can always check the summary against the evidence.
Guarding against drift
Enrichment is additive and timestamped rather than overwriting. Reference data — sanctions lists, allocations, taxonomies — is re-synchronised from the authority on a schedule instead of being edited in place, so local copies cannot quietly diverge from the source of truth. Attribution is recorded with a confidence level and the reporting it rests on, and inferred relationships are labelled as inferred. When a source retracts or corrects, the correction propagates rather than leaving a stale assertion behind.
What this means for you
You can put a finding from this platform in front of a regulator, a court, a board or a partner agency and show where each element came from. That is the standard the tooling is built to — because in this work, being confidently wrong is more damaging than being usefully uncertain.
By the numbers
The taxonomy this entry belongs to is not a marketing list — it is the actual structure of the platform: 52 mission domains, 52 intelligence disciplines and 65 data points, each with a live dashboard behind it. Supporting that: 18 indicator types, 14 playbooks, 16 AI skills, 18 export formats and a 30-step automated pipeline.
This particular entry connects directly to 10 intelligence disciplines, 9 mission domains, 6 closely related entries — every one of them a tag you can follow, and a dashboard you can open.
Questions analysts actually ask
Can I use a fake account to see a private profile?
Not without specific authorisation, which most private, corporate, academic and journalistic actors do not have. Fictitious accounts breach platform terms, which can render evidence inadmissible and can cost your organisation access. Engaging a subject through such an account raises entrapment and harassment issues, and in several jurisdictions covert online personas are a regulated investigative technique requiring a formal authorisation and a legend policy. Where content is not publicly visible, the lawful route is a preservation and production request to the platform naming the account identifier.
How do I detect a coordinated inauthentic network?
Look at structure and timing rather than content. The most robust indicators are tight account creation clustering, synchronised posting within narrow time windows, near identical follower sets acquired in the same period, reused or stock profile imagery, low content originality with high amplification, and consistent activity gaps matching a single working day. No single indicator is sufficient. Build the case from several structural features, state the baseline you compared against, and be aware that ordinary communities and marketing operations can produce superficially similar patterns.
What does posting time analysis actually tell me?
With a large sample it indicates the hours during which an operator is active, which constrains the plausible timezone band. It does not give a location. It is confounded by scheduling tools, automation, travel, shift work and multi operator accounts, and it is frequently ruined by analysts failing to note that platforms render times in the viewer's timezone. Use hundreds of posts rather than dozens, state the timezone convention of your data source explicitly, present a distribution rather than a point conclusion, and corroborate with independent evidence.
How should I preserve profile evidence?
Capture at the moment of viewing using a tool that records the full page, its hash, the retrieval time and the URL, and simultaneously create an independent third party archive so the record does not depend solely on your own system. Record the tool and version. Screenshots alone are weak and readily challenged because they carry no provenance. For video and images, download the original file where the platform permits and hash it. Document any content that failed to capture, since gaps acknowledged are far better than gaps discovered later.
The account is deleted. Is the evidence gone?
Not necessarily. Check web archives, which often hold earlier profile states, and search for reposts, quote posts, screenshots circulated by others and cached search results. Platform held data survives deletion for a period and can be obtained under legal process, which is a strong argument for sending a preservation request early. Third party monitoring services and research datasets sometimes retain content. What you cannot do is reconstruct evidence you never captured, which is why preservation precedes analysis in every competent workflow.
How do I avoid collecting special category data?
Recognise that you will collect some inadvertently and manage it deliberately. Group membership, event attendance, follow relationships and profile imagery routinely reveal religion, political opinion, sexuality or health. Define the question narrowly so collection is bounded, review the record before circulation and remove anything not necessary for the purpose, and where retention is necessary record the specific lawful condition relied on. Avoid analytical steps that infer these attributes as a matter of course, such as enumerating every group an individual belongs to.
Is scraping public profiles lawful?
It is contested and jurisdiction dependent, and platform terms almost always prohibit it. Courts have reached differing conclusions on whether accessing public pages contrary to terms engages computer misuse law, but breach of terms carries separate consequences: account termination, loss of access, and challenges to the admissibility of resulting evidence. Data protection law applies independently, since scraped profile data is personal data requiring a lawful basis. Prefer official research interfaces and legal process, keep manual collection proportionate, and record the method used for every item.
Standards, frameworks and further reading
Work that references a recognised framework is easier to defend, easier to hand over, and easier for a partner to consume:
- GDPR Articles 5, 6, 9 and 14 with equivalent national law, governing lawful basis, minimisation, special category conditions and transparency for profile data.
- EU Digital Services Act transparency and researcher access provisions, defining platform obligations and vetted researcher data access routes.
- Berkeley Protocol on Digital Open Source Investigations, setting collection, verification, preservation and analyst documentation standards.
- ISO/IEC 27037 and ISO/IEC 27042, covering identification, preservation and analysis of digital evidence including web captures.
- National regimes governing covert online activity and undercover investigative techniques, which control use of fictitious personas.
- Association of Internet Researchers ethical guidelines, the reference framework for research involving identifiable online users.
- Editors code of practice and equivalent journalistic standards, governing privacy, accuracy and right of reply in identification.
- Platform terms of service and developer policies, which constrain automated collection and affect evidential admissibility.
References
Primary sources and authoritative references for this entry. Publishers revise and retire material, so treat the retrieval date as part of the citation and re-check before relying on any of it in a formal product.
- Berkeley Protocol on Digital Open Source Investigations — UN OHCHR. Standard for collection, verification and preservation of online material for accountability proceedings.
- Internet Archive Wayback Machine — Internet Archive. Historic capture service used to recover deleted or edited profile content. (archived copy — the publisher moved or withdrew the original)
- Meta transparency centre — Meta. Documentation of platform transparency reporting, researcher access and legal process routes.
- DSA Transparency Database — European Commission. Repository of platform content moderation statements of reasons under the Digital Services Act.
- Digital Forensic Research Lab — Atlantic Council. Published methodology and case studies on coordinated inauthentic behaviour detection.
- Stanford Internet Observatory — Stanford University. Research outputs on platform manipulation, network coordination and takedown datasets.
- InVID and WeVerify verification toolkit — InVID and WeVerify projects. Open toolkit for image and video provenance verification.
- Online investigation resources — Bellingcat. Published verification methodology and case studies for social media investigation.
- Internet research ethical guidelines — Association of Internet Researchers. Reference ethics framework for research involving identifiable online users.
- ISO/IEC 27037 digital evidence guidance — ISO. International standard for identification, collection and preservation of digital evidence.
Link integrity: every reference above was verified with a live request when this page was generated. Where a publisher had moved or withdrawn a document, the link was repointed at a preserved copy in the Internet Archive and marked as archived. Anything with no reachable copy anywhere had its link removed rather than left to rot — the source is still credited, it simply cannot be linked.
Put it into practice
The Quantus Intel threat intelligence platform operationalises this entry: captures and preserves profile evidence with metadata analysis, cross-platform linkage and integrity hashing under case controls. Explore the platform, or browse the rest of the library by following any tag above.