Security datasets that
never go stale.

Human-in-the-loop labeled data across six threat categories. Every sample reviewed by a certified security analyst — not an algorithm — then refreshed continuously as the threat landscape evolves.

Every label reviewed by a security analyst
Custom labeling — your data stays private
Refreshed daily, weekly, or on demand

Human analysts,
not algorithms.

Every sample in our datasets passes through a four-step human review process before publication. No ML model makes the final labeling call — a certified security analyst does. Every label ships with an analyst ID and full audit trail.

01

Automated Extraction

Honeypots, threat feeds, crawlers, and partner submissions identify candidate samples 24/7 and queue them for analyst review.

02

Analyst Primary Review

A certified SOC analyst triages each candidate: confirms the category, assigns an initial label, and adds a confidence score with reasoning.

03

Peer Verification

A second independent analyst reviews every label before publication. Disagreements escalate to a senior reviewer for adjudication.

04

Published to Dataset

The labeled sample ships with analyst ID, confidence score, reasoning tags, and timestamp — full audit trail on every label.

156 certified security analysts
24/7 global analyst coverage
SOC-2 audit trail on every label
Labels QA'd weekly for accuracy drift

6 threat categories. Analyst-verified. Continuously refreshed.

Most Popular

Phishing URLs & Domains

Analyst-verified

Human-in-the-loop labeled phishing and benign URLs with full feature annotations. Each sample reviewed by a security analyst. Updated daily with new campaigns detected by our honeypot network.

28.4M samples
Daily
CSV · Parquet · JSON Lines
Labels
phishingbenignsuspicious
Included features
URL structuredomain ageregistrarredirect chainspage content hash

Malware PE Binaries

Analyst-verified

Windows PE binary metadata labeled by malware analysts through a peer-review process. Includes file-level features and behavioral sandbox tags. Hashes only — no raw binaries.

2.4M samples
Weekly
CSV · Parquet
Labels
malwarebenignpuaransomwaretrojanworm
Included features
PE header fieldsimport tablesection entropybehavioral tagshash (SHA-256)
New

Malicious Browser Extensions

Analyst-verified

Chrome and Firefox extension metadata with analyst-assigned malice labels. Includes permission abuse patterns and code-level features from static analysis.

18.2K extensions
Weekly
CSV · JSON Lines
Labels
maliciousadwaretrackingbenignsuspicious
Included features
manifest permissionsJS code featuresCWS ratinginstall countupdate history

Malicious Network Traffic

Analyst-verified

Labeled network flow samples from honeypots and sandbox environments. Each flow reviewed by a network security analyst. Covers C2 communication, data exfiltration, lateral movement, and port scanning.

4.1M flow samples
Weekly
CSV · Parquet
Labels
c2exfiltrationlateral-movementbenignscanning
Included features
flow featurespacket timingport/protocolgeoASN

JavaScript Threat Snippets

Analyst-verified

Labeled JavaScript snippets extracted from web pages and browser extension content scripts. Analyst-verified. Includes obfuscated samples with deobfuscation tags.

890K samples
Weekly
JSON Lines · Parquet
Labels
maliciousobfuscatedcryptominerskimmerbenign
Included features
AST featuresentropyAPI call sequencesobfuscation score
Private

Custom Labeling Service

Analyst-verified

Bring your own unlabeled data. Our security analysts label it to your schema using the same human-in-the-loop review pipeline. Your data is never shared with third parties or used to train our models.

your data
On demand
Any format
Labels
your-labels
Included features
any feature set you need

What changed in 2026.

Attackers shipped four major technique changes in four months. A static dataset from December would miss all of them. Here's what our analysts caught — and how we responded:

Jan
Attackers adopt QR-code phishing to bypass URL filters
Analyst response: New 'qr-phishing' label added. 12K samples ingested and analyst-reviewed.
Feb
Domain generation algorithm (DGA) pattern shift detected
Analyst response: DGA label updated. 40K domain samples refreshed with analyst sign-off.
Mar
Zero-day exploited via malicious PDF attachments
Analyst response: PDF threat samples added. Existing benign PDFs re-evaluated by senior analysts.
Apr
AI-generated phishing lures reach 34% of campaigns
Analyst response: AI-generated label added. Detection guidance updated with analyst reasoning.

1,000 free analyst-verified labels per month.
No credit card.

Free tier includes access to the Phishing URLs dataset with daily analyst refresh. Upgrade to unlock all 6 datasets, custom labeling, and higher refresh cadences.