Security datasets that
never go stale.
Human-in-the-loop labeled data across six threat categories. Every sample reviewed by a certified security analyst — not an algorithm — then refreshed continuously as the threat landscape evolves.
Human analysts,
not algorithms.
Every sample in our datasets passes through a four-step human review process before publication. No ML model makes the final labeling call — a certified security analyst does. Every label ships with an analyst ID and full audit trail.
Automated Extraction
Honeypots, threat feeds, crawlers, and partner submissions identify candidate samples 24/7 and queue them for analyst review.
Analyst Primary Review
A certified SOC analyst triages each candidate: confirms the category, assigns an initial label, and adds a confidence score with reasoning.
Peer Verification
A second independent analyst reviews every label before publication. Disagreements escalate to a senior reviewer for adjudication.
Published to Dataset
The labeled sample ships with analyst ID, confidence score, reasoning tags, and timestamp — full audit trail on every label.
6 threat categories. Analyst-verified. Continuously refreshed.
Phishing URLs & Domains
Human-in-the-loop labeled phishing and benign URLs with full feature annotations. Each sample reviewed by a security analyst. Updated daily with new campaigns detected by our honeypot network.
Malware PE Binaries
Windows PE binary metadata labeled by malware analysts through a peer-review process. Includes file-level features and behavioral sandbox tags. Hashes only — no raw binaries.
Malicious Browser Extensions
Chrome and Firefox extension metadata with analyst-assigned malice labels. Includes permission abuse patterns and code-level features from static analysis.
Malicious Network Traffic
Labeled network flow samples from honeypots and sandbox environments. Each flow reviewed by a network security analyst. Covers C2 communication, data exfiltration, lateral movement, and port scanning.
JavaScript Threat Snippets
Labeled JavaScript snippets extracted from web pages and browser extension content scripts. Analyst-verified. Includes obfuscated samples with deobfuscation tags.
Custom Labeling Service
Bring your own unlabeled data. Our security analysts label it to your schema using the same human-in-the-loop review pipeline. Your data is never shared with third parties or used to train our models.
What changed in 2026.
Attackers shipped four major technique changes in four months. A static dataset from December would miss all of them. Here's what our analysts caught — and how we responded:
1,000 free analyst-verified labels per month.
No credit card.
Free tier includes access to the Phishing URLs dataset with daily analyst refresh. Upgrade to unlock all 6 datasets, custom labeling, and higher refresh cadences.