PhiShark2026: A Multi-Layer Active-Web Raw-Evidence Dataset for Phishing Website Research

2026-08-24Cryptography and Security

Cryptography and Security
AI summary

The authors created a large dataset of phishing and safe websites that keeps all the original information, like webpage content, URLs, and internet details, instead of just simplified features. This helps researchers study phishing in more detail and adjust to changing tricks by phishers. They also used a method to avoid confusing website hosting details when sites share servers. Their analysis found clear differences between phishing and legitimate sites in things like domain age and security settings. The dataset is designed to be easy to check and reuse for future phishing research.

phishing websitesdatasetURLHTML contentTLS certificatesDNSsecurity headershosting infrastructuredomain maturityweb measurement
Authors
Furkan Çolhak, Ferhat Demirkıran, Hasan Dağ, Alexander Iliev
Abstract
Phishing websites are short-lived and rapidly changing, yet many phishing datasets reduce observations to URLs or precomputed features, constraining researchers to predefined representations and discarding the underlying evidence needed to derive alternative features, apply new extraction methods, examine cross-layer relationships, and reanalyze observations as phishing techniques evolve. This study addresses this limitation with a multi-layer active-web dataset comprising 67,502 scans, including 33,387 phishing observations from operational feeds and 34,115 screened benign reference observations. The corpus preserves raw evidence across HTML content and screenshots, URL and redirect behavior, HTTP and security headers, compliance files, TLS certificates, DNS and domain registration, open ports, geolocation and accessibility measurements, and network infrastructure, while explicitly recording unavailable evidence rather than treating it as negative observations. To avoid misleading infrastructure attribution on shared platforms, the study applies a hosting-aware evidence model that masks provider-owned infrastructure signals for free-hosted tenant pages while retaining meaningful page- and transport-level evidence. Characterization reveals systematic differences between phishing and benign websites across web-resource usage, domain maturity, mail and policy configuration, security headers, and infrastructure context. By preserving raw artifacts together with acquisition metadata and explicit evidence availability, the corpus provides an inspectable and reproducible foundation for future phishing measurement and dataset research.