Addressing Labelled Data Scarcity: Taxonomy-Agnostic Annotation of PII Values in HTTP Traffic using LLMs
Researchers from arXiv have published a new study addressing the scarcity of manually labelled data in automated privacy audits for web and mobile applications. Current detectors for Personally Identifiable Information (PII) leakage in HTTP traffic are often limited by fixed label taxonomies and dependence on scarce real-user data. This paper introduces a multi-stage pipeline leveraging Large Language Models (LLMs) to perform taxonomy-agnostic annotation of PII values within HTTP message bodies. The proposed method combines deterministic pre-processing with label-level classification, instance-level value annotation, and output validation. To facilitate evaluation without compromising user privacy, the authors also developed an LLM-based generator for synthetic HTTP traffic with validated PII annotations. Evaluations across three distinct PII taxonomies demonstrate that the pipeline accurately detects PII types and extracts corresponding values. The findings suggest that LLMs offer a flexible foundation for traffic annotation and labelled data creation, enabling better adaptability to evolving privacy definitions and domains without relying on sensitive real-world captures.
Wire timeline
Addressing Labelled Data Scarcity: Taxonomy-Agnostic Annotation of PII Values in HTTP Traffic using LLMs
Researchers from arXiv have published a new study addressing the scarcity of manually labelled data in automated privacy audits for web and mobile applications. Current detectors for Personally Identifiable Information (PII) leakage in HTTP traffic are often limited by fixed label taxonomies and dependence on scarce real-user data. This paper introduces a multi-stage pipeline leveraging Large Language Models (LLMs) to perform taxonomy-agnostic annotation of PII values within HTTP message bodies. The proposed method combines deterministic pre-processing with label-level classification, instance-level value annotation, and output validation. To facilitate evaluation without compromising user privacy, the authors also developed an LLM-based generator for synthetic HTTP traffic with validated PII annotations. Evaluations across three distinct PII taxonomies demonstrate that the pipeline accurately detects PII types and extracts corresponding values. The findings suggest that LLMs offer a flexible foundation for traffic annotation and labelled data creation, enabling better adaptability to evolving privacy definitions and domains without relying on sensitive real-world captures.
cs.AI updates on arXiv.org