NaiAD: A Comprehensive Dataset for Data-Driven LLM Advertising Research
Researchers have introduced NaiAD, the first comprehensive dataset designed specifically for Large Language Model (LLM) native advertising. Addressing the challenge of balancing platform revenue with user experience, NaiAD comprises 58,999 carefully constructed ad-embedded responses paired with user queries. The dataset is structured around theoretically grounded metrics that separately evaluate user and commercial utility. To overcome dimensional collinearity in aligned LLMs, the team proposes a decoupled generation pipeline that produces structurally diverse samples, ranging from those that disentangle stakeholder utilities to those with uniform strength across dimensions. Additionally, score labels are calibrated using a Variance-Calibrated Prediction-Powered Inference (VC-PPI) framework to align automated scoring with human annotations. Mechanistic analyses indicate that successful ad integration relies on four distinct semantic strategies. Models utilizing NaiAD can internalize these strategies to simultaneously enhance both user and commercial value while allowing independent control via in-context learning. This initiative establishes NaiAD as foundational infrastructure for the future development of LLM-native advertising systems, marking a significant step in data-driven AI research.
Wire timeline
NaiAD: A Comprehensive Dataset for Data-Driven LLM Advertising Research
Researchers have introduced NaiAD, the first comprehensive dataset designed specifically for Large Language Model (LLM) native advertising. Addressing the challenge of balancing platform revenue with user experience, NaiAD comprises 58,999 carefully constructed ad-embedded responses paired with user queries. The dataset is structured around theoretically grounded metrics that separately evaluate user and commercial utility. To overcome dimensional collinearity in aligned LLMs, the team proposes a decoupled generation pipeline that produces structurally diverse samples, ranging from those that disentangle stakeholder utilities to those with uniform strength across dimensions. Additionally, score labels are calibrated using a Variance-Calibrated Prediction-Powered Inference (VC-PPI) framework to align automated scoring with human annotations. Mechanistic analyses indicate that successful ad integration relies on four distinct semantic strategies. Models utilizing NaiAD can internalize these strategies to simultaneously enhance both user and commercial value while allowing independent control via in-context learning. This initiative establishes NaiAD as foundational infrastructure for the future development of LLM-native advertising systems, marking a significant step in data-driven AI research.
cs.AI updates on arXiv.org