PrepBench: Evaluating the Gap in Natural-Language-Driven Data Preparation
Researchers have introduced PrepBench, a new benchmark designed to evaluate the capabilities of Large Language Models (LLMs) in natural language-driven data preparation. While traditional tools rely on graphical user interfaces, LLMs offer a potential paradigm shift by allowing users to specify data transformation intents directly through natural language. However, existing code generation benchmarks fail to address key challenges such as ambiguous user intents, imperfect real-world data, and the need for interpretable workflow validation. PrepBench addresses these gaps by assessing three core capabilities: interactive disambiguation, preparation code generation, and code-to-workflow translation. Derived from the Preppin' Data Challenges, the benchmark features diverse domains with tasks requiring complex Python solutions, often exceeding 100 lines of code. Evaluation results indicate that despite recent advancements in AI, state-of-the-art LLMs still struggle to fully realize this paradigm shift. PrepBench serves as a principled tool for measuring current limitations and identifying critical challenges in automating data preparation workflows through natural language interactions.
Wire timeline
PrepBench: Evaluating the Gap in Natural-Language-Driven Data Preparation
Researchers have introduced PrepBench, a new benchmark designed to evaluate the capabilities of Large Language Models (LLMs) in natural language-driven data preparation. While traditional tools rely on graphical user interfaces, LLMs offer a potential paradigm shift by allowing users to specify data transformation intents directly through natural language. However, existing code generation benchmarks fail to address key challenges such as ambiguous user intents, imperfect real-world data, and the need for interpretable workflow validation. PrepBench addresses these gaps by assessing three core capabilities: interactive disambiguation, preparation code generation, and code-to-workflow translation. Derived from the Preppin' Data Challenges, the benchmark features diverse domains with tasks requiring complex Python solutions, often exceeding 100 lines of code. Evaluation results indicate that despite recent advancements in AI, state-of-the-art LLMs still struggle to fully realize this paradigm shift. PrepBench serves as a principled tool for measuring current limitations and identifying critical challenges in automating data preparation workflows through natural language interactions.
cs.AI updates on arXiv.org