Data Language Models: A New Foundation Model Class for Tabular Data
Researchers have introduced the Data Language Model (DLM), a new class of foundation models designed to natively understand tabular data without requiring preprocessing or serialization. Unlike current approaches such as gradient-boosted trees or existing tabular foundation models, DLMs process raw cell values directly, similar to how language models interpret text. The paper presents Schema-1, the first DLM, featuring 140 million parameters and trained on over 2.3 million synthetic and real-world datasets. Schema-1 demonstrates superior performance compared to gradient-boosted ensembles, AutoML stacks, and other tabular models in row-level prediction benchmarks. It also achieves lower reconstruction errors in missing value imputation than classical statistical methods and large language models, highlighting the importance of structural understanding over encoded world knowledge. Additionally, Schema-1 can reliably identify industry sectors from unseen datasets, a capability previously unavailable in tabular models. This innovation aims to eliminate preprocessing pipelines, providing a native tabular understanding layer for AI systems, agents, and vertical applications.
Wire timeline
Data Language Models: A New Foundation Model Class for Tabular Data
Researchers have introduced the Data Language Model (DLM), a new class of foundation models designed to natively understand tabular data without requiring preprocessing or serialization. Unlike current approaches such as gradient-boosted trees or existing tabular foundation models, DLMs process raw cell values directly, similar to how language models interpret text. The paper presents Schema-1, the first DLM, featuring 140 million parameters and trained on over 2.3 million synthetic and real-world datasets. Schema-1 demonstrates superior performance compared to gradient-boosted ensembles, AutoML stacks, and other tabular models in row-level prediction benchmarks. It also achieves lower reconstruction errors in missing value imputation than classical statistical methods and large language models, highlighting the importance of structural understanding over encoded world knowledge. Additionally, Schema-1 can reliably identify industry sectors from unseen datasets, a capability previously unavailable in tabular models. This innovation aims to eliminate preprocessing pipelines, providing a native tabular understanding layer for AI systems, agents, and vertical applications.
cs.AI updates on arXiv.org