Dataforge: Agentic Platform for Autonomous Data Engineering
Xinyuan Wang, Hongyu Cao, Kunpeng Liu, Yanjie Fu
arXiv:2511.06185·cs.AI·Published 2025-11-09·Updated 2026-02-16
The growing demand for artificial intelligence (AI) applications in materials discovery, molecular modeling, and climate science has made data preparation a critical but labor-intensive bottleneck. Raw data from diverse sources must be cleaned, normalized, and transformed to become AI-ready, where effective feature transformation and selection are essential for robust learning. We present Dataforge, an LLM-powered agentic data engineering platform for tabular data that is automatic, safe, and non-expert friendly. It autonomously performs data cleaning and iteratively optimizes feature operations under a budgeted feedback loop with automatic stopping. Across tabular benchmarks, it achieves the best overall downstream performance; ablations further confirm the roles of routing/iterative refinement and grounding in accuracy and reliability. Dataforge demonstrates a practical path toward autonomous data agents that transform raw data from data to better data.
TopicsLarge Language Models & Materials
Tagsmaterials-discovery
arXiv categoriescs.AI
arXiv abstract pagePDF