Dataforge: Agentic Platform for Autonomous Data Engineering

Xinyuan Wang, Hongyu Cao, Kunpeng Liu, Yanjie Fu

arXiv:2511.06185·cs.AI·Published 2025-11-09·Updated 2026-02-16

The growing demand for artificial intelligence (AI) applications in materials discovery, molecular modeling, and climate science has made data preparation a critical but labor-intensive bottleneck. Raw data from diverse sources must be cleaned, normalized, and transformed to become AI-ready, where effective feature transformation and selection are essential for robust learning. We present Dataforge, an LLM-powered agentic data engineering platform for tabular data that is automatic, safe, and non-expert friendly. It autonomously performs data cleaning and iteratively optimizes feature operations under a budgeted feedback loop with automatic stopping. Across tabular benchmarks, it achieves the best overall downstream performance; ablations further confirm the roles of routing/iterative refinement and grounding in accuracy and reliability. Dataforge demonstrates a practical path toward autonomous data agents that transform raw data from data to better data.

TopicsMaterials Science & Condensed Matter

Tagsmaterials-discovery

arXiv categoriescs.AI

arXiv abstract pagePDF