Back to papers
March 19, 2026cs.LGcs.AIstat.MEIntermediate
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
An Luo, Jin Du, Xun Xian, Robert Specht, Fangqiao Tian, Ganghua Wang, Xuan Bi, Charles Fleming, Ashish Kundu, Jayanth Srinivasa, Mingyi Hong, Rui Zhang, Tianxi Li, Galin Jones, Jie Ding
AI-Generated Summary
AgentDS is a benchmark that tests how well AI agents and humans working together can solve real-world data science problems across industries like healthcare, retail, and manufacturing. The research found that while AI agents alone perform poorly on these domain-specific tasks, human-AI collaboration produces the best results, showing that human expertise remains crucial even as AI becomes more advanced.
HF Upvotes
4
Difficulty
Intermediate
Categories
cs.LG, cs.AI, stat.ME
AI Tags
AI agentsbenchmarkinghuman-AI collaborationlarge language modelsdata sciencedomain-specific tasksautomation