Large language models identify machine learning task types from datasets

Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation

Machine LearningArtificial Intelligence

Summary

Knowing the type of machine learning task and data is important to build correct ML models but is usually done by hand. The authors tested whether large language models can figure this out automatically from dataset information when only the main feature to predict is given. They created a new set of 625 public datasets to evaluate this idea. Their results show that these models do better than existing automatic tools, especially when handling different kinds of data and when using smaller models for local use.

What this means in practice

  • For automl platform developers: Enhance automated pipeline tools by integrating language models to automatically identify task types from dataset features.
  • For data engineering teams: Use lightweight local language models to classify prediction tasks from datasets in resource-limited environments.

Authors

Petros Tsialis, Steffen Limmer, Tobias Rodemann, Martin Heckmann

Abstract

Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the target feature is provided by the user. Together with our LLM-based system we also release an annotated benchmark comprising 625 public tabular and time series datasets. We evaluate the proposed approach in three settings: (i) tabular datasets in comparison with established AutoML heuristics, (ii) cross-domain evaluation across tabular and time series datasets, and (iii) a practical deployment scenario using smaller local models. The results show consistent advantages for LLM-based task type identification, with increasing difficulty in heterogeneous and resource-constrained settings. LLM-based approaches outperform AutoGluon in the tabular setting, reaching 0.98 F1 macro compared to 0.93. In the cross-domain setting, the best model achieves 0.90 F1 macro, while smaller locally deployable models reach 0.75, indicating a trade-off between deployment feasibility and accuracy.