Pretrained model infers programs fast to generate and analyze data

From Data to Program: Fast & Direct Generative Program Inference from Empirical Data

Machine LearningArtificial Intelligence

Summary

Creating models that understand data usually needs fitting them to that specific data, which can take time. The authors introduce PRODiGI, a model trained beforehand that can quickly turn data into explicit programs you can run to generate or analyze similar data. These programs are easy to inspect and use without relying on the original model. PRODiGI also allows fine-tuning to better match new data, improving quality. This approach speeds up data modeling while keeping results interpretable.

What this means in practice

  • For data engineers: Generate reusable programs from data that speed up sampling and density evaluation for database-driven applications.
  • For machine learning practitioners: Quickly produce explicit generative programs from datasets to better understand and adjust data distributions during model training.

Authors

Simon Klüttermann, Xueying Ding, Leman Akoglu

Abstract

Estimating probability densities from a finite set of samples typically requires dataset-specific model fitting. We introduce PRODiGI, a pretrained data-to-program model that infers an explicit, executable generative program in a single forward pass. Pretrained on synthetic datasets paired with their ground-truth programs, PRODiGI accommodates diverse generative families and data dimensionalities through template prediction and non-autoregressive program parameter decoding. Its inferred programs support direct sampling, density and score evaluation, and inspection independently of the pretrained model. We further introduce program-space fine-tuning, which refines differentiable program parameters by matching generated and empirical samples while keeping model parameters intact. Experiments show that PRODiGI achieves lower average density and score MAE than existing pretrained models, while offering multi-fold speedups over its closest competitors. Program-space fine-tuning further reduces generation MMD by 84%. By turning empirical data into explicit, reusable programs, PRODiGI introduces a new direction for fast, interpretable tabular generative modeling.