Fuzzy modeling improves synthetic data with clear reasoning and privacy
Fuzzy Distribution Modeling for Synthetic Tabular Data Generation with Causality Preservation
Artificial Intelligence
Summary
When real-world data is hard to get, making up fake data helps train AI safely. The authors show a new way to create this fake data using fuzzy logic, which handles unclear or missing information well. Their method makes it easy to understand how the data was made and keeps important connections between different features. It works well across tests, keeping data useful and private while letting users ask 'what if' questions.
What this means in practice
- •For medical data teams: Generate realistic patient data while preserving privacy and interpretability for training healthcare AI models when real data is limited.
- •For financial risk analysts: Create synthetic financial records that respect complex dependencies and rules, enabling safer risk modeling without exposing sensitive client data.
Authors
Michael Vasilakakis, Dimitris K. Iakovidis
Abstract
Synthetic tabular data generation provides an effective alternative for the training of machine learning models when real-world data is limited or inaccessible. However, the heterogeneous, non-smooth, and incomplete nature of tabular data poses fundamental challenges to conventional probabilistic and deep generative models, where their interpretability remains limited. This paper proposes a novel fuzzy distribution modeling methodology for synthetic tabular data generation based on fuzzy sets theory. Feature distributions are represented using fuzzy sets and feature dependencies are modeled through Fuzzy Cognitive Maps, resulting in a low-parameter, and an interpretable data representation. Synthetic samples are generated by sampling fuzzy concepts rather than raw values, enabling native support for mixed data types, missing values, and domain constraints. The methodology further supports linguistic queries and IF-THEN reasoning, facilitating transparent simulation of decision-making processes. Experimental results on benchmark datasets demonstrate competitive performance with respect to utility, fidelity and privacy compared to state-of-the-art methods, while offering substantially improved interpretability. These results establish fuzzy distribution modeling as a principled and effective approach for synthetic tabular data generation in fuzzy systems and decision support applications.