Large language models enable privacy-safe mental distress prediction across surveys

LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys

Machine LearningArtificial IntelligenceCryptography and Security

Summary

Mental health surveys vary a lot, making it hard to combine data from different places while keeping people’s answers private. The authors created a system using a large language model that turns different survey questions into a common language format. This lets groups build better mental distress prediction tools without sharing raw data. Their approach is more accurate and needs much less computing power on clients’ side than other methods.

What this means in practice

  • For mental health data teams: Combine diverse mental health survey data from multiple institutions without sharing raw responses to improve prediction models.
  • For privacy engineers: Deploy split learning with language model encoders to protect survey data privacy while enabling efficient model training.

Authors

Md Khalid Syfullah, Alvi Ataur Khalil

Abstract

Rising societal and lifestyle complexity has been linked to a growing prevalence of mental distress worldwide. Educational institutions, workplaces, clinics, etc. collect large volumes of mental health survey data to understand and reduce this burden. Collaborative analysis of such data could yield effective generalizable predictive models. Privacy constraints and varied survey designs (i.e., different questions, scales, and formats) hinder direct integration. We propose a schema-aware split learning (SL) framework that preserves privacy, using a large language model (LLM) as a shared semantic encoder to harmonize heterogeneous survey schemas across institutions. We serialize each survey record into a natural-language description, unifying disparate survey schemas into a common format. The LLM is fine-tuned for mental distress assessment via Low-Rank Adaptation (LoRA) and partitioned across client and server. Clients retain the raw survey responses locally and run only a lightweight front-end, so original records never leave the institution that collected them. The resource-intensive backbone runs on the server, minimizing client-side computation. Using LLaMA-3.2-3B-Instruct, the framework attains an average ANLS of 0.708 with only 2,000 training samples, surpasses federated learning (FL) in eight of nine settings, and cuts per-client computation by three orders of magnitude, while generalizing to unseen datasets. Overall, it enables accurate, privacy-preserving, and resource-efficient collaborative learning from heterogeneous mental health survey data.