A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes

2026-08-24Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionHuman-Computer Interaction
AI summary

The authors created a new dataset called 3DTongueQA that connects muscle activity in the tongue to detailed 3D tongue shapes using a simulation model. This allows them to track how specific muscle activations cause tongue movements, which earlier real-world imaging methods couldn’t label directly. They tested their data with machine learning models that predict muscle states and tongue shapes, showing good accuracy even with variations like different languages. Their work supports building systems that understand tongue movement in both structured formats and natural language questions, without being limited to one model type.

articulatory corpusreal-time MRIelectromagnetic articulographytongue muscle activationfinite-element modelArtiSynth3D tongue meshbiomechanical simulationstructured predictionnatural language question answering
Authors
Seungho Eum, Unsang Park
Abstract
Existing articulatory corpora based on real-time MRI and electromagnetic articulography capture tongue shape and motion but do not provide traceable labels for the muscle-driven process that generated an observed configuration. We introduce a simulator-grounded data-construction framework and instantiate it as 3DTongueQA. Controlled 11-dimensional muscle activations are mapped to fixed-topology tongue meshes with the ArtiSynth Badin finite-element model, converted into structured biomechanical records, and rendered as deterministic QA on muscle state, geometry, and target-directed change. We screen 295,157 configurations, retain 295,115 valid meshes, and construct 891,156 QA records per language. Language naturalization changes only surface form and is verified against the source records; English and Korean instantiations demonstrate construction-level portability. A swappable SpiralNet++--Qwen3-8B baseline reaches 62.9 $\pm$ 9.2 Muscle EM, 74.0 $\pm$ 0.2 Value Accuracy, and 65.9 $\pm$ 4.7 Direction EM, while mismatching the paired mesh reduces Muscle EM to 2.2; a dataset-leakage-controlled anchor-held-out model retains 80.4--98.6\% of the full-inventory scores on unseen anchors. Task-specific structured readouts further reach 88.7 $\pm$ 0.7 Muscle EM and 93.3 $\pm$ 1.0 Direction EM. These complementary results show that the constructed supervision supports both efficient structured prediction and heterogeneous natural-language QA rather than being tied to a particular decoder architecture.