How Reliable Are Multimodal Signals of Conversational State? Evidence from Remote Dyadic Collaborative Tasks
2026-07-20 • Computation and Language
Computation and Language
AI summaryⓘ
The authors studied how to measure mental effort and power dynamics in conversations using features from speech, language, and interaction during video calls. They found that language features best predict mental effort but only work well within the same task, while acoustic features often reflect who is speaking rather than the conversation itself. Interaction-based features were the most reliable across different tasks and speakers. Also, detecting power roles from overall conversation patterns was very difficult. The authors suggest carefully adjusting for speaker differences and using multiple evaluation methods to choose good conversation features.
cognitive loadconversational powermultimodal featurespredictive accuracycross-task generalizabilitytest-retest reliabilitylinguistic featuresacoustic featuresinteraction featuresspeaker normalization
Authors
Tahiya Chowdhury
Abstract
Measuring conversational states such as cognitive load and conversational power from multimodal behavior requires characteristic features that are not only predictive but also reliable across task contexts. We present a three-dimensional evaluation framework assessing predictive accuracy, cross-task generalizability, and test-retest reliability, applied to interactional, acoustic, and linguistic features extracted from dyadic conversations during collaborative tasks performed over a video-conferencing platform (AVCAffe dataset; 53 dyads, 9 tasks). Our results show that no single feature family dominates all three dimensions. Linguistic features show the highest predictive accuracy for cognitive load but collapse under cross-task evaluation, revealing sensitivity to task-specific vocabulary. Additionally, acoustic reliability, often reported as evidence of feature stability, degrades once speaker identity is controlled, confirming that standard prosodic features measure vocal characteristics rather than conversational state. Interaction features provide the only genuinely reliable signal, unchanged after speaker normalization. Interestingly, classifying power role remained near chance baseline across all conditions, indicating limitations of task-level aggregated behavior for predicting power role in conversation. Our findings reveal three insights: (1) linguistic features predict best but generalize poorly across task contexts; (2) acoustic reliability collapses to near-zero once speaker identity is controlled, challenging standard evaluation practice; and (3) interaction features provide the only genuinely reliable signal, with floor dominance predicting within-dyad cognitive load asymmetry. These results argue for speaker normalization and multi-dimensional evaluation as prerequisites for context-aware, robust multimodal feature selection in conversational systems.