AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques

2026-08-10Human-Computer Interaction

Human-Computer InteractionMultimedia
AI summary

The authors address challenges in learning from long audio and video materials by creating three AI tools that help learners consume, understand, and imitate content more efficiently. One tool speeds up speech playback based on how hard it is to understand, another summarizes videos without losing key information, and the last one helps with pronunciation practice by analyzing speech and visuals. Their tests show these tools save time and improve learning without sacrificing comprehension. This work offers ways for AI to make skill learning faster and more interactive.

speech recognitionplayback speed adjustmentvideo summarizationmultimodal learningpronunciation practicedeep learningacoustic modelingskill acquisitionconfidence intervalsphoneme
Authors
Kazuki Kawamura
Abstract
Audio and video have become major learning media, but learners face two persistent challenges: the time cost of consuming long-form content sequentially and the lack of scalable feedback for imitation-based skill acquisition. This dissertation proposes an AI-guided learning framework that supports three interconnected stages: Consume, Understand, and Imitate. It develops and evaluates three systems. AIxSpeed dynamically adjusts audio playback speed at the phoneme level using speech-recognition-model confidence as a proxy for listening difficulty. FastPerson generates multimodal video summaries that preserve visual and auditory information and lets learners switch between summarized and full versions by chapter. Profy learns proficiency from largely unannotated speech data and visualizes classifier-relevant regions and model-derived acoustic distances to support pronunciation practice. Technical and user evaluations show that AIxSpeed achieved average playback factors of 1.30x on LibriSpeech and 1.29x on UME-ERJ and received higher mean opinion scores than matched constant-speed playback; FastPerson reduced viewing time by 53% with no statistically significant difference in quiz scores compared with normal playback; and Profy showed an observed improvement in pronunciation intelligibility, with non-overlapping pre- and post-practice confidence intervals. Together, these systems demonstrate how deep learning can support efficient content consumption, multimodal understanding, and repeated skill practice while retaining learner access to the original material.