STELLA: Efficient Sensor-to-LLM Translation for On-Device Human Activity Recognition

2026-07-03Machine Learning

Machine LearningArtificial Intelligence
AI summary

The authors present STELLA, a method that helps wearable devices recognize human activities efficiently by converting sensor data into a smaller, manageable form for large language models (LLMs). Instead of changing the LLM itself, STELLA uses a lightweight tokenizer to simplify sensor signals while preserving important patterns. This lets the system run fully on the device, keeping data private and responding quickly. They also show STELLA can personalize the recognition for each user by updating only the tokenizer with small user-specific data. Tests on multiple datasets show STELLA improves accuracy and runs in real-time on mobile hardware.

Human Activity RecognitionEdge DevicesLarge Language ModelsSensor TokenizationInertial SensorsOn-device PersonalizationLatencyPrivacyF1 ScoreTime-series Data
Authors
Nirhoshan Sivaroopan, Albert Zomaya, Kanchana Thilakarathna
Abstract
HAR is increasingly expected to run continuously on edge devices, yet recent LLM-based methods remain hard to deploy: raw sensor prompts are long, cloud inference adds latency and privacy risk, and fine-tuned LLM pipelines turn general-purpose models into task-specific classifiers. We present STELLA, an efficient sensor-to-LLM translation framework for on-device HAR that shifts the burden from LLM adaptation to sensor tokenization. A lightweight hierarchical tokenizer compresses an entire multi-channel inertial window into a fixed set of compact latent sensor tokens, which are projected into the embedding space of a frozen pretrained LLM and combined with a natural-language prompt for label scoring. This preserves activity-relevant temporal and cross-channel structure while keeping LLM-side computation predictable across sensor configurations. STELLA also supports on-device personalization, adapting only the lightweight tokenizer on small amounts of user-specific labelled data and augmenting inference with a local retrieval context, keeping the LLM, user data, and retrieval on device. Across seven public HAR datasets and eight benchmark settings, STELLA achieves new state-of-the-art performance, improving over prior methods by up to 11.83% F1; on-device personalization yields up to a further 21.91% F1 as user data accumulates after deployment. STELLA also outperforms representative time-series tokenizers under the same LLM pipeline and achieves real-time inference under practical mobile and edge budgets, showing that efficient sensor tokenization is a practical path toward accurate, private, and personalized LLM-based HAR on edge devices.