Aspect based sentiment analysis improves with typed decisions on multiple languages
Decide, Don't Generate: Competitive Dimensional ABSA with Jev's Typed Decisions
Computation and Language
Summary
Aspect-based sentiment analysis (ABSA) helps computers understand what people think about specific parts of products or services in text. Usually, this involves generating text, which can be complex and slow. The authors show that simply making yes/no and score decisions, without generating text, works very well for this task. They use a fixed model that answers typed questions and calibrate it with some added coefficients, achieving better results than existing big models. This approach works across several languages and reduces errors in understanding feelings and extracting related info.
What this means in practice
- •For customer experience teams: Automate detailed sentiment detection about product features in multiple languages without complex text generation models.
- •For sentiment analysis software developers: Build efficient aspect sentiment systems by integrating typed decision outputs from frozen models calibrated with lightweight coefficients, improving accuracy and speed.
Authors
Yiqun Zhang, Peidong Wang, Zihan Wang, Shi Feng
Abstract
Aspect-based sentiment analysis (ABSA) has largely turned to text generation. We show that competitive dimensional ABSA does not need it. Using Jev, a frozen model that answers typed questions with rubric scores, label probabilities, and yes/no judgments, we decompose all three tasks of SemEval-2026 Task III Track A into such decisions and align them with the annotation scheme through 488 coefficients fitted on CPU, with no text generation and no backbone tuning. On valence-arousal regression over ten corpora in six languages, the system reaches 1.0645 RMSE, the lowest aggregate error of any participating system. On triplet and quadruplet extraction, it reaches 52.09 and 44.06 continuous F1, above fine-tuned Llama-3.3-70B and GPT-OSS-120B baselines. Analyses and ablations show where the accuracy comes from: supervised calibration roughly halves the raw regression error, exact valence-arousal would add only 4.5 F1 to extraction, and the learned combination of span-boundary evidence, not any single signal, carries the extraction systems.