Week beginning 7 September 2026

Commercial implications of this week’s papers

289 of this week’s computer science papers describe a result that could become a product or a service. Open a paper for the summary and the full list of uses.

The headline and the lines under it are written by an AI system from the abstract. They can be incomplete or wrong. The paper is the authority.

Image system improves bright highlights by extending dynamic range gradually

Recurrent Dynamic Range Extension

Commercial implications: For mobile camera app developers: Allows selling camera apps that produce superior HDR photos on common devices by gradual brightening techniques.

Fri 11 SeptComputer Vision and Pattern RecognitionGraphics
The gist
Capturing very bright parts of a scene well is hard with normal cameras. The authors created a method that improves images step-by-step by brightening them a little at a time. Their system learns how to enhance images gradually and can be run multiple times to reveal details in extremely bright areas. This approach works with common RAW photos and produces realistic results even in difficult lighting. They also made the training smarter by remembering past improvements to reduce mistakes.
Open 2609.13135v1

UID preserving method improves clinical event timelines from discharge summaries

Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication

Commercial implications: For medical software developers: Enables advanced clinical timeline features that can be sold as enhancements to electronic health record systems.

Fri 11 SeptArtificial Intelligence
The gist
Medical records often list events out of order or miss details, making it hard to understand a patient's treatment timeline. The authors created a system that links each event in a doctor's notes back to its source and uses both text and record data to build accurate timelines. They also built a tool that checks and compares these timelines against the original records. Their approach recovered many more events and matched expert doctors’ timeline assessments closely.
Open 2609.13062v1

Kraken improves speech translation by using low-bitrate tokens and source speech input

Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning

Commercial implications: For translation service developers: Enables new commercial voice translation products with improved naturalness and speaker fidelity for global communication tools.

Fri 11 SeptComputation and LanguageSound
The gist
Speech-to-speech translation systems help convert spoken words in one language to another while keeping the speaker's voice and tone. The authors show that using simpler, compressed speech tokens makes it easier for large language models to process speech accurately. They also add a tool that recreates the sound while listening to the original speech for better voice and emotion preservation. Their approach, called Kraken, was trained on a large, diverse set of spoken languages and outperforms other similar models in translation quality and naturalness.
Open 2609.13045v1

Attention quantization speeds tabular foundation model inference without accuracy loss

Attention Quantization for Tabular Foundation Models

Commercial implications: For cloud service providers: This enables faster and cheaper cloud AI services specializing in tabular data tasks by reducing computational costs.

Fri 11 SeptMachine LearningArtificial Intelligence
The gist
Running big computer models on tables of data can be slow. The authors found that focusing on speeding up a specific part called attention, by using a method called quantization, makes these models run faster. They changed how parts of the model convert data to a simpler form without losing accuracy. Their approach made the models run up to 1.7 times faster while keeping their performance. This helps use these models more efficiently in real-world tasks.
Open 2609.13031v1

Cardiac MRI segmentation improved for rare single ventricle defects

SV-Cine: Diagnosis-Conditioned Segmentation of Single Ventricle Physiology via Generative Data Augmentation

Commercial implications: For medical device developers: It enables development of advanced imaging software for congenital heart disease diagnostics, supporting better treatment planning.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Single ventricle heart defects are rare and different in each patient, which makes it hard for computers to understand MRI images of their hearts. The authors created a new way to generate more heart images to teach computers, and they built a smarter system that understands the type of heart defect a patient has when analyzing images. Their system did better than previous methods at identifying and measuring parts of the heart in these patients. This helps doctors by providing more accurate information from MRI scans despite limited real patient data.
Open 2609.12997v1

Sign language translation improves using pose motion features and T5 models

Investigating Temporal Motion Features for Pose-to-Text Indian Sign Language Translation

Commercial implications: For mobile app developers: Enables sale of consumer-facing sign language translation apps with improved accuracy using novel motion feature integration.

Fri 11 SeptComputer Vision and Pattern RecognitionComputation and Language
The gist
Translating Indian Sign Language into English text can be tricky because it relies on understanding body movements. The authors tested how well different sizes of a language model called T5 worked when combined with pose data. They also added extra information about how the poses change over time, called motion features. This addition made the translations better, especially with the smallest model. Their approach ranked 5th in a recent competition and the code is available for others to use.
Open 2609.12993v1

Virtualized 5G network demo shows reliable throughput using OpenAirInterface

Virtualized 5G Tesbed using OpenAirInterface: Tutorial and Benchmarking Tests

Commercial implications: For industrial wireless integrators: Allows creating customizable private 5G solutions with adaptable software and hardware, enabling new product offerings for industrial clients.

Fri 11 SeptNetworking and Internet Architecture
The gist
Building and testing 5G networks usually requires expensive hardware and complex software. This paper explains how a free software platform called OpenAirInterface, combined with accessible radio hardware, can create a virtual 5G network. The authors provide step-by-step instructions and show that this setup delivers consistent and repeatable performance. This approach helps developers experiment with 5G technologies without relying on proprietary equipment.
Open 2609.12972v1

Generative method finds people from text descriptions without labels

Generative Retrieval for Unsupervised Text-Based Person Search

Commercial implications: For security system developers: Enables deployable person search tools that reduce costly manual labeling, useful for security product vendors.

Fri 11 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Finding pictures of a person from a description usually needs many labeled examples, which take a lot of work. The authors propose a new two-step approach that first creates detailed and varied text descriptions from unlabeled images, then uses these to train a system to match text to images more accurately. They also introduce a large new dataset with detailed text annotations to help improve this kind of search. Their experiments show this method works well even without any labeled training data.
Open 2609.12965v1

StepAudio 3 Gen enables versatile audio creation from text inputs

StepAudio 3 Gen Technical Report

Commercial implications: For game audio teams: Enables integrated audio pipelines for games that can sell richer sound experiences with fewer assets and less manual design.

Fri 11 SeptSound
The gist
Creating different kinds of audio like speech, singing, sound effects, and music all with one tool is hard. The authors built StepAudio 3 Gen, which can create many types of sounds from text without needing extra training for each audio type. It uses a new way to break down and predict audio pieces that keeps both meaning and sound details. This method lets StepAudio 3 Gen produce high-quality voices and other sounds all in a single system.
Open 2609.12945v1

Parallel training improves covid diagnosis model speed and accuracy

Parallel Training Using a CNN-DNN Architecture for Accelerated Development of Diagnostic Models

Commercial implications: For medical imaging companies: Enables development of commercial diagnostic software optimized for speed and accuracy using large medical image data.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Training deep learning models to detect diseases in medical images can take a long time, especially with large and growing datasets like during a pandemic. The authors collected CT scans from COVID-19 and other pneumonia patients and used a combined CNN-DNN approach that breaks images into parts to speed up training by running in parallel. This method achieved similar or better accuracy while reducing training time significantly. Their approach helps create diagnostic tools faster, which can be important in urgent health situations.
Open 2609.12902v1

Robot estimates 3d center of mass from gentle push on unknown objects

Before the Tipping Point: Force-Guided Active Perception for Shape-Agnostic Estimation of 3D Centers of Mass

Commercial implications: For automated warehouse teams: Enables development of robotic warehouse systems that can reliably handle diverse, unmodeled inventory with minimal manual intervention.

Fri 11 SeptRobotics
The gist
It is hard for robots to figure out the balance point and weight of strange or oddly shaped objects without picking them up. The authors came up with a method where a robot gently pushes an object just enough to tip it slightly, then helps it settle back. By measuring the forces and angles during this motion, the robot can calculate the object's center of mass height and weight with good accuracy. This process works even without knowing the object’s shape beforehand, and it avoids making the object fall over completely.
Open 2609.12894v1

Biomedical image segmentation improved by focusing on uncertain boundaries

Beyond Accuracy: Uncertainty-Guided Boundary Refinement for Reliable Biomedical Image Segmentation

Commercial implications: For medical device developers: Enables devices with improved trustworthiness in boundary detection for clinical diagnostics, a competitive feature for medical imaging products.

Fri 11 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Segmenting biomedical images accurately means not just labeling regions correctly but also precisely outlining important structures like cells. The authors propose a method that first makes a standard prediction and then refines only the uncertain boundary areas based on multiple uncertainty measures. This approach slightly improves accuracy and boundary clarity in blood-smear images. Although it doesn’t improve overall confidence calibration, it provides a transparent way to trust and improve boundary segmentation.
Open 2609.12892v1

MGAvatar improves realistic head avatars with hybrid geometry representation

MGAvatar: Mesh-Bound Gaussians for Head Avatar Geometry and Appearance Modeling

Commercial implications: For game developers: Enables production of high-fidelity, personalized avatar assets enhancing player immersion and customization features in games.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Creating realistic 3D head models is hard because existing methods often rely on basic face shapes that miss personal details like hair or clothes. The authors developed MGAvatar, which combines a flexible mesh with Gaussian blobs to better capture detailed head shapes and appearances. They also introduced new ways to handle how the face changes with different expressions and views, making the avatars look more consistent and detailed. Tests showed MGAvatar generates higher quality and more lifelike head models than previous methods.
Open 2609.12850v1

Self-supervised pre-training improves eye disease models with little data

Self-supervised Pre-training Helps Retinal Disease Progression Modelling Most When Data Is Scarce

Commercial implications: For clinical software developers: Enables more accurate prognostic tools for AMD based on limited follow-up images, which can be marketed to eye care providers.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
It's hard to study how eye diseases like age-related macular degeneration (AMD) get worse over time because the right kinds of patient images are rare. The authors found that training a computer program first on many single images taken at one time (instead of changes over time) can help it learn useful information. When there isn’t much data showing disease progression, this pre-training helps the model predict AMD worsening better than starting from scratch. This improvement depends more on how the model is trained beforehand, rather than on the size or similarity of the training images.
Open 2609.12834v1

Image generation balances emotion and content through new reinforcement method

Balancing Emotional Alignment and Semantic Consistency in Image Generation via Reinforcement Learning with Valence-Arousal Anchoring

Commercial implications: For advertising creatives: Enables new marketing visuals with targeted emotions, increasing ad effectiveness and brand appeal.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Generating pictures that match both what they show and the feelings they should convey is tricky because changing the emotion can unintentionally change the content. The authors developed a new approach that uses a kind of learning called reinforcement learning combined with an emotional map based on feelings called valence and arousal. This approach helps the system create images that better match the intended emotions while keeping the scene and objects consistent. They tested it on thousands of prompts and demonstrated it improves emotional accuracy without losing important picture details.
Open 2609.12830v1

Dual attention AI improves cervical cancer screening with new dataset

A Dual Cross-Attention Framework for Colposcopic CIN Grading and Swede Score Prediction Using a New Multi-Center Dataset

Commercial implications: For medical imaging software developers: Enables creation of AI-powered colposcopy software products for clinical use addressing cervical cancer screening.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Cervical cancer screening can be hard because it needs expert doctors and the tests are often subjective. The authors created a smart computer method that looks at special images from multiple places to better grade the risk of cervical disease and give clinical scores. They also gathered a new collection of images and expert labels to help train and test this method. Their system works better than older ones, especially at telling disease severity and scoring important clinical factors. This could help doctors in places with fewer experts screen and treat patients more accurately.
Open 2609.12827v1

Lightweight method improves fusion of polarization and intensity images

LG-PF: Lightweight Confidence-Guided Polarization Image Fusion

Commercial implications: For imaging hardware developers: Enables advanced polarization image fusion with efficient hardware use, improving product capabilities for manufacturers of polarization imaging systems.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Polarization image fusion blends brightness and texture from regular images with special details that reveal material properties, but unreliable parts can cause problems. The authors propose LG-PF, a simple system that selectively adds trustworthy polarization details while avoiding unstable ones. They also created a new dataset to test this fusion across many kinds of scenes. LG-PF is efficient, requires little computing power, and works well across different environments without extra training.
Open 2609.12787v1

Automated tools enable video editing from simple previews to trailers

Unified Agentic Video Editing Across Levels of Complexity and Creativity

Commercial implications: For marketing teams: Enables production of customized promotional videos for brands and agencies, improving content marketing efficiency.

Fri 11 SeptArtificial IntelligenceHuman-Computer InteractionMultimedia
The gist
Editing videos can be a complex and creative process that involves many choices. The authors developed methods that allow automated tools to help edit videos in different ways, from short scene previews to full cinematic trailers. Their work explores how these tools can balance creativity and complexity in video editing. They also evaluated how well these automated edits work and considered what it means for the future of video editing.
Open 2609.12769v1

Vision language models assist spatial navigation for visually impaired users

Assisted Spatial Cognition Through Vision-Language Models

Commercial implications: For assistive technology developers: Enables new commercial navigation apps offering real-time, detailed spatial assistance for visually impaired and neuro-divergent users.

Fri 11 SeptArtificial Intelligence
The gist
Many people who are visually impaired or have different cognitive needs find it hard to understand and navigate their surroundings. This paper presents a new system that uses phone cameras to create detailed 3D maps of places in real time. The system combines smart language and vision AI models to describe spaces clearly and accurately, helping users understand where things are around them. The authors tested their system in many settings and showed it works well for giving detailed directions and object locations.
Open 2609.12747v1

Physics-guided synthetic ultrasound aids skin layer segmentation

Physics-Guided Synthetic High-Frequency Ultrasound Generation for Skin Layer Segmentation

Commercial implications: For dermatology device makers: Enables more accurate skin diagnostic tools by supplementing limited annotated ultrasound data, improving commercial skin imaging devices.

Fri 11 SeptMachine LearningComputer Vision and Pattern Recognition
The gist
High-frequency ultrasound helps see the layers of skin without cutting it, but there isn't enough detailed data showing all the skin layers. The authors created fake ultrasound images using physics-based models of skin layers to generate lots of training data. When a computer learned from these fake images and then from real ones, it got better at identifying skin layers in real ultrasound images. This approach shows promise but needs more work to make the fake images look more like real ones.
Open 2609.12735v1

Neural system separates contaminants to score muscle signal quality

Prism-SQA: An Interpretable and Adaptable Neural Framework for Surface Electromyography Quality Assessment

Commercial implications: For clinical device makers: Enables clearer and adaptable sEMG quality feedback in medical devices, improving user trust and regulatory acceptance.

Fri 11 SeptMachine Learning
The gist
Surface electromyography (sEMG) signals used in medical tests can get messed up by different types of noise, making analysis tricky. The authors developed Prism-SQA, a new method that breaks down these signals to pull out clean parts from several kinds of noise, showing exactly how each noise type affects quality. This helps doctors understand the quality ratings instead of relying on unclear black-box models, and lets them adjust quality rules without retraining. Their tests showed Prism-SQA works as well or better than existing methods while offering clearer explanations.
Open 2609.12724v1

Digital pen and ai capture handwriting on regular paper precisely

Write on Paper and Get the Online Digital Trace:\newline A New Era for Handwriting

Commercial implications: For note-taking app developers: Enables development of pen-based note apps that capture digital traces accurately, enhancing product features for consumers.

Fri 11 SeptMachine Learning
The gist
Writing on regular paper feels natural and helps people remember better, but turning that writing into a digital form usually needs special devices or paper. The authors designed a smart pen with sensors and programs that use artificial intelligence to track handwriting accurately on any paper. This system doesn't need extra setup or special surfaces, making it easier to write by hand and keep digital records at the same time. Their solution processes handwriting in real time, bridging traditional note-taking and digital capture.
Open 2609.12702v1

Pipeline generates personalized photorealistic advertising images using AI

I Am AdMan: A Pipeline for Automatic Generation of Personalized Advertising Imagery

Commercial implications: For advertising technology teams: Enables selling personalized ad generation services for marketing firms that want to automate visually customized campaigns.

Fri 11 SeptArtificial IntelligenceHuman-Computer Interaction
The gist
Personalized ads usually match products to people but often look generic. The authors created a system called AdMan that uses customer info and AI to make custom ad images that look real. These ads change based on who the customer is, the product, and can be made automatically at large scale. The system's quality varies depending on the product and the AI model used, showing both promise and current challenges. This work helps understand how AI might fully automate personalized advertising images someday.
Open 2609.12694v1

Dual path network improves wearable blood pressure estimation accuracy

SIFPBPNet: A Dual-Path Network for Wearable and Cuffless Blood Pressure Estimation via Individualized Steady-state Representation

Commercial implications: For wearable device developers: Enables the creation of consumer-friendly wearable blood pressure monitors that improve accuracy and personalization compared to current models.

Fri 11 SeptMachine Learning
The gist
Blood pressure is an important health measure but usually needs a cuff to check, which isn’t convenient all the time. This paper presents a new computer method that uses data from wearable devices to estimate blood pressure without a cuff by learning both steady and changing patterns in the data. Their method also learns from longer-term history to understand personal differences, making it more accurate than existing approaches. The authors tested it on a large dataset and showed it works better for predicting blood pressure continuously and comfortably.
Open 2609.12690v1

Construction health recommendations adapt to trust and worker needs

Personalized and Trust-Aware Health Recommendation Policies for a Construction Workplace

Commercial implications: For industrial wearable device makers: Enables development of smart wearable systems for workplaces that improve personalized health recommendations and compliance.

Fri 11 SeptInformation Retrieval
The gist
Construction workers face health risks like fatigue and heat stress that can harm their safety and productivity. The authors created a model that looks at how workers’ health and trust in health advice change over time, and how trust affects whether workers follow recommendations. Their approach uses smart strategies to decide when and how often to give health advice personalized to each worker. This helps balance keeping workers healthy, productive, and willing to trust advice at the same time.
Open 2609.12679v1

Detecting fake news videos using keyframes and evidence fusion

Detecting and Explaining Fake News Short Videos with Multimodal Content and Real-World Evidence

Commercial implications: For social media platforms: Enables selling fake news filtering tools to platforms needing to moderate user video content at scale.

Fri 11 SeptComputer Vision and Pattern RecognitionMultimedia
The gist
Short videos online can spread fake news quickly, and it can be hard to spot which videos are false. The authors developed a new way to pick important video frames by looking at changes in both pictures and text seen in the video. They then use two separate computer checkers, one that looks at the video content and another that finds real-world proof, to decide if the video is fake. Their system combines these checkers’ opinions to detect fake news videos better and explain why.
Open 2609.12678v1

Deepfake detection improves with pulse and face movement analysis

Beyond Ambiguous Visual Cues: Studying Physiological Disruptions and Cross-Modal Inconsistencies in Deepfake Videos

Commercial implications: For mobile app developers: Enables selling verification apps or services with enhanced detection of face-swapped videos leveraging joint physiological and behavioral cues.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Deepfake videos are fake videos often created by swapping or changing faces in videos. The authors show that these fake videos disrupt natural body signals like heartbeats seen in face color changes. They created a method that looks at both the heartbeat signals and facial movements together to better spot fakes. Their approach performs better than methods looking at just one of these signs and works well even on new datasets.
Open 2609.12668v1

Laughter helps people relive and reflect on positive moments

Reconstruction and Reflection of Positive Experiences through Resurfacing Laughter-indexed Everyday Moments

Commercial implications: For mobile app developers: Enables development of consumer well-being apps that passively collect meaningful moments for emotional reflection and memory enhancement.

Fri 11 SeptHuman-Computer Interaction
The gist
Many happy moments in daily life aren't recorded because people don’t think to save them. The authors found that laughter can act as a natural signal to mark these moments without much effort. They created a device called LaughAnchor that detects laughter and collects related information to help people remember and think about these enjoyable times later. Users said the system helped them feel the emotions again and understand their habits and relationships better. This approach lets people control how they keep and interpret their memories.
Open 2609.12642v1

Steerable full duplex speech models improve conversational control and timing

SteerDuplex: Steerable Duplex Speech Dialogue Models

Commercial implications: For voice assistant developers: Allows creation of more natural and user-controllable conversational agents for smart devices and customer service.

Fri 11 SeptArtificial IntelligenceComputation and Language
The gist
People want voice assistants and dialogue systems to talk more naturally and follow instructions about how to speak, like changing tone or speed. The authors found that current full-duplex speech models, which can listen and talk at the same time, were missing this ability to change their speaking style reliably. They created SteerDuplex, a model trained to adjust conversation style and timing based on user instructions, and tested it on a new benchmark called SteerBench. Their model showed big improvements in controlling voice style and handling turn-taking smoothly, though some issues remain with incomplete responses.
Open 2609.12623v1

TraceMind predicts how users absorb AI content during co-writing

TraceMind: Predicting User Information Uptake from Low-Cost Interaction Traces during Human-LLM Content Co-Generation

Commercial implications: For content collaboration platforms: Enables enhanced interactive AI writing assistants that ensure users comprehend generated content, increasing trust and usability.

Fri 11 SeptHuman-Computer Interaction
The gist
AI can help people write content together, but users might not fully understand or remember the AI-generated parts. The authors created TraceMind to track small pieces of information and predict if users recognize them, using simple data from how users interact with the writing process. They tested TraceMind with 62 people and found it worked better than other methods by looking at how users engage over time. This work helps build tools that know what users actually understand, not just what they accept from AI.
Open 2609.12600v1

Ultra-widefield octa dataset and network improve retinal vessel segmentation

An Ultra-Widefield Swept-Source OCTA Dataset and a Polar-Gated Mamba Network for Retinal Vessel Segmentation

Commercial implications: For ophthalmology software developers: This paper enables advanced retinal analysis software capable of better vessel detection for eye care providers.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Retinal scans cover large areas but lack good public data for analyzing blood vessels accurately. The authors created WOIVES, a new large set of retinal images with detailed vessel markings to help this problem. They also developed PG-Mamba, a model that looks at the images from different angles to better find blood vessels. This model worked better than others in measuring blood vessel features precisely.
Open 2609.12574v1

RoofLang enables AI to design faster large language model inference systems

RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems

Commercial implications: For cloud service providers: Enables selling faster and more responsive AI inference services to enterprises using advanced LLM architectures found by RoofLang.

Fri 11 SeptDistributed, Parallel, and Cluster ComputingArtificial Intelligence
The gist
Optimizing how large language models (LLMs) generate answers is tricky because current AI improvements rely on measuring existing software, which limits new ideas. To fix this, the authors created RoofLang, a specialized programming language that describes computing tasks in a way that allows AI to explore entirely new system designs. Using RoofLang, they found specific LLM designs that can be 3.5 to almost 40 times faster than others, mainly by improving memory use. The system also automatically discovered better designs that further boosted speed and responsiveness on real hardware.
Open 2609.12551v1

Meddies improves clinical data privacy with multilingual PII detection

Meddies-PII: A Multilingual Framework for Personally Identifiable Information Extraction in Clinical De-identification

Commercial implications: For healthcare software developers: Enables creation of privacy-compliant healthcare software products that automatically detect PII across numerous languages.

Fri 11 SeptComputation and LanguageArtificial Intelligence
The gist
Protecting patient privacy in medical records requires identifying personal information correctly. Creating real training data for this is expensive, and previous synthetic alternatives were less detailed. To solve this, the authors made a huge set of one million fake clinical documents in 17 languages, carefully checked for accuracy. They trained a model to find personal details in these documents and found it outperformed other methods across many test sets. They will share their dataset, model, and tools publicly for others to use.
Open 2609.12544v1

Graph transformer model improves molecular chirality identification

$\text{GSF-}χ$: Global Stereochemical Fields for Chiral Graph Transformers

Commercial implications: For pharmaceutical developers: Enables development of drug candidates with improved chiral selectivity and safety profiles through enhanced molecular encoding.

Fri 11 SeptMachine Learning
The gist
Molecules can have mirror-image forms called enantiomers that behave differently, especially in chiral environments like the human body. The authors introduced a new method called GSF-χ that helps computers better detect these subtle differences by considering global stereochemistry, rather than focusing on single atoms. Their method respects molecular symmetry and reflection properties to predict molecular behaviors more accurately. This model shows improved results on tasks involving molecular rotations and chiral signals compared to previous approaches.
Open 2609.12532v1

Mindspeller uses EEG and tasks to guide job role suggestions

Mindspeller Neuroprofiling. How task performance, EEG, and association evidence support O*NET-based role guidance

Commercial implications: For human resources technology developers: Enables new commercial career guidance products that integrate EEG-driven cognitive profiling for more tailored role suggestions.

Fri 11 SeptComputers and Society
The gist
Matching people to job roles is complicated, especially when trying to fit their cognitive strengths. The authors created Mindspeller, a tool that combines brainwave data from EEG and performance on mental tasks to suggest suitable job types. It also uses self-reports and word associations to explain motivations but relies on EEG and task results to choose roles. This approach is meant to support conversations about fit, not make hiring decisions or predict success.
Open 2609.12501v1

Clinical knowledge improves AI recognition of physiotherapy exercises

PhysioAI: Clinical Knowledge-Guided Semantic Supervision for Skeleton-Based Physiotherapy Action Recognition

Commercial implications: For physiotherapy device developers: Enables development of more accurate home physiotherapy tracking products that use AI trained with expert clinical input.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Tracking physiotherapy exercises automatically helps people do rehabilitation at home where a therapist can't always be present. Existing AI methods often struggle because rehab exercises are subtle and vary a lot, especially for people with motor impairments. The authors created PhysioAI, which teaches the AI using detailed clinical descriptions of exercises during training, helping it learn better. This approach improves how well the AI can recognize rehabilitation movements just from skeleton data without extra info when actually used.
Open 2609.12491v1

Adaptive design improves agent behavior in complex environments

Adaptive Agent Design

Commercial implications: For financial algorithm developers: Enables creation of advanced adaptive trading systems that better handle partial observability and complex market behaviors, offering competitive advantage.

Fri 11 SeptArtificial IntelligenceComputer Science and Game Theory
The gist
Some computer programs called agents have to make decisions based on past actions and observations, but the situations they face can be complicated and not follow simple rules. The authors studied how an agent can learn both how it moves between internal states and how it acts to get the best results using existing data. They showed that a learning method called soft Q-learning can find good solutions even when the environment does not follow simple assumptions, and they explored ways to improve how the agent changes its internal state transitions in partly observable settings. This work helps understand how to create smarter agents that adapt to complex and uncertain situations.
Open 2609.12486v1

Audio moment retrieval improved by better feature extraction and detection

Overview and Meta-Analysis of DCASE 2026 Challenge Task 6: Audio Moment Retrieval from Long Audio

Commercial implications: For audio content platforms: Enables selling advanced audio search features to podcast and media streaming services to enhance user engagement.

Fri 11 SeptSound
The gist
Audio moment retrieval means finding specific moments in a long audio recording based on a text query. The authors describe a challenge where teams tried to locate such moments using advanced computer methods. The results show that improving the way audio and text features are matched and detecting the right moment boundaries helped a lot. However, this problem is still difficult, and even the best systems only found about half of the correct moments.
Open 2609.12484v1

TripPattern improves text watermarks without hurting quality

TripPattern: A Pattern-based Text Watermarking Method for Large Language Models

Commercial implications: For ai service operators: Enables watermarking features integrated into AI-generated content services for compliance and authenticity verification.

Fri 11 SeptArtificial Intelligence
The gist
Large language models can produce text that might be hard to tell apart from human writing, so people create watermarks to detect machine-generated text. Existing watermark methods can make the text sound less natural because they push the model toward certain words. The authors propose TripPattern, which splits words into three groups and uses patterns between two groups while allowing neutral words freely. This keeps the text natural while embedding detectable patterns. Their tests show TripPattern keeps text quality good and still finds machine-generated text reliably.
Open 2609.12472v1

Adaptive layer reduces false wake ups in voice assistants

Not All Speech Is Intent: Adaptive Self-Correcting Inference Layer for Post-ASR False Wake-Up

Commercial implications: For smart home device manufacturers: Enables smarter voice control products with fewer errors, appealing directly to consumers seeking reliable smart home interactions.

Fri 11 SeptComputation and LanguageArtificial Intelligence
The gist
Sometimes voice assistants mistakenly think they’ve been called when a person actually didn’t mean to speak to them, causing annoying errors. The authors developed a new system called ASCIL that listens again after the assistant is triggered and uses clues like hesitation or silence to check if the wake-up was intentional. It learns from past mistakes and adjusts itself over time without needing people to label the errors manually. Their tests showed ASCIL decreased these false alarms by over half while keeping the assistant’s responses quick and accurate.
Open 2609.12469v1

Hieronym improves function naming in stripped binary code

Hieronym: Leveraging Hierarchical Multi-Source Information for Function Renaming in Stripped Binary

Commercial implications: For software security teams: Allows creation of advanced binary analysis products that offer better function name recovery for security audits.

Fri 11 SeptSoftware Engineering
The gist
Stripped binaries remove helpful function names to protect or reduce programs, making it hard for people to understand the code. Hieronym is a new tool that uses a large language model to suggest accurate and readable names for these functions by combining different information sources from the program. It works well across several computer architectures and optimization levels, making reverse engineering easier and more reliable. The authors also tested Hieronym on real malware, showing it can help in security situations.
Open 2609.12457v1

Glioma tumor changes forecasted using MRI anchored model updates

Observation-Anchored Selective Assimilation for Longitudinal Tumor-State Proxy Forecasting in Post-Treatment Glioma

Commercial implications: For clinical software developers: The method can be embedded in clinical imaging software to provide enhanced tumor forecasting features for hospitals and diagnostic centers.

Fri 11 SeptMachine LearningArtificial Intelligence
The gist
Doctors use MRI scans over time to track brain tumors after treatment, but it’s hard to predict how the tumor will change. The authors created a method that updates tumor predictions by carefully combining new MRI data with previous estimates. Their approach keeps observed tumor details as anchors and selectively adjusts other areas, which helps maintain prediction accuracy compared to simpler methods. This technique may support better monitoring of tumor changes over time using MRI images.
Open 2609.12435v1

Large synthetic dataset enables robots to fold and unfold t-shirts

FoldNet++: a Large-Scale Synthetic Dataset for Robotic T-Shirt Folding and Unfolding

Commercial implications: For automation integrators: Enables sale of adaptable laundry-folding robots for commercial laundries and apparel businesses, improving efficiency and reducing manual labor.

Fri 11 SeptRobotics
The gist
Folding and unfolding T-shirts is hard for robots because clothes are soft and change shape easily. The authors created a huge fake dataset with many different T-shirts, environments, and robot types to teach robots how to handle this task. They used special points marked on the T-shirts to guide the robot actions and trained computer programs to do the folding. Their programs, trained only on this fake data, worked well when tested on real T-shirts the robots had never seen before.
Open 2609.12433v1

OphBiWSSD improves eye surgery action detection with faster modeling

OphBiWSSD: Scaling Temporal Action Localization in Ophthalmic Surgeries with Bidirectional Weight-tied State Space Duality

Commercial implications: For medical device software developers: Enables advanced surgical intelligence products for eye surgery centers by providing fast, accurate action recognition critical for intraoperative guidance.

Fri 11 SeptComputer Vision and Pattern Recognition
The gist
Eye surgeries involve very quick and precise movements that are hard to analyze because current computer models struggle to understand long and detailed sequences without slowing down too much. The authors created OphBiWSSD, a method that looks at surgical videos both forwards and backwards in time, which helps it spot important moments more accurately and faster than previous systems. This new approach uses a clever way to handle information efficiently, making it possible to monitor surgical steps closely without using too much computer memory. Their tests show it works better than earlier methods on a benchmark for eye surgery videos. This research could help build smarter tools that assist surgeons during operations.
Open 2609.12409v1

Hexapod legs evolve walking patterns independently for better coordination

Decentralized Evolution of Hexapod Gaits with Independent Leg Controllers

Commercial implications: For multi-legged robot manufacturers: Enables production of more efficient legged robots by simplifying gait design and improving walking stability.

Fri 11 SeptArtificial IntelligenceRobotics
The gist
Making six-legged robots walk can be complicated because all legs must work together smoothly. This paper shows how evolving each leg's movement pattern separately, without a central controller, lets good walking styles emerge on their own. The authors used simulation and a real robot to test this idea and found that it led to more stable and adaptable walking compared to methods where legs evolve together. This approach also makes it easier to improve robots with many legs by breaking down a complex problem into simpler parts.
Open 2609.12400v1

Odin speeds up encrypted llama 3 inference on nh100 gpu

An Open-Source End-to-End FHE Implementation for Privacy-Preserving Llama 3 8B Inference

Commercial implications: For cloud service operators: This technique enables cloud providers to offer privacy-preserving AI services, attracting clients concerned about data confidentiality.

Fri 11 SeptCryptography and Security
The gist
Cloud services that run large language models usually require sending your text so the provider can see it, which risks your privacy. The authors developed Odin, a system that lets these models run on encrypted input without revealing your data. They improved how data is packed and processed inside encryption, making inference much faster and using less memory. Odin runs Llama 3 8B on a powerful NVIDIA GPU much quicker than previous methods, with the same privacy protections.
Open 2609.12378v1

ChronicleRec compresses long user history for better recommendations

ChronicleRec: Pre-training Temporally Anchored Tokens for Lifelong User Modeling

Commercial implications: For e-commerce platform engineers: Enables retail platforms to sell better-targeted ads and recommendations by efficiently modeling long user behavior without heavy computation.

Fri 11 SeptInformation Retrieval
The gist
Many apps and websites try to understand what users like by looking at their long history of actions, but processing all that data takes too much time and computer power. The authors propose ChronicleRec, which squashes a user’s long sequence of actions into a shorter, ordered summary that still keeps important signals and the timing of events. They designed the method to look only at past actions before each query, learning from different recent time windows to capture interests over time. Their tests show this method improves how recommendations match user intent while running faster online.
Open 2609.12375v1

Occupation-focused benchmark tests large language models on real work tasks

ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge

Commercial implications: For hr tech developers: Enables creation of reliable hiring assessment products that use occupation-specific question-answering powered by language models.

Fri 11 SeptComputation and LanguageArtificial Intelligence
The gist
It can be hard to test how well AI programs understand real jobs because good questions are expensive and rare. The authors created ORQA, which uses trusted job-related websites to make real work questions for AI to answer. They tested many popular AI models and found some do quite well on healthcare jobs but struggle on others like office support. This method helps see what kinds of professional knowledge these AIs really have and where they still need improvement.
Open 2609.12366v1

AI system reduces false ICU alarms while limiting missed alerts

Certified AI Triage of ICU Alarms

Commercial implications: For medical device developers: Enables safer, marketable ICU alarm systems with provable performance guarantees for false alarm reduction.

Fri 11 SeptMachine Learning
The gist
Too many alarms in intensive care units can be false, making it hard for staff to notice real emergencies quickly. The authors developed a method that either keeps, ignores, or delays alarms to reduce false alerts while ensuring very few real emergencies are missed. They provide a mathematical guarantee that the number of missed real alarms stays below a chosen limit with high confidence. Their method performs well compared to other systems on a standard test, reducing false alarms by nearly 75% while only silencing about 1.5% of genuine ones.
Open 2609.12365v1

Socially assistive robots help anxiety homework with cloud web support

A Deployable Architecture for Robot-Mediated Tasks (DART): Evaluation in Socially Assistive Robot-Guided Cognitive Behavioral Therapy Exercises

Commercial implications: For home healthcare device makers: Enables production of low-cost therapy robots with cloud features for in-home mental health support suitable for commercial sale.

Fri 11 SeptRobotics
The gist
It can be hard and costly for robots to do complex health tasks over time. The authors created DART, a system that connects a simple robot with a web app and cloud servers, so it can show images, accept user answers, and store progress. They tested this setup with students using a robot to help with anxiety exercises and found it reduced stress and was easy to use. The system worked well both in the lab and for several weeks at home, though users wanted better speech, visuals, and syncing.
Open 2609.12349v1

Unified system generates 3D motion for humans and animals

UniMo: Unifying Human and Animal Motion Generation

Commercial implications: For game developers: Enables new motion generation tools for gaming studios to produce varied creature animations efficiently from text commands

Fri 11 SeptComputer Vision and Pattern RecognitionGraphics
The gist
Generating realistic 3D movements for different animals is hard because animals have different body shapes compared to humans, and data about animal movements is scarce. The authors created a new system called UniMo that represents motion using points instead of skeleton types, making it work for many species. They also made a large new dataset with many motion examples and text descriptions for both humans and animals. Their method performs well across several tests, showing it's possible to generate motions for both humans and animals in one system.
Open 2609.12342v1

Disengagement-aware simulators improve evaluation of AI tutors

Simulating Disengaged Students to Evaluate LLM-based Tutors

Commercial implications: For educational software companies: Enables development of marketable AI tutors with validated performance for diverse learner behaviors.

Fri 11 SeptMachine Learning
The gist
Sometimes students using AI tutors lose focus or try to trick the system. The authors created computer models that mimic different types of disengaged student behaviors so AI tutors can be tested more realistically. Their models matched human judgments well and helped reveal how various AI tutors perform with different student behaviors. This approach helps developers understand and improve AI tutoring systems before they are used by real learners.
Open 2609.12331v1

Wearable system decides when to act based on personal signals

Affective Agent: On-Device Personalized Intervention Reasoning for Wearable Systems

Commercial implications: For wearable device engineers: Enables companies to sell smart wearables that act on personal data privately on-device with adaptable, user-specific decisions.

Fri 11 SeptArtificial IntelligenceMachine Learning
The gist
Wearable devices can sense how we feel but figuring out when and how to help or alert us is tricky, especially without relying on the cloud. The authors created Affective Agent, a system that uses a small built-in language model combined with body signals, surroundings, and past behavior to make smart decisions on the device itself. It learns about each person by updating stored memories instead of retraining the whole system. They tested it in simulated indoor environments and found it improves when and how the device suggests interventions.
Open 2609.12322v1

Decision-Flow sampling improves reasoning in language models without retraining

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

Commercial implications: For ai product developers: Enables enhanced AI assistance products by improving reasoning accuracy while avoiding expensive model retraining.

Fri 11 SeptMachine Learning
The gist
Large language models can solve problems by reasoning through multiple steps, but finding the best solution path is tricky without extra training. This paper shows a way to explore many reasoning paths during model use, scoring and choosing the best complete answers instead of deciding step-by-step. The authors introduce Decision-Flow Sampling, which finds better chains of reasoning already present in the model, boosting accuracy without costly retraining. This method works well on multiple tests and models, meaning that smarter searching alone can unlock better reasoning from existing systems.
Open 2609.12317v1

Function name sequences enable real-time blockchain attack detection

Function Name Is All You Need to Detect Blockchain Application Attacks

Commercial implications: For transaction monitoring services: Enables new or improved commercial real-time security monitoring products that scale across many contracts without source access.

Fri 11 SeptCryptography and Security
The gist
Blockchain apps called decentralized applications (dApps) sometimes have bugs that attackers exploit, causing people to lose money. Detecting these attacks is hard because it often needs special rules or access to the app’s code, which isn’t always available. The authors show that just looking at the names of functions called in a transaction can reveal what the transaction does and catch attacks. They built a tool named TxLucent that uses these function names to detect attacks quickly and accurately without needing the app’s code or manual rules.
Open 2609.12315v1

Wrist sensors estimate whole-body movement with hybrid physics ai model

Hybrid Physics-AI Framework of Body Center of Mass Dynamics from Wrist-Worn Sensors

Commercial implications: For fitness technology developers: Enables wearable fitness trackers to provide precise whole-body movement metrics, enhancing product accuracy and attractiveness.

Fri 11 SeptArtificial Intelligence
The gist
Measuring whole-body movement is tricky when only using wrist sensors, which don’t capture the entire body’s motion. The authors developed a simpler physics-based model to estimate body center of mass movement from wrist sensors. They improved this by combining the physics model with neural networks to better handle real data and noisy measurements. This hybrid approach led to more accurate and robust estimates of whole-body motion during walking and standing up. Their work shows that mixing physics knowledge with AI can help wearable sensors better understand body dynamics.
Open 2609.12304v1

Thermodynamical AI method improves evolving code and descriptions

T-GADE: Thermodynamical Generative-AI-Driven Evolution of LLM Artifacts

Commercial implications: For ai service providers: Enables commercial AI platforms to offer improved automated coding and optimization services with thermodynamics-enhanced evolution.

Thu 10 SeptArtificial IntelligenceNeural and Evolutionary Computing
The gist
Combining evolutionary computing with large language models (LLMs) can create better sets of structured outputs like paired descriptions and code. The authors propose a method called T-GADE that uses a physics-inspired selection process to evolve these paired artifacts while keeping diversity in the population. They tested this approach on a bin-packing problem and showed it reduces waste compared to previous methods. Their results confirm that this thermodynamics-based selection helps find better solutions and reuse good results.
Open 2609.12286v1

Browser tool creates printable braille and tactile storybooks on demand

Tact: A Zero-Cost, Browser-Based Pipeline for On-Demand Tactile Braille Storybooks

Commercial implications: For consumer 3d printer manufacturers: This paper enables development of accessible braille book printing solutions marketed to home 3D printer users interested in educational accessibility.

Thu 10 SeptHuman-Computer InteractionComputers and Society
The gist
Braille reading skills among blind children are declining partly because making illustrated braille books requires special tools and trained people. The authors created Tact, a tool that anyone can use in a web browser to turn spoken or typed stories into braille pages with matching raised pictures. This tool works without needing an account, internet connection, or extra costs, and it uses simple 3D printers common at home. The authors describe how they built and tested the system, including how the braille and images are made, how it handles different languages, and the ethical rules for using it.
Open 2609.12272v1

Context improves multi-hop question answering on disease knowledge graphs

Repair Before Reinforce: Context-Augmented Knowledge Graph Reasoning for Multi-Hop Question Answering

Commercial implications: For medical chatbot developers: Enables selling more accurate, context-aware medical chatbots using augmented knowledge graphs for better multi-hop question answering.

Thu 10 SeptComputation and LanguageArtificial Intelligence
The gist
Answering complex questions often requires connecting several facts instead of just one. The authors propose training language models not only on individual fact triples from knowledge graphs but also including related supporting facts for better context. They tested this approach on disease-specific knowledge graphs for Gastroparesis and Diabetes and used a repair process to fix mistakes on simpler facts. Adding context helped the models answer harder multi-step questions more accurately, especially after reinforcement learning.
Open 2609.12230v1

Automated 3d camera measures thickness of bioprinted tissue constructs

An Automated Thickness Evaluation Procedure Using an Integrated Structured Light 3D Camera in a Robotic Bioprinting Framework

Commercial implications: For 3d printing service providers: Enables sale of automated thickness inspection products for bioprinting companies seeking quality assurance solutions.

Thu 10 SeptRobotics
The gist
Measuring how thick bioprinted tissues are is important for growing cells properly, but current methods lack precision and automation. This paper presents a fully automated way to measure the thickness of tissue shapes made by bioprinters using a 3D camera that scans the surface and color images to detect the tissue boundaries. The authors combine image segmentation and robot data with 3D point clouds to get very accurate thickness measurements even for complex shapes. They tested their method in virtual simulations and on real bioprinted samples, achieving errors below 0.06 millimeters.
Open 2609.12206v1

Physics aware AI improves identification of 2D quantum materials

QuPAINT: Physics-Aware Multimodal Reasoning for Quantum Material Characterization

Commercial implications: For quality control teams: Enables selling advanced inspection tools that precisely detect material layers for electronics manufacturers ensuring product reliability.

Thu 10 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceMachine Learning
The gist
Identifying ultra-thin 2D quantum materials using microscope images is hard because these materials look different depending on the lab and setup used. The authors created a computer system called QuPAINT that uses physics knowledge and synthetic training images to better find and understand these tiny materials in real microscope pictures. They also made a large test dataset to check the system's accuracy. QuPAINT works better than earlier methods, especially at detecting single-layer flakes, and stays reliable even when shown new types of materials.
Open 2609.12202v1

Real-time music source separation runs efficiently on low-power audio DSP

Real-Time Music Source Separation on a Low-Power Audio DSP

Commercial implications: For audio device manufacturers: Enables new features in portable music and audio equipment that require live music separation with constrained hardware.

Thu 10 SeptSound
The gist
Separating different sounds from a music mix in real time usually needs powerful computers, but the authors show how to do it on a small, low-power audio processor. They found existing methods don’t fit the strict memory and speed limits of typical audio hardware. By changing how the model is trained and adding a new filter that controls delay, they made a system that works quickly and nearly as well as bigger setups. This opens the door for better music processing in small devices.
Open 2609.12201v1

Machine learning predicts pedestrian counts from city map features

Estimating Pedestrian Volumes from GIS-Derived Built-Environment Features: A Machine Learning Framework

Commercial implications: For smart city software developers: Enables new predictive traffic management products that can be sold to municipalities for improved pedestrian monitoring and planning.

Thu 10 SeptMachine Learning
The gist
Transportation planners need to know how many people walk at different intersections to make roads safer, but counting pedestrians manually is hard and expensive. The authors created a computer model that uses information from maps about streets and buildings to predict how many people walk at certain times in Portland. Their model does better than the standard method, cutting down prediction errors by around 12 to 19 percent. This helps make better decisions about where to focus on pedestrian safety without needing lots of manual counting.
Open 2609.12173v1

Neural models improve distant speaker diarization in noisy conditions

Neural Multichannel Distant Speaker Diarization with Heavy-tailed Source Separation Model

Commercial implications: For voice assistant developers: Better speaker separation enables commercial voice assistants to handle overlapping speech and distant voices more reliably.

Thu 10 SeptSoundArtificial Intelligence
The gist
Separating and identifying who is speaking in a room with multiple people talking at once and from a distance is hard. The authors improved a method that uses multiple microphones and advanced math models to better separate voices and label who spoke when. They replaced the usual assumptions about voice signal patterns with more flexible ones that better fit real speech. Their experiments show this change reduces mistakes in identifying speakers.
Open 2609.12154v1

Guide suggests user preferences by smart questions during talks

GUIDE: Generative Utility Inference and Decision Engine

Commercial implications: For financial advisors: Enables advisory platforms to offer tailored portfolio recommendations using guided preference inference, improving client satisfaction and retention.

Thu 10 SeptMachine Learning
The gist
People have many preferences that are hard for computers to understand, especially when these preferences involve many different factors. The authors created GUIDE, a computer system that chats with users and asks smart questions to quickly learn what they like. It uses both clever sampling methods to pick questions and rules about the world to better guess preferences. When they tested GUIDE on choosing investment portfolios, it made better recommendations early on compared to other methods. This helps computers make choices more closely matched to what users really want.
Open 2609.12137v1

Robot learns to insert bending rods precisely with world model guidance

RodForesight: A World Model Enhanced Diffusion Policy for Slender and Material Agnostic Rod Insertion

Commercial implications: For industrial automation teams: Enables manufacturing robots to handle delicate rod insertions more reliably, reducing defects and downtime.

Thu 10 SeptRobotics
The gist
Inserting thin, bendy rods into tight spaces is hard because the rod's tip can move unpredictably when it bends. This paper presents RodForesight, a method that first moves the rod close to the hole using cameras and then carefully uses a special AI approach to predict and correct rod alignment. This method tries out multiple possible moves before picking the best one, leading to more successful insertions than earlier techniques. The authors show that this approach improves success rates on the task by combining vision, prediction, and decision-making.
Open 2609.12103v1

Foundation models improve glucose forecasting when adapted with diet data

Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting

Commercial implications: For diabetes device manufacturers: Enables better predictive algorithms for commercial CGM devices that help patients manage blood sugar based on meal information.

Thu 10 SeptMachine Learning
The gist
Continuous glucose monitors track blood sugar levels frequently, which helps predict short-term changes important for managing diabetes. The authors tested powerful general time-based models on glucose data but found they only worked well after some fine-tuning. Adding information about what people ate also improved predictions, especially after meals. Their study shows these models need specific adjustments for glucose data and that diet details add valuable clues for better forecasts.
Open 2609.11872v1

Speech language models improve reasoning accuracy with real time self correction

RetroThinker: Enabling Retrospective Thinking in Speech LLMs

Commercial implications: For voice assistant developers: RetroThinker allows production voice assistants to provide more accurate real-time spoken responses with self-corrected reasoning.

Thu 10 SeptArtificial IntelligenceComputation and Language
The gist
Speech language models can understand spoken words faster and keep voice details better than converting speech to text first. But they are not as good as text-based models at solving tricky problems quickly. The authors created RetroThinker, which lets a speech model check and fix its own thinking steps while listening. This approach helps the model be more accurate without slowing it down much. Tests showed RetroThinker improved problem-solving scores by 11% with similar speed.
Open 2609.11864v1

Model-aware schedules improve image generation quality and efficiency

Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport

Commercial implications: For graphics software developers: This paper enables better AI-based image generation features that can be sold as improved graphics software functions.

Thu 10 SeptMachine LearningArtificial Intelligence
The gist
Making computer programs that create images often involves mixing noise and data in careful steps. The authors found a new way to plan these steps by considering how well the program predicts the data at each stage. This approach finds better schedules that improve the image quality and require fewer steps. They tested the idea on popular image generation methods and datasets, seeing consistent improvements without extra training work.
Open 2609.11842v1

MotionQ improves wifi gesture recognition under changing conditions

MotionQ: Operator-Conditioned Motion Quotients for Cross-Observation WiFi Gesture Recognition

Commercial implications: For smart home device teams: Allows selling gesture-controlled smart home products that work reliably across different home environments.

Thu 10 SeptHuman-Computer Interaction
The gist
WiFi gesture recognition systems often struggle when the environment changes, like when people move or devices are rearranged. The authors show that these changes change how motion is detected, not just how it looks, which previous methods overlooked. They designed MotionQ, which adapts to these changes by focusing on what motion cues are truly observable from different viewpoints. This makes gesture recognition more reliable even when the setup varies.
Open 2609.11818v1

Whisper tools improve speech transcripts from videos in seven languages

Whisper-Based Speech Transcription from Videos Across Multiple Languages for Cross-Cultural Understanding

Commercial implications: For media localization teams: Enables enhanced subtitle services for worldwide video content by reducing transcription errors using fine-tuned automated tools.

Thu 10 SeptComputation and Language
The gist
Understanding people from different cultures often needs tools that can turn spoken words in videos into text in many languages. The authors looked at how well a speech recognition system called Whisper works on videos in seven languages and found it makes quite a few mistakes at first. But by using a little bit of extra training data, the system gets better and creates clearer transcripts. They also share the data and recordings they used so others can try to make the system even better.
Open 2609.11772v1

Reflex informed learning improves muscle driven human walking control

Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion

Commercial implications: For rehabilitation device developers: This paper enables adaptive neuromuscular control for wearable aids that improve gait quality by modulating reflex gains based on real-time muscle and state feedback.

Thu 10 SeptRoboticsGraphicsMachine Learning
The gist
Controlling human-like walking using muscle models is hard because it requires movements to be both realistic and adaptable to changes like muscle weakness or bumps. The authors combine a basic reflex system with reinforcement learning to adjust key muscle-related controls dynamically. This approach helps produce walking patterns that look more natural, stay balanced between legs, and handle disturbances without retraining. The resulting control method mimics how humans might adjust their muscle reflexes when walking under different conditions.
Open 2609.11733v1

Continuous-time modeling improves speech emotion tracking in text to speech

Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations

Commercial implications: For text-to-speech developers: Enables more natural and style-expressive synthetic voices for virtual assistants and audiobooks, improving user engagement.

Thu 10 SeptSoundArtificial Intelligence
The gist
When computers turn text into spoken words, they need to decide how long to say each part, like syllables or sounds. Usually, these lengths only change the timing but not the details of how the speech sounds. The authors used a new math tool called neural controlled differential equations to model the speech sounds continuously over time. This lets the speech better capture emotions and timing simultaneously, making the computer voices sound more expressive and natural. They showed that this method can track emotional intensity more accurately without losing quality.
Open 2609.11725v1

High-fidelity 3D human scans enable better avatar images

Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need

Commercial implications: For game developers: Enables production-quality avatar creation tools that improve realism for game characters, reducing manual modeling costs.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
To create realistic digital avatars of clothed people, 3D body scans need to be precisely aligned to a standard body model. The authors show previous methods failed to do this well, limiting the quality of avatar images. They introduce AvaImg, a new technique that sharply improves this precise alignment and adds fine surface details. Their method produces avatar textures so close to real scans that they can be efficiently processed by existing 2D image models.
Open 2609.11722v1

Motion-consistent model improves detection and trajectory forecasting

MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View Images

Commercial implications: For autonomous vehicle engineers: Enables more accurate and socially-aware trajectory forecasting, critical for commercial autonomous driving products to navigate complex environments safely.

Thu 10 SeptComputer Vision and Pattern RecognitionRobotics
The gist
Predicting where cars and people will move next is important for self-driving cars. The authors worked with a prior model named DeTra that combined seeing objects and guessing their future paths but was hard to access publicly. They rebuilt DeTra and added ways to teach the model about how objects really move and how they interact with others nearby, without slowing down predictions when in use. Their improved version, MC-DeTra, better guesses the future movement of dynamic road users while keeping or improving how well it spots them in the first place.
Open 2609.11717v1

Random forest classifies real and imagined motor EEG signals accurately

Electroencephalography Signal Analysis for Human Activities Classification: A Solution Based on Machine Learning and Motor Imagery

Commercial implications: For assistive technology developers: This enables affordable thought-controlled assistive tools for people with physical disabilities, improving independence.

Thu 10 SeptNetworking and Internet Architecture
The gist
Understanding brain signals when people move or imagine moving can help develop tools controlled by thoughts. The authors used a machine learning method called Random Forest to identify whether brain signals come from actual movements or imagined ones and also which body part is involved. They tested their method on two types of EEG devices, including a consumer-level one, and it worked well. However, brain activity patterns vary between people, which makes classification harder.
Open 2609.11695v1

AI model management moves beyond storage with learnware concept

Learnware and AI Model Management System

Commercial implications: For ai development companies: This enables selling and licensing AI models through a specification-based system that maintains data privacy while allowing discoverability.

Thu 10 SeptMachine Learning
The gist
Managing AI models today is like storing files without organizing them to work well together. The authors propose upgrading from just saving AI models to managing them as 'learnware,' which combines a model with a detailed description called a specification. This helps systems identify, reuse, and link models from different developers without needing access to the original training data. Their Learnware Dock System is a step toward smarter AI collaboration and reuse.
Open 2609.11656v1

Cardiac phase detection improves with simple model for heart cycle timing

Self-Supervised Cardiac Phase Detection via Single-Parameter Latent Orbits

Commercial implications: For cardiology device engineers: This enables commercial ultrasound device makers to offer automated heart phase detection features improving diagnostic speed and accuracy.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Identifying key moments in heartbeats from ultrasound images is important but often varies between doctors. The authors made a new method that focuses on representing the heartbeat as a simple repeating signal with just one main variable. This approach helps the model learn clear patterns for the heart’s timing without needing labels, making it easier to spot important phases in the heartbeat. Their model matched or beat previous methods while being simpler and faster to train.
Open 2609.11650v1

ZipCodec compresses speech using ultra-low frame rate and bitrate

ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

Commercial implications: For assistive technology builders: Enables product features for assistive communication devices that require efficient, low-latency speech coding operable on standard hardware.

Thu 10 SeptSoundArtificial IntelligenceMachine Learning
The gist
Compressing speech for streaming usually requires sending many small pieces quickly, which is hard to do well at very low frame rates. The authors present ZipCodec, a new technology that sends fewer pieces of speech data per second while keeping sound quality good enough to understand. They built ZipCodec using advanced methods like a special transformer design and clever data compression techniques. ZipCodec works fast enough to run in real time on regular computers and is better than other tools at similar data rates.
Open 2609.11642v1

Vidu S2 enables real-time editable interactive video with avatars

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Commercial implications: For live streaming teams: Enables sale of advanced live streaming software that offers real-time avatar and scene customization for content creators.

Thu 10 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Creating and changing videos in real time is hard. The authors show Vidu S2, a tool that can generate live video of digital characters at 720p quality and lets users change styles, clothes, characters, or backgrounds while the video plays. It improves on an earlier version by allowing fast updates and more complex user commands like dancing. This system lets users interact instantly with video content, opening up new possibilities for live editing and avatar use.
Open 2609.11638v1

Modular manufacturing systems optimized faster with inverse models

Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models

Commercial implications: For factory automation vendors: Enables building and selling advanced reinforcement learning control software tailored for modular flexible manufacturing.

Thu 10 SeptArtificial IntelligenceMachine Learning
The gist
Modular factories can be hard to control because they have many separate parts working together in different ways. The authors introduced a new method that helps robots and machines learn better how to run these modular systems by teaching the control software how actions translate into changes. They use a special kind of learning called reinforcement learning but improve it by separating how machine actions link to results. This makes training faster and the systems perform better in practice.
Open 2609.11615v1

Gait recognition improved with multimodal data and unified identity encoding

MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous Modalities

Commercial implications: For security system developers: Enables new multimodal biometric security products that work across lighting and visibility conditions by combining diverse sensor data into one system.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Recognizing people by the way they walk, called gait recognition, usually uses video images or simplified outlines. The authors point out that walking generates many different types of data like infrared, depth, radar, and more, which are rarely studied together. They created a big dataset called MMGait that combines all these types of data to compare and test gait recognition across different sensors. They also developed a system called OmniGait++ that can learn from and combine multiple sensor types at once, working well whether it uses one type of data or many. Their work shows it’s possible to have one system that adapts to different sensor combinations and still identify people accurately.
Open 2609.11601v1

TimelyRAG improves question answering on changing documents with overlapping updates

TimelyRAG: Semantic-Temporal Hybrid Retrieval for Time-Critical Question Answering in Overlapping-Evolving Documents

Commercial implications: For legal technology teams: Enables legal tech companies to build tools offering more reliable, up-to-date regulation answers, improving compliance services.

Thu 10 SeptInformation Retrieval
The gist
Information like laws and policies often change over time by updating existing documents rather than replacing them entirely. This makes it hard for computer systems to find the right and most recent version to answer questions accurately. The authors developed TimelyRAG, a method that considers both the meaning and timing of updates to better match questions with the correct document version. They also created a benchmark called TimelyQABench to test these challenges in regulation-heavy texts, showing that accounting for document timing greatly improves question answering accuracy.
Open 2609.11572v1

Pre-trained convolutional networks show varying accuracy in melanoma detection

A Comparative Evaluation of Pre-trained Convolutional Neural Networks for Melanoma Detection

Commercial implications: For medical ai developers: Enables creation of targeted AI-powered diagnostic tools that can be commercialized for healthcare providers.

Thu 10 SeptComputer Vision and Pattern RecognitionArtificial Intelligence
The gist
Melanoma, a dangerous type of skin cancer, is hard to identify early because skin lesions can look very similar. The authors compared different pre-trained convolutional neural networks (CNNs), which are computer models trained to recognize images, to see how well they detect melanoma using two types of skin images. They found that the best CNN models achieved accuracy between 71% and 84%, but performance varied depending on the type of image used. This study helps identify which CNN models might work best for different skin image types to support doctors in diagnosing melanoma.
Open 2609.11550v1

Post-training enables fine-grained natural language control of speech emotion and timing

Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

Commercial implications: For audiobook producers: Enables production companies to offer more expressive audiobook voices controlled by natural language prompts.

Thu 10 SeptSound
The gist
Many text-to-speech systems can talk but have trouble making their voice change emotions or speaking speed in small parts of a sentence. This paper shows how to improve existing speech models by teaching them after they are built, so they can understand simple instructions about feeling and pace for different segments of speech. The researchers fine-tune these models using new techniques to make the changes accurate while keeping the person’s voice and words clear. This method does not need extra parts during speaking, making it easier to use.
Open 2609.11523v1

Model generates images and layouts together for better design templates

Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates

Commercial implications: For advertising agencies: Enables selling customized, harmonious ad template generators that speed up creative production for marketers.

Thu 10 SeptComputer Vision and Pattern RecognitionArtificial IntelligenceGraphics
The gist
Creating graphic design templates usually means making a background image and placing items on top one after another. The authors found this sequential method misses how backgrounds and layouts affect each other. They built a new model that creates the background and layout at the same time, letting them interact during the process. This approach keeps the realistic look of designs and lets users guide the results without retraining the model.
Open 2609.11519v1

Prototype-based method improves visible-infrared person matching

Prototype Matters: Modality-unified Prototype Self-distillation for Unsupervised Visible-infrared Person Re-identification

Commercial implications: For security system developers: Enables building improved multispectrum identification software for security firms by enhancing accuracy without costly labeled data.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Matching people captured in visible light with those in infrared is hard because the images look very different. The authors show that using shared group examples called prototypes helps computers learn better links between the two kinds of images. They also refine this process by letting the system teach itself, leading to better understanding inside and across both image types. Their combined method improves matching accuracy without needing labeled training data.
Open 2609.11514v1

Multi subject video generation gains precise control and better identity consistency

Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation

Commercial implications: For visual effects studios: Enables studios to produce higher-quality controlled character animations for films or games, improving workflows with less manual editing.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Generating videos with multiple people or subjects is hard because it’s difficult to control how closely the video matches the input references and to avoid mixing up who is who. The authors studied how a type of AI model focuses on different parts of the video and found it naturally highlights each subject’s location. Using this, they created a method that guides the model during generation to keep subjects clear and consistent. Their method also uses rewards during training to prevent the model from drifting away from the subjects. This leads to videos that better maintain each subject’s identity and allow users to control the quality without needing to retrain the model.
Open 2609.11507v1

Dataset links reviewer comments to paper evidence for better grounding

ReGround: Grounding Reviewer Comments in Multimodal Evidence

Commercial implications: For academic paper management platforms: Enables product features that streamline peer review management and improve author-reviewer communication in submission platforms.

Thu 10 SeptComputation and LanguageInformation Retrieval
The gist
Sometimes reviewers of scientific papers give comments that point to specific parts of the paper, but it's very hard to find exactly where in the paper those comments relate. The authors created a large collection of examples where reviewer comments are linked to the evidence in the paper, using author replies to help find the exact parts. They tested many ways to find these links and found it is a hard problem, especially because papers often have text, figures, and tables that all give clues. This new dataset helps computers learn to better understand and connect reviewer comments to the right parts of scientific papers.
Open 2609.11460v1

Avatar motion generated naturally from speech text and path inputs

Multi-Modal Controlled Coherent Motion Generation

Commercial implications: For game developers: Enables game studios to produce more immersive characters by synchronizing motion with speech and text inputs in real time.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
People often walk and talk at the same time naturally, but making 3D avatars do this realistically is hard. The authors developed a new method that can mix different types of input—like speech sounds, text descriptions, and movement paths—to create smooth, lifelike avatar motions. Their approach builds the motion step-by-step, letting each input contribute separately before combining them in a way that looks natural. This method works well even when the inputs aren't perfectly matched, producing more realistic movements than earlier techniques.
Open 2609.11439v1

Fast real-time light rendering using gaussian mixtures

Gaussian Light Transport

Commercial implications: For game developers: Allows game studios to deliver high-quality lighting with low memory use and fast rendering suited for consumer hardware.

Thu 10 SeptGraphics
The gist
Rendering realistic lighting in computer graphics is usually slow because it models complex interactions between light and surfaces. This paper presents a new method that represents light behavior as a mix of Gaussian functions considering position, direction, surface features, and materials. This representation lets the authors estimate lighting more quickly and efficiently by directly solving the rendering equation. As a result, their method can produce fast, view-independent lighting effects using less memory than neural network approaches, enabling real-time rendering.
Open 2609.11430v1

Quantum strategies solve blocked path puzzle in interferometers

The Quantum Plumber's Problem

Commercial implications: For quantum communication device makers: Enables devices that detect communication channel failures more accurately and securely, improving product quality.

Thu 10 SeptInformation Theory
The gist
The paper looks at a tricky problem where one path in a quantum setup is blocked, and someone called the "quantum plumber" wants to find out which path it is with certainty. The authors explore different ways to figure this out using special setups called interferometers, including a more complex one with three paths. They test their ideas by simulating many attempts and also study a version where the blockage is replaced by a detector that doesn't destroy the particle. This work helps show how quantum methods can be designed to solve these detection problems more effectively.
Open 2609.11416v1

Brain pace estimates reveal early brain aging linked to impairment

Brain-PACE: A Deep Siamese MRI Framework for Modelling Longitudinal Brain Acceleration

Commercial implications: For medical device developers: Enables creation of enhanced diagnostic software for neurodegenerative diseases to sell to hospitals and clinics.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Brain health can be estimated by predicting brain age from MRI scans. The authors developed Brain-PACE, a method that looks at how fast the brain ages over time by comparing pairs of MRI scans. They found that a faster brain aging pace was linked to worse thinking and daily functioning in people with mild cognitive problems, and to higher levels of a brain protein involved in Alzheimer's disease. Their method improved accuracy over previous approaches by learning from brain images in a smarter way.
Open 2609.11378v1

Vision transformer improves sewer defect classification with lightweight models

Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification

Commercial implications: For embedded system developers: Enables compact, efficient AI products for automated sewer inspection that can be sold to utilities or inspection service providers.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Sewer systems sometimes have different types of damage that need to be found to keep them working well. Existing computer programs find it hard to quickly and accurately recognize multiple types of sewer problems from images. The authors created a new way to analyze sewer photos using a vision Transformer that combines details at different levels, making it more accurate. They also built smaller models that work well with limited computing power, suitable for on-site inspections. Their tests showed these models can reliably detect sewer defects and work better than earlier methods, even when there is less training data.
Open 2609.11375v1

Ramamba net improves detecting focus of hearing in noisy places

RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection

Commercial implications: For hearing device developers: Enables advanced hearing aids to decode user auditory attention more accurately, improving user experience in noisy environments.

Thu 10 SeptArtificial Intelligence
The gist
People trying to hear one speaker in noisy places use brain signals to detect who they are focusing on. EEG brain signals alone miss some important clues, so this paper looks at combining EEG with eye movement signals (EOG) for better detection. The authors created RAMamba-Net, which mixes these signals in a smart way that pays attention to how reliable each signal is at each moment. Their system improves accuracy and is more robust when signals are noisy or change. This might help devices like hearing aids understand what users want to hear more reliably.
Open 2609.11372v1

Cross-modal augmentation improves brain state decoding accuracy

Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding

Commercial implications: For neurotechnology product teams: This enables the development of advanced brain-monitoring devices selling better decoding accuracy to healthcare and consumer neurotech markets.

Thu 10 SeptArtificial Intelligence
The gist
Understanding brain states by combining different types of brain signals is tricky because data is limited. The authors developed a method called CoMA-DiT that uses each type of signal to help generate more training data for the other, improving learning. This approach led to better predictions of brain states like attention and emotions than existing methods. It also helps the model understand how different brain signals relate and interact.
Open 2609.11341v1

OCT tracking improves motion accuracy using predictive landmark updates

Predictive Multi-Landmark OCT Tracking for Increased Motion Robustness

Commercial implications: For medical device developers: Enables development of medical navigation devices that track instruments precisely at high speeds, enhancing surgery safety and outcomes.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Tracking devices inside the body during movement is tricky because fast motion can confuse the tracking system. The authors found a way to predict where multiple landmarks will be during quick motion, making the system more reliable. Their approach keeps errors below 1 millimeter even when objects move fast and many landmarks are tracked one after another. This could help technologies that rely on precise tracking in 3D space work better under challenging conditions.
Open 2609.11330v1

Mi-ripple reduces digital ripple artifacts in ai edited images

Mi-Ripple: Restoring Images Degraded by Iterative AI Editing

Commercial implications: For photo editing software developers: It enables photo software companies to offer enhanced AI editing tools that produce cleaner images, differentiating their products in the market.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Images edited repeatedly by AI can develop distracting grid-like or grainy textures called digital ripple. The authors present Mi-Ripple, a method that carefully identifies and reduces these unwanted patterns while keeping the important parts of the image intact. They do this by separating the different types of artifact textures and using special filters and cleaning steps that remove the distortions without losing detail. This approach results in visibly cleaner images and measurable reductions in artifact levels.
Open 2609.11317v1

Diffusion vision-language models improve answers by adapting reasoning length

Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

Commercial implications: For chatbot developers: Enables building more efficient and accurate conversational AI products that adjust reasoning per question complexity, improving user experience.

Thu 10 SeptArtificial Intelligence
The gist
Some AI models answer questions by slowly refining their guesses step by step. But using the same number of steps for all questions can cause problems: simple questions may get overthought, while complex questions may be cut short. The authors studied this problem with a model called LLaDA-V and created a way to watch how the answer changes during thinking. Their system decides when to stop early, keep going normally, or focus on detailed reasoning, based only on signals seen without knowing the true answer. Tests showed this approach gave better, more reliable answers across different types of questions.
Open 2609.11315v1

Greek automatic lyric transcription improves with adapted Whisper models

Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation

Commercial implications: For music streaming services: Enables lyric display products for Greek music listeners, offering content-enhanced streaming services.

Thu 10 SeptComputation and LanguageSound
The gist
Transcribing song lyrics automatically is harder than regular speech recognition because songs have melody, rhythm changes, and background music. This difficulty is even greater for Greek songs, which haven't had much research done before. The authors tested different versions of a model called Whisper to see how well it can transcribe Greek song lyrics. They found that bigger models work better, and training the models with a mix of related tasks helps smaller models. Their best adapted model cut errors significantly, creating the first benchmark for this task in Greek.
Open 2609.11302v1

Analogue memory hardware powers bio-inspired probabilistic decision making

Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1

Commercial implications: For hardware engineers: This paper enables creation of analogue memory-based chips capable of brain-like uncertain decision-making for AI hardware markets.

Thu 10 SeptArtificial IntelligenceMachine Learning
The gist
Animals make decisions by combining what they sense with what they already believe, even when unsure. The authors explain how brain-like random activity in neurons and connections can model this by exploring many possible states to learn and decide. They found that certain noisy electrical memory chips can mimic this process efficiently in computers. This could help build machines that learn and make decisions in a way similar to the brain, handling uncertainty naturally.
Open 2609.11281v1

Few-shot learning techniques compared for network attack detection

Few-Shot Learning for Network Intrusion Detection: Methods, Datasets, and Performance

Commercial implications: For network security vendors: Enables development of intrusion detection products capable of quickly adapting to emerging threats with minimal data.

Thu 10 SeptCryptography and Security
The gist
Detecting new types of cyberattacks on computer networks is difficult because there are usually very few examples available to learn from. The authors reviewed studies on few-shot learning, a method that helps systems learn from only a few examples, applied to this problem. They found that approaches like meta-learning and convolutional neural networks are most common, and certain datasets are frequently used for evaluation. However, differences in experimental setups and missing details make it hard to fairly compare the methods.
Open 2609.11275v1

Xiaomi model transcribes speakers separately in noisy group settings

Xiaomi-CocktailASR-1 Technical Report

Commercial implications: For call center operations: Enables selling advanced transcription services that isolate individual speakers in multi-speaker recordings for improved customer insights.

Thu 10 SeptSoundComputation and Language
The gist
Hearing and understanding speech when many people talk at once is a tough problem for computers, like trying to focus on one conversation at a noisy party. The authors introduce Xiaomi-CocktailASR-1, a new speech recognition system that can pick out what a single target speaker is saying even when others are talking too, without needing to separate all voices first. It also knows when the target speaker isn't speaking and doesn't produce confusing text. This model is competitive when only one person talks and works well in multi-speaker situations, advancing the ability to understand overlapping speech.
Open 2609.11274v1

Spiking neural network predicts cancer nerve invasion with less energy

SCINTILLA-SNN: A Spiking Multi-Scale Selective Aggregation Network for Perineural Invasion Prediction

Commercial implications: For medical device developers: Enables development of specialized low-power diagnostic imaging devices or software for hospitals and clinics using this paper's spiking network approach.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Predicting whether certain cancer has spread along nerves before surgery is helpful but hard because signs on MRI scans are very small and hard to spot. The authors created a special type of neural network that processes MRI scans more efficiently by focusing on important small areas using a brain-inspired spiking method. Their method not only predicts this nerve invasion better than usual methods but also uses much less computing energy. They tested it on 10 years of patient data and showed good accuracy and energy savings.
Open 2609.11237v1

Spiking neural networks partition inputs with richer patterns than ReLU nets

Polyhedral Geometry of Time-to-First-Spike Neural Networks

Commercial implications: For neuromorphic hardware developers: This work supports building neuromorphic chips that utilize spike timing for improved computational expressivity, enabling novel AI hardware products.

Thu 10 SeptMachine Learning
The gist
This paper explores how spiking neural networks, which use timing of neuron spikes to process information, can represent input-output relationships differently than traditional neural networks. The authors studied a model where each neuron’s firing time is determined by patterns of input spike order, forming complex regions in input space. They mathematically describe these regions as shapes called polyhedra and show that spiking networks can create more varied partitions of inputs than typical ReLU networks. This suggests spiking networks can be more expressive in how they map inputs to outputs.
Open 2609.11227v1

Sampling strategies affect accuracy of web security studies

You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements

Commercial implications: For cybersecurity product developers: Enables new scanning tools that deliver cost-effective, unbiased security assessments by using adaptive sampling methods.

Thu 10 SeptCryptography and Security
The gist
Web security researchers often study large lists of websites to find security problems, but checking every site can be too expensive. Instead, they look at a smaller sample, but it hasn’t been clear if common ways of picking these samples give a true picture. The authors studied different sampling methods and found that picking the top popular sites misses many important details and can bias results. They suggest using a probability-based approach that gives more reliable estimates and works well even when you don’t know how common a security issue is.
Open 2609.11218v1

Ai chatbot detects stress and supports wellness for pakistani students

An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning

Commercial implications: For mental health app developers: Enables development of mental health apps tailored for Pakistani students using AI and multilingual chatbots, a niche not addressed by existing Western-focused tools.

Thu 10 SeptArtificial Intelligence
The gist
Pakistani university students face many unique pressures that affect their mental health. The authors created a chatbot that uses artificial intelligence to spot signs of stress from student responses and then offers wellness advice. This chatbot understands cultural details and communicates in English, Urdu, and Roman Urdu to better help these students. The system uses a machine learning model that achieved nearly 90% accuracy in identifying stress levels. One key stress factor found was the relationship between teachers and students, showing the importance of cultural context in mental health tools.
Open 2609.11199v1

FST Pay ensures safe teen payments with strict real-time checks

FST Pay: Deterministic Safety-Gated Architecture for Youth Digital Payments

Commercial implications: For digital payment providers: Enables payment platforms to offer youth accounts with built-in safety features that comply with financial regulations.

Thu 10 SeptSoftware EngineeringCryptography and Security
The gist
Digital payment systems now let teens access money instantly, which helps them learn about finances but also risks impulsive buys and scams. The authors present FST Pay, a system that ensures teen payments follow strict, clear rules before approval. It blocks or flags risky transactions using six fixed safety checks, involving spending limits, guardian approval, timing, and device security. AI is only used afterward to explain transactions, not to influence payment decisions.
Open 2609.11195v1

Benchmark tests large language models on software version rules

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

Commercial implications: For ai coding assistant makers: This enables AI assistants to reliably handle version management tasks, making the product more trustworthy and useful.

Thu 10 SeptArtificial IntelligenceSoftware Engineering
The gist
Large language models (LLMs) like GPT often need to decide if software version numbers match certain rules, but it was never checked how well they understand these rules. The authors made a new benchmark called SemVerBench with questions from three major coding systems and tested six top LLMs. They found common predictable mistakes related to specific version rules, and showed some models do much better than others. They also showed that giving models a little extra hint fixes many mistakes, suggesting the problem is using knowledge, not lacking it. They recommend that coding tools should let specialized software handle version checking instead of relying on LLMs alone.
Open 2609.11180v1

Terms txt protocol enables web crawler deals with identity and payment

terms.txt: A Consent and Compensation Protocol for Agentic Web Access

Commercial implications: For cloud service operators: Enables development of paid API access and licensing solutions for AI platform customers using extensive web data.

Thu 10 SeptNetworking and Internet ArchitectureArtificial IntelligenceCryptography and Security
The gist
The web traditionally allowed search engines to crawl websites freely while sending users back in return. However, as AI programs now crawl much more and use many pages per user visit, this old deal is breaking down. The authors created terms.txt, a new file format that lets websites specify rules for different web users, including who they are, what they want, and what they should pay. Their system includes security features and can track if these rules are followed, all with minimal delay added to web requests.
Open 2609.11152v1

Automatic evaluation method measures human interpreter quality accurately

Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

Commercial implications: For language pedagogy software makers: Enables development of language learning apps that provide detailed, rubric-aligned scoring for simultaneous interpreting practice.

Thu 10 SeptComputation and Language
The gist
Simultaneous interpreting is when people translate speech live, but it's hard to automatically measure how well they do it. The authors created a dataset with human scores on meaning, delivery, and timing for many short interpreted speech segments. They found that current methods mix up these evaluation parts and don't match human judgments well. They developed a new machine learning model that better keeps these evaluation aspects separate and aligns more closely with human ratings, making it useful for giving interpreters helpful feedback.
Open 2609.11131v1

Multi-view method improves 3d shape generation with guided noise control

ReconPlusGen: Injecting Reconstruction Prior into Multi-view 3D Generation through Noise Inversion and Modulation

Commercial implications: For 3d content creators: Creates higher fidelity 3D models useful for games, VR experiences, and digital asset production.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Creating 3D shapes from multiple pictures is a tricky job because the computer has to guess parts it can't see. The authors present a way to guide this process by predicting a rough 3D shape first and then mixing that knowledge carefully into a generative method that normally produces random results. This helps the computer keep the important parts from the pictures while still being creative to fill in missing details. Their method lets the 3D shapes look more accurate and realistic.
Open 2609.11129v1

ProMediConv sets benchmark for AI legal dispute mediators

ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation

Commercial implications: For legal technology developers: Enables creation of legal AI mediation products that improve dispute resolution efficiency and reduce mediator training costs.

Thu 10 SeptComputation and Language
The gist
Mediating disputes is important but training skilled mediators is hard and slow. The authors created ProMediConv, a new way to test AI systems that handle mediation, using real legal cases with detailed notes on strategies and behaviors. They also designed a better method for measuring how well these AI agents influence the conversation over time. Their study reveals challenges current AI models face in handling complex, multi-person legal talks. This work offers a useful foundation and standard to improve AI tools for resolving conflicts.
Open 2609.11101v1

Muscle-driven simulation creates realistic sprinting without real examples

Learning Realistic Athletic Sprinting Without Demonstrations

Commercial implications: For game developers: Enables realistic athletic animations for games without expensive motion capture, improving visual quality and reducing development cost.

Thu 10 SeptGraphics
The gist
High-speed running motions are hard to create realistically in simulations without using actual video or motion capture data. The authors built a fast computer system that uses detailed muscle models and learns to run quickly just by trying different motions and getting rewards for speed and safety. This means their system can generate realistic sprinting and athletic movements like side-stepping or backpedaling without needing to watch real athletes. The simulated running looks a lot like real runners according to data comparisons.
Open 2609.11083v1

Quantum reservoir computing predicts molecular properties accurately

Coherent Floquet quantum reservoirs for molecular property prediction

Commercial implications: For drug developers: Enables more accurate and efficient molecular screening tools for pharmaceutical companies developing new drugs.

Thu 10 SeptMachine Learning
The gist
Predicting how molecules behave is important for designing new materials and medicines, but it can be complex. The authors use a quantum system that evolves in a special repeating way to process information about molecules. This system turns molecular data into fixed-size feature patterns that a classical computer can interpret to predict things like how molecules interact with inhibitors or cross the blood-brain barrier. Their approach performs better than some classical methods and keeps useful information even when there is noise in real quantum experiments.
Open 2609.11071v1

Acoustic and prosodic cues improve speech turn end detection accuracy

Less can be More: What Aspects of Speech Drive End-of-Turn Detection

Commercial implications: For voice assistant developers: Enables creation of smoother AI assistants that better detect when users finish speaking, enhancing user experience and market competitiveness.

Thu 10 SeptArtificial IntelligenceSound
The gist
Knowing when someone finishes talking is important for smooth conversations with AI. The paper shows that listening to sound patterns and intonation helps computers better guess when a person stops speaking. Surprisingly, understanding the words doesn’t improve this guess and may actually cause errors. The authors found that focusing on how something is said, rather than what is said, leads to faster and more reliable detection of turns in conversation.
Open 2609.11066v1

EMMI reduces edge communication for multimodal AI by 32 times

EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression

Commercial implications: For mobile device developers: Enables new or improved cloud-assisted multimodal AI apps on mobile devices with constrained bandwidth and power.

Thu 10 SeptMachine LearningDistributed, Parallel, and Cluster Computing
The gist
Running large AI models that understand images, text, and sensor data on small devices like smartphones or cameras is very hard because these models need a lot of computer power and memory. The paper presents EMMI, a way to shrink and combine data from different sensors right on the device, so only a small summary is sent to a bigger server for understanding. This method cuts the data sent over the network by 32 times and still keeps the AI’s accuracy similar, making it much faster to get answers under low network conditions. It also keeps private data safer by not sending raw sensor information.
Open 2609.11058v1

Dynamic privacy protection boosts usefulness of large language models

Demystifying the Privacy-Utility Trade-off in LLM Interactions

Commercial implications: For ai product developers: Enables building privacy-aware AI products that protect user data without sacrificing performance, appealing to privacy-conscious markets.

Thu 10 SeptArtificial IntelligenceCryptography and Security
The gist
Large language models help with many tasks but need a lot of personal information to work well, which risks privacy. The authors found that privacy methods that treat all data the same harm usefulness a lot. They discovered that what to hide, how to hide it, and how information fits together depends on the user's goal and the task. Using this, they built a system that smartly protects privacy while keeping the model's helpfulness much higher than before.
Open 2609.10992v1

Enhanced chest X-ray detection with fractal pixel transformation technique

Exponential Pixelating Integral transform with dual fractal features for enhanced chest X-ray abnormality detection

Commercial implications: For medical device manufacturers: Enables development of advanced X-ray analysis software that improves accuracy and speed of lung disease diagnosis for hospitals and clinics.

Thu 10 SeptComputer Vision and Pattern Recognition
The gist
Detecting lung diseases from chest X-rays is hard because images can be unclear and noisy. The authors created a new method that brightens the important parts of the image and uses special fractal math shapes to highlight lung features. Then, a smart computer system sorts the images to identify different lung illnesses accurately. Their tests showed this method works better than previous approaches in spotting problems in chest X-rays.
Open 2609.10988v1

Contact guided retargeting preserves human object interactions on diverse characters

ReCHOIR: Contact-guided Human Object Interaction Retargeting to Diverse Characters

Commercial implications: For animation studios: Enables scalable production of diverse character animations with consistent object use, improving animation pipelines.

Thu 10 SeptGraphics
The gist
When people move and interact with objects, their motions and how they touch those objects matter. The authors created a system called ReCHOIR that can take a recorded human movement with object use and adapt it so different virtual characters perform it naturally. Their method keeps the important body movements and the contact with the object consistent, even if the new character's body is quite different. This helps make animations where characters of different shapes can interact with objects realistically.
Open 2609.10982v1