Papers for

cloud platform teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Efficient attention method speeds up large context processing

SANTA++: Sampling Attention through Representative Keys

Abstract: Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusion probability. This importance sampling correction estimates attention over the full cache, with a sampling budget that lets us trade memory reads for accuracy. Remarkably, with 32 or 64 sampled teams, SANTA++ uses 16% to 22% of dense attention's KV reads and retains 94% to 99% of the dense-attention baseline's scores on LongBench v2 and HELMET's retrieval-augmented generation subset, and 85% to 91% on RULER, with Qwen2.5-7B-Instruct at 32K context. With 31 sampled teams, our GPU implementation delivers a $1.69\times$ attention speedup over the dense FlashAttention baseline at 32K context. By reducing the number of cache entries read, SANTA++ in principle complements architectures with compressed KV representations, such as multi-head latent attention. Our kernels are available at: https://github.com/OPUSLab/santapp-kernel-demo.git.

Mon 28 SeptMachine LearningComputation and Language
The gist
Processing long sequences of information in AI models can be slow because they have to look at everything at once. The authors introduce SANTA++, which cleverly picks smaller representative parts to pay attention to without having to check everything. This approach keeps results very close to full processing but reads much less data, making it faster and more efficient. Their tests show it works well on tasks requiring understanding of very long context.
Open → 2609.35629v1

Relational AI personas boost user engagement and goal diversity

AI Persona, Service Consumption, and User Intent Entropy: Field Experimental Evidence from an LLM Platform

Abstract: Problem definition: Firms deploying large language model services must decide how their AI communicates, not just what it can do. We examine how a relational persona - warmer, more empathetic and more engaging than a non-relational persona - affects service consumption and the evolution of user objectives. Methodology/results: In a randomized field experiment with 9,586 newly registered users, we hold the underlying model and service capabilities constant. The relational persona increases interactions (sessions, +8.1%; duration, +10.6%; chat rounds, +24.2%; intent entropy, +5.8%) and outputs (files, +12.3%; distinct goals, +12.1%). Effects vary by entry intent. First-session effects are insignificant for Task Execution users. Socialization and Knowledge Exploration users show similar increases in chat rounds: Socialization increases intent entropy without more outputs, whereas Knowledge Exploration increases outputs without higher intent entropy. Modeling intent dynamics as a transition process, we find higher intent transition entropy for Socialization (+11.8%) but higher intent continuation probability for Knowledge Exploration (+12.8%), suggesting greater conversational breadth and persistence, respectively. In subsequent use, the relational persona increases aggregate chat rounds and outputs across all entry intents. Session count rises by 12.2% for Task Execution and 49.4% for Socialization, but not significantly for Knowledge Exploration. Effects on session count and intent entropy strengthen over time, whereas output effects remain stable. Managerial implications: AI persona is an operational design lever, not merely a presentation feature. Because more interactions do not uniformly generate more outputs, firms should evaluate interactions and outputs separately and consider matching persona to user intent, especially when added interactions consume costly computing resources.

Sun 20 SeptHuman-Computer Interaction
The gist
Choosing how an AI talks to users can change how people use it, not just what it does. The authors found that when an AI sounds warmer and more empathetic, users interact more, explore more topics, and produce more work. Different kinds of users respond differently: some chat more broadly, others stick with their goals longer. This means companies should think carefully about the AI’s personality to match what users want.
Open → 2609.23274v1

Distributed database testing tool finds new query processing bugs

Distribution-Aware Distributed Database Testing (Extended Version)

Abstract: Distributed database management systems (DDBMSs) introduce new challenges for assessing their reliability due to distribution-specific characteristics that affect query execution and optimization. Existing testing approaches, largely designed for centralized DBMSs, often fail to explore diverse distributed execution behaviors and suffer from low executability of generated test queries, thereby limiting their effectiveness in bug detection. We propose DAT (Distribution-Aware Testing), a novel automated approach for detecting query-processing bugs related to distribution strategies and distributed optimizations in DDBMSs, by systematically leveraging distribution-aware information throughout the testing pipeline. DAT builds on a set of techniques that capture diverse combinations of logical schemas and data distribution strategies, and performs guided query mutation to trigger a wide range of distributed query execution behaviors and optimizations, while improving query executability via historical feedback. We implement our approach in a tool, DistRanger, and evaluate it on four widely used production DDBMSs. It uncovers 31 previously unknown bugs, including 28 related to distributed query processing and optimization, and outperforms state-of-the-art testers.

Wed 16 SeptDatabases
The gist
Distributed databases split data across many machines, making it harder to test if they work correctly. Existing tests often miss problems because they don’t cover how data is spread out and handled. The authors created DAT, a new method that understands how data is distributed and changes queries to explore more cases. Their tool, DistRanger, tested four popular distributed databases and found 31 unknown bugs, mostly involving how queries are processed and optimized across machines.
Open → 2609.18501v1