Papers for
ai tool integrators
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Benchmark reveals varied strengths of language models on quantum tasks
QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks
Abstract: We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find that overall rankings mask substantial per-task variation. The Spearman correlation between overall and per-task rankings is statistically insignificant for 4 out of the 11 tasks included in this benchmark. A 2-parameter Item Response Theory (IRT) model validates measurement quality, and prompt sensitivity analysis confirms ranking robustness across prompt conditions. All tasks are auto-verifiable via execution, thus not requiring any manual evaluation. We make the code and data publicly available on Huggingface.
Benchmark evaluates AI email agents on enterprise productivity tasks
EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks
Abstract: Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and productivity scenarios across 16 task categories. The benchmark couples a typed email API specification with provider-neutral naming, a deterministic synthetic Enron-inspired corpus, and a scenario suite whose topic selection was informed by aggregate task-intent telemetry from an interactive prototype. Its hybrid evaluation protocol combines 258 executable static assertions with 211 LLM rubrics. We evaluate eight LM configurations on a fixed single-user corpus. The best-performing configuration passes only 33.5% of scenarios despite 99.7% of its tool calls completing without an observed API failure, with pass rates varying substantially across task categories. This gap shows that valid tool execution is not equivalent to task completion. EmailBench provides a self-contained environment for end-to-end email-agent evaluation, with broader tool coverage, multi-persona testing, and repeated-run evaluation as future work areas.
Interactive authoring tools improve science data visualization and analysis
"Here Be Sharks!": Enhancing Scientific Communication and Analysis through Authoring Interactivity
Abstract: We report on a case study for designing authoring environments for interactive visualizations to enhance scientific work. We conducted a workshop and prototype review with a group of marine biologists. When it came to visualizing their data, participants identified challenges in conveying their research accurately and completely as well as and analyzing it with ease. Based on our findings, authoring environments should consider the context of scientific work aspects such as domain expertise, collaboration culture, and publication traditions. We propose consideration of non-traditional programming languages and environments for scientific work, discuss ways of facilitating interactive visualizations for scientists, and examine ways of meaningfully integrating AI. The critiques, artifacts, and reactions from the scientists along with our analysis and discussion inform how computational tools should be designed for scientific work.