Papers for

ai tool integrators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Benchmark reveals varied strengths of language models on quantum tasks

QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks

Abstract: We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find that overall rankings mask substantial per-task variation. The Spearman correlation between overall and per-task rankings is statistically insignificant for 4 out of the 11 tasks included in this benchmark. A 2-parameter Item Response Theory (IRT) model validates measurement quality, and prompt sensitivity analysis confirms ranking robustness across prompt conditions. All tasks are auto-verifiable via execution, thus not requiring any manual evaluation. We make the code and data publicly available on Huggingface.

Mon 28 SeptArtificial IntelligenceMachine Learning
The gist
Evaluating how well large language models understand quantum computing, the authors created a test covering many different quantum tasks like building circuits and fixing errors. They found that a model that does well overall might not do well on some specific tasks, showing that these skills vary a lot. The benchmark runs tasks automatically, so no manual checking is needed, and their methods ensure reliable results. They also shared the test and data openly for others to use.
Open → 2609.35581v1

Benchmark evaluates AI email agents on enterprise productivity tasks

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

Abstract: Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and productivity scenarios across 16 task categories. The benchmark couples a typed email API specification with provider-neutral naming, a deterministic synthetic Enron-inspired corpus, and a scenario suite whose topic selection was informed by aggregate task-intent telemetry from an interactive prototype. Its hybrid evaluation protocol combines 258 executable static assertions with 211 LLM rubrics. We evaluate eight LM configurations on a fixed single-user corpus. The best-performing configuration passes only 33.5% of scenarios despite 99.7% of its tool calls completing without an observed API failure, with pass rates varying substantially across task categories. This gap shows that valid tool execution is not equivalent to task completion. EmailBench provides a self-contained environment for end-to-end email-agent evaluation, with broader tool coverage, multi-persona testing, and repeated-run evaluation as future work areas.

Fri 25 SeptArtificial Intelligence
The gist
Handling enterprise email involves tasks like retrieving information, managing schedules, and coordinating multiple steps accurately. The authors created EmailBench, a set of 206 email-related tasks designed to test how well AI agents can perform these real-world workflows using a synthetic email dataset. They tested eight AI configurations and found that even the best one completed only about a third of tasks correctly, showing that just completing API actions doesn’t mean the task was truly done. This benchmark helps developers measure and improve AI email assistants in a controlled, realistic setting.
Open → 2609.31906v1

Interactive authoring tools improve science data visualization and analysis

"Here Be Sharks!": Enhancing Scientific Communication and Analysis through Authoring Interactivity

Abstract: We report on a case study for designing authoring environments for interactive visualizations to enhance scientific work. We conducted a workshop and prototype review with a group of marine biologists. When it came to visualizing their data, participants identified challenges in conveying their research accurately and completely as well as and analyzing it with ease. Based on our findings, authoring environments should consider the context of scientific work aspects such as domain expertise, collaboration culture, and publication traditions. We propose consideration of non-traditional programming languages and environments for scientific work, discuss ways of facilitating interactive visualizations for scientists, and examine ways of meaningfully integrating AI. The critiques, artifacts, and reactions from the scientists along with our analysis and discussion inform how computational tools should be designed for scientific work.

Tue 8 SeptHuman-Computer Interaction
The gist
Scientists often find it hard to show their data clearly and analyze it easily. This paper looks at how creating interactive tools for writing and showing data can help scientists, especially marine biologists. The authors found that tools should fit the scientists’ specific ways of working and sharing results. They also suggest using new kinds of programming and adding artificial intelligence to make these tools better.
Open → 2609.08386v1