Papers for

software developers in ai

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Evolving recommender systems with reusable skill modules improves predictions

EvoSkillRec: Skill-Genome Evolution for Recommender Architecture Discovery

Abstract: Modern recommender systems advance not only by scaling data and parameters, but also by encoding task-specific inductive biases through architecture, including sparse feature interactions for click-through rate (CTR) prediction, temporal attention for sequential recommendation, and expert routing for multi-task learning. However, these biases are typically human expert designed or searched within predefined operator spaces. Although Recent LLM-driven code evolution expands this space, unconstrained edits often produce invalid or ineffective architectures, underuse established architecture design knowledge, and fail to preserve successful innovations for reuse. We introduce EvoSkillRec, a promotion-and-reuse framework for cumulative recommender architecture evolution. It first decomposes recommenders into atomic executable skills and represents architectures as typed skill genomes, with each skill equipped with input--output types, semantic annotations, and implementation code. We then evolve models with different tasks through two coupled spaces: a constrained skill--space that mutates, recombines, specializes, and reuses validated skills, and an open-ended code--space in which LLM planners and synthesizers invent new skill modules using prior evolution traces and accumulated experience. An autoresearch controller evaluates candidates, diagnoses failures, retrieves relevant skills, promotes validated innovations into the skill library, and adaptively allocates the proposal budget between the two spaces. Extensive experiments on CTR prediction, multi-task learning, and multi-domain learning, including resource-constrained co-optimization of predictive quality and model FLOPs utilization in generative ranking models, consistently demonstrate the effectiveness of our proposed EvoSkillRec.

Mon 28 SeptInformation Retrieval
The gist
Recommender systems help suggest things you might like, but building the best system can be tricky because they need special designs for each kind of task. The authors developed EvoSkillRec, a method that breaks down recommenders into small reusable skill pieces, which can be evolved and combined to create better architectures automatically. It uses both guided changes and new code invention powered by AI to improve the designs step by step. Their tests show this approach works well on various recommendation tasks, balancing prediction quality and computational cost.
Open → 2609.34552v1

Small reasoning models improve by asking bigger models for help

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

Abstract: Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.

Mon 28 SeptArtificial IntelligenceComputation and Language
The gist
Thinking longer doesn’t always help small language models solve hard problems if they lack the right knowledge. The authors found that these models tend to circle around known answers instead of discovering new ones on their own. They created FlyBy, a method where small models try to reason first, then ask bigger, smarter models for help only when needed. This approach improves problem solving while keeping computing costs low.
Open → 2609.34327v1

A unified theory explains how people learn concepts and chunks

A Unified Account of Concepts and Chunks

Abstract: Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are nearly disjoint, which poses a challenge for unified theories of cognition. In this paper, we review Cobweb, a computational account of categorization and concept formation, and propose an extended theory that incorporates chunks and their acquisition. The theory makes no commitments about modality, applying to any experience that decomposes into elements and relations among them. We also present \trellis/, an implementation of this theory, and illustrate its application to learning context-free grammars, which we adopt as a testbed because they involve both concept-like and chunk-like elements. In addition, we report experimental results on three synthetic grammars that demonstrate the system's ability to represent syntactic knowledge, use it to parse and generate sentences, and learn compositional structures from sample parses. We conclude by discussing related work on concepts and chunks, along with directions for future research in the area.

Thu 24 SeptComputation and LanguageArtificial Intelligence
The gist
People learn by recognizing categories (concepts) and familiar patterns (chunks), but these ideas have been studied separately. The authors review a system called Cobweb that models how concepts form, and they extend it to include chunks too. They created a tool named Trellis to test this idea using grammar rules, showing it can learn language-like patterns from examples. This work helps understand how we combine different kinds of knowledge in our minds.
Open → 2609.30414v1

Models organize reusable functions by linking parameters to components

The Ups and Downs of Backprop Weights

Abstract: Backpropagation (BP) has driven the remarkable success of modern deep learning by enabling large hierarchical networks to learn complex functions end-to-end. Yet it does not by itself determine how parameters should be organized so that functional components can be reused and adapted selectively. For example, object recognition and motion prediction may depend on overlapping parameter sets, making them difficult to isolate or modify independently. We call this condition weight entanglement. Modern architectures dynamically select which parts of a network process each sample: nonlinearities gate units, attention selects interactions, and Mixture-of-Experts architectures route inputs to modules. Yet such selection does not ensure that the same functional component remains linked to an identifiable parameter set across samples. We propose weight operators: parameterized modules that implement reusable functional components and can be composed at inference to form the function required by each sample. Learning proceeds in two stages: the model first infers the required operator composition, then updates only the selected operators' parameter sets. Vector Networks (VNs) provide one implementation. They couple operator selection to local error-driven updates within each layer and show that learned operators can be reused in combinations absent from training while updates remain restricted to the selected parameter sets. This provides a basis for testing functional parameter identifiability: whether an operator remains linked to the same functional component during learning. We argue that functional parameter identifiability may provide an organizing principle for models that systematically reuse and recombine learned functions while adapting only the components that need to change.

Fri 18 SeptMachine LearningArtificial Intelligence
The gist
Deep learning models learn by adjusting many parameters, but sometimes these parameters overlap across tasks, making it hard to change one function without affecting another. The authors identify this problem as weight entanglement and propose a way to organize parameters into reusable modules called weight operators. Their approach lets models select and update only the relevant modules for each task, keeping functions more separate and adaptable. This could help make learned skills easier to reuse and modify independently in future AI systems.
Open → 2609.22554v1