Papers for

content creators

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Framework makes human and AI understandings match better in tasks

Alignment Games: A Framework for Conceptual Repair in Human-AI Collaboration

Abstract: The meaning of a concept in use is shaped by the situation, task, goals, and prior knowledge. For example, a request to make a poster "visually appealing for a five-year-old" might evoke bright colors and cartoon imagery for one collaborator, but less text, bold shapes, and visual simplicity for another. We call such task-relevant differences conceptual misalignment. We introduce Alignment Games, a framework for making these differences visible and repairable during human-AI interaction. Drawing on theories of situated conceptualization, we characterize task-specific conceptual frames in terms of relevant attributes, values, relations, constraints, and priorities. We then define alignment moves that intervene on the situation, the reasoning used to interpret it, or the resulting frame. Through examples from educational content generation, creative coding, and argumentative writing, we show how these moves can be composed into repair sequences and derive design principles for supporting task-sufficient conceptual alignment at runtime.

Mon 28 SeptHuman-Computer InteractionArtificial Intelligence
The gist
People can have different ideas about what a task means, which causes confusion when working with AI. The authors introduce a method called Alignment Games that helps show and fix these differences while humans and AI collaborate. This method looks at what matters in a task, like priorities and rules, and allows changing either the situation or the way the AI thinks to reach better agreement. They demonstrate how this approach helps in creating educational materials, code, and arguments. This makes it easier for AI tools to understand what people really want.
Open → 2609.35197v1

Diffusion transformers enhance real-world image super-resolution efficiently

Fill2SR: Repurposing Inpainting Diffusion Transformers for Real-World Super-Resolution

Abstract: Recent real-world image super-resolution (SR) methods often adapt text-to-image (T2I) backbones with ControlNet-style branches or spatial conditioning tokens, which increases memory and computes with resolution and often constrains training to a fixed scale. We propose Fill2SR, which repurposes a masked-inpainting Diffusion Transformer for SR without extra spatial branches. Our Inpainting-Interface Evidence Adapter (IIEA) writes the low-quality (LQ) observation into the native masked-image slot under a full-image mask, turning inpainting into a reverse-degradation conditional rectified flow trained with LoRA-only tuning. We further introduce RCDT, an offline pipeline that distills degradation descriptors from unpaired real images and transfers them onto clean targets using frozen open-source models. Fill2SR supports mixed-resolution training up to QHD and yields stable performance across $512/1024/2048$ outputs. On synthetic benchmarks, our base model with IIEA achieves the best LPIPS on DIV2K and LSDIR; adding RCDT trades a small LPIPS drop for consistently stronger no-reference quality on RealLQ250 and RealPhoto60. Fill2SR remains memory-predictable, running $1536^2$ inference on a single 32GB GPU and extending to multi-megapixel outputs via tiled restoration.

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
Improving low-quality images to higher resolutions is usually costly in memory and limited by fixed scales. The authors introduce Fill2SR, which cleverly reuses an image-inpainting diffusion transformer to enhance images without extra memory-heavy parts. They also develop a way to learn about real image degradations and apply that knowledge to improve clean images. Their method works well on different image sizes and runs efficiently on common hardware.
Open → 2609.33582v1

Cognitive multi-agent system creates persuasive videos like humans

CogenPVG: Cognitive-Enhanced Reflective Multi-Agent Framework for Persuasive Video Generation

Abstract: Persuasive video generation (PVG) is a valuable yet under-explored research topic. Despite the significant advances in multimodal content generation, AI-empowered automated creation of human-made-like videos with substantial persuasiveness remains a formidable challenge. In this paper, we propose CogenPVG, a novel Cognitive-Enhanced reflective multi-agent framework tailored for Persuasive Video Generation task. Given the topic and stance from the user, we decouple the sophisticated generation process into four sequential stages: argument reasoning, storyboard planning, asset creation, and post-editing, imitating the workflow of human video producers. To ensure high persuasiveness, each stage is equipped with a pair of generator and critic agents, following a reflective refinement scheme grounded in a solid psychological theory of persuasion, the Elaboration Likelihood Model (ELM). In the argument reasoning stage, we generate highly logical and credible reasoning thoughts under the guidance of critical thinking theory, enabling cognitive enhancement via the central route of the ELM. For the other three stages, we generate and optimize multimodal assets, assembling them into a persuasive video guided by theories of heuristics, as the peripheral route of the ELM. To the best of our knowledge, CogenPVG is the first work focused on general persuasive topics, without being confined to commercial purposes. Extensive experiments and comprehensive analysis demonstrate that our framework achieves the best persuasion performance, thereby proving the effectiveness of our proposed multi-agent framework for the PVG task.

Tue 22 SeptMultimediaArtificial Intelligence
The gist
Making videos that can persuade people is hard for computers. The authors created a system called CogenPVG that works like a team of agents, each doing a step like thinking of arguments, planning scenes, making content, and editing. These agents check and improve each other's work using ideas from psychology about how people get convinced. This approach makes videos that feel logical and appealing, much like those made by humans.
Open → 2609.25821v1

Netflix reveals new way to measure recommendation impact accurately

Beyond Raw Engagement: A Counterfactual Observability Framework for Recommender Systems at Netflix

Abstract: Understanding the performance of large-scale recommender systems remains an underexplored challenge, especially for content creators and model developers. The raw engagement signals available to them, such as views and clicks, conflate content quality, model behavior, presentation bias, and audience reach, making it hard to attribute outcomes to the right cause. In this work, we present a general evaluation framework that enhances observability across multiple recommender systems at Netflix and demonstrate its effectiveness through several production deployments. The framework treats recommender-system observability as a counterfactual measurement problem: estimating what the recommender would have done, and what engagement would have followed, in the absence of a specific content item or model decision. We articulate three stakeholder-centered observability principles for content creators and model developers, and propose measurement methodologies covering bias reduction, relativity, and incrementality, applicable to both single-stage and cascading recommender systems and serving both audiences from a single measurement foundation.

Sat 19 SeptInformation RetrievalArtificial Intelligence
The gist
Netflix has a new way to figure out how well its recommendation system really works. Often, just looking at clicks or views doesn’t tell the full story because many things can influence those numbers. The authors created a method that estimates what would have happened if a certain show or algorithm choice hadn't been made. This helps creators and engineers understand the real reasons behind user engagement more clearly. It works across different recommendation stages and helps reduce bias in measuring results.
Open → 2609.22747v1

Ownership matters when people lead AI-assisted everyday tasks

Ownership in AI-Assisted Everyday Tasks

Abstract: When does work done with AI still feel like ours? As AI becomes woven into everyday tasks, we must examine what happens to our sense of ownership and contribution when a machine shares in producing what we make. We report an exploratory qualitative survey in which participants were asked to describe two recent, self-selected tasks completed with AI: one that felt like their own and one that did not. We find that felt ownership depends on the process of collaboration: people disown work when they merely approve AI's suggestions, but retain ownership when they lead, iterate, or rewrite. Ownership can also extend to settings where people own the vision for a project but not the execution; respondents reported high ownership on tasks they could not have completed without AI. Loss of personal voice and a lack of comprehension of the output both erode ownership. Finally, willingness to disclose AI use is often decoupled from actual pride or ownership, and instead shaped by community norms and fear of credit erasure. We propose several research directions as a result of these findings to promote AI development that supports people's sense of authorship over their own lives.

Thu 17 SeptArtificial IntelligenceHuman-Computer Interaction
The gist
This paper looks at when people still feel like the work they do with AI is really theirs. The authors find that people feel ownership when they actively guide or change the AI’s work, but less so when they just agree with what the AI suggests. People also keep ownership when they have the original idea, even if AI does much of the work. Losing a personal touch or not understanding the AI output can make people feel less connected to the work. Finally, whether people tell others they used AI often depends on social norms or worries about losing credit, not just pride in the work itself.
Open → 2609.20658v1