Papers for

social science research teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

PaperDoctor improves scientific paper feedback with evidence and action

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

Abstract: Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and does not scale. To shift automated paper assessment from a judge to a diagnostician, we introduce PaperDoctor, an agent framework for pre-submission feedback with three key innovations. First, a holistic hierarchical framework evaluates writing, layout, references, code, theory, prior work, and experiments through three layers: L1 surface screening, L2 typed verifiers that route each claim to the appropriate evidence, and L3 reproducers that rerun experiments by priority. Second, each finding contains an observation, a pointer to specific evidence such as a sentence, equation, or code line, and a revision suggestion, making critiques auditable and actionable. Third, PaperDoctor selectively rebuilds and reruns experiments based on claim importance and compute budget, surfacing reproducibility gaps and quantitative limitations that are invisible from the manuscript alone. We evaluate PaperDoctor on 30 in-progress papers, yielding 70.6% agreement and all positive holistic scores, and on 40 manuscripts across machine learning, natural science, and social science, covering human- and AI-authored papers with code. Overall, PaperDoctor produces more auditable feedback than human and other agentic reviewers, pairs critiques with concrete suggestions by design, and complements dimensions often overlooked by human reviewers. We also develop an interactive interface that lets authors browse findings grounded in their paper. PaperDoctor reframes automated paper assessment as diagnosis rather than verdict, taking a concrete step toward AI advisors for more rigorous AI-assisted scientific discovery.

Tue 15 SeptComputation and LanguageMultiagent Systems
The gist
Checking scientific papers before they are published can be slow and hard to do well with only humans. The PaperDoctor system helps by automatically finding problems in a paper’s writing, experiments, and references. It points to exact parts of the paper or code where issues appear and suggests how to fix them. This helps scientists improve their papers better and faster, making research clearer and more reliable.
Open → 2609.16995v1

AI coders debate and agree to improve qualitative coding accuracy

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding

Abstract: The utility of AI in multi-coder qualitative coding has been widely discussed, yet little empirical evidence exists to delineate the contexts in which it performs reliably. We address this gap by quantifying the effectiveness of multi-agent LLM coding across varied qualitative datasets, revealing key contextual and structural factors that mediate coding outcomes. We developed a literature-informed baseline pipeline that enables AI agents to independently code, debate, and reconcile disagreements. Results revealed that coding accuracy depends on factors such as codebook length, qualitative data similarity, and agent disagreement. Notably, intense and unresolved debates between agents led to higher accuracy. Our analysis showed that while LLMs emulate many human discussion behaviors, they lack adaptive responsiveness to context. From these findings, we offer design recommendations for building automated coding systems. Our open-source AI discussion dataset and methodological framework lay the groundwork for advancing the design of AI-mediated automated thematic analysis.

Thu 10 SeptHuman-Computer InteractionArtificial Intelligence
The gist
Figuring out how AI can help people label text data is tricky because it depends on many things. The authors studied how multiple AI agents code, argue, and come to agreements when labeling text from different sources. They found that longer instructions, how similar the data is, and how much the AI agents disagree affect how well the AI does. Surprisingly, harder debates between AI agents sometimes meant the final answers were more accurate. The authors also noticed AI behaves a bit like humans in discussions but doesn't adapt well to changing contexts.
Open → 2609.11109v1