Mutual evaluation method enables truthful reporting without peer comparisons
Mutual Evaluation and Supervision without Peers
Computer Science and Game TheoryInformation Theory
Summary
This paper looks at how a worker and a critic can report honestly on completed tasks without needing to compare results with others. The authors propose a system where multiple versions of a task are done independently, and the critic gives scores based on patterns across these reports. This removes the need for direct peer comparisons or having a known correct answer. The method can calculate unbiased information scores while considering strategic behavior of both the worker and the critic.
What this means in practice
- •For crowdsourcing platform designers: Create task evaluation mechanisms that encourage honest responses without relying on comparisons to other workers or known answers.
- •For data quality engineers: Develop quality control processes using replicated task runs that measure information validity without requiring ground truth labels.
A theory result. No direct application yet.
Authors
Zachary Robertson
Abstract
This article introduces mutual evaluation of a replicable task worker and a critic that incentivizes truthful reporting, both modeled as strategic agents. The critic chooses a finite-valued rule that induces an evaluation score on joint report laws. Their common payoff is analyzed through regret relative to the unrestricted critic envelope. The critic rule is distinct from the evaluation score. This class enables a peer-free information elicitation mechanism using conditionally independent replications of a worker on the same task. This replication-loop mechanism implements a type-agreement payoff using same-task replications and new-task samples. In contrast to the peer-prediction and scoring-rule literature, implementations are shown that produce unbiased Pearson and Shannon information scores without requiring peers, a ground-truth reference, or likelihood-ratio estimation. A valid binary critic also can be represented by shared finite type annotations of worker returns. One runtime restriction is that the number of required replicas is random and can depend on the critic rule. Other timing effects, such as commitment and reoptimization, yield distinct incentives, connecting the framework to variational peer prediction. This mechanism class illustrates why strategic considerations matter for both critic and worker agents.