Papers for

benchmark developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Culture language and region annotations enhance web data benchmarking

FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing

Abstract: Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally situated phenomena at the level of languages, regions, and locale-specific practices, while web-scale corpora are usually organized only by language. A shared culture-language-region layer makes these resources comparable, enabling audits of whether a target cultural phenomenon is represented in pretraining data, evaluated by benchmarks or both. To this end, we introduce FineWeb-CLaR, a large-scale annotated dataset derived from FineWeb and FineWeb-2 that places web documents on a shared culture-language-region axis for corpus auditing and benchmark alignment. FineWeb-CLaR annotates the full 30.9B-document collection from FineWeb and FineWeb-2 with URL-derived region labels and cultural-topic provenance. Our region resolver assigns a non-empty region to 25.61% of documents (7.92B). For cultural-topic analysis, we induce locale-specific topics and project them onto the 14 leaves of the Cultural Taxonomy of Liu et al. (2025), producing Locale Topic Distributions (LTDs) for corpus-side comparison. We also annotate 277 cultural NLP benchmarks with the same taxonomy, language coverage, and region coverage. Together, these resources enable direct comparison between corpus-side pretraining evidence and benchmark-side evaluation coverage.

Mon 21 SeptComputation and Language
The gist
It’s hard to tell if language AI models understand different cultures because the training data and tests are labeled differently. The authors created FineWeb-CLaR, a big dataset that tags web documents with culture, language, and region information, making it easier to compare what’s in the training data with what tests measure. They also labeled many benchmarks to match this system, so it’s clearer if models truly cover cultural topics. This helps check if AI learns from or is tested on diverse cultural content.
Open → 2609.25298v1

Policy ambiguity causes unreliable agent evaluation scores

Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error

Abstract: Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assumption through silence, ambiguity, or contradiction, admitting multiple defensible readings that a single gold trajectory cannot capture. Auditing two $τ^2$-bench domains, we develop a taxonomy of such policy loopholes and show that affected tasks produce unreliable scores: they lower scores across different models in different ways and make every model less consistent across repeated trials. A cross-domain comparison reveals that exploitability requires both policy ambiguity and tool permissiveness: when policy complexity exceeds what tools can enforce, agents resolve gaps inconsistently and scores become unreliable. Policy specification quality sets the ceiling on evaluation quality. Benchmark developers should audit policies before collecting gold annotations.

Sun 13 SeptComputation and Language
The gist
Agent benchmarks often assume that a set of rules (policies) clearly say what an agent should do in every situation. This paper shows that real policies written in natural language can be vague or contradictory, allowing multiple right answers that benchmarks don’t capture well. As a result, scores given to agents can fluctuate and be misleading. The authors analyze tasks where policy ambiguity mixes with tool limits, causing inconsistent agent behavior and evaluation. They suggest that carefully checking policies before judging agents can improve evaluation reliability.
Open → 2609.14400v1

Protocol choices strongly affect hardware trojan detection accuracy

Protocol effects on feature-based hardware-Trojan detection across Trust-Hub families

Abstract: Trust-Hub reuses host circuits: several files differ mainly in the inserted Trojan. When gates from sibling variants enter both training and test folds, a detector can benefit from host logic it has already seen. We measure that effect instead of proposing another classifier. The corpus contains 49,124 gates from 16 netlists grouped into five host families. We left the parser, 36 gate features, class weighting, model settings, threshold, and family-level aggregation unchanged and altered one choice: the test boundary. The three settings draw test gates from the pooled corpus, withhold a complete netlist, or withhold every variant of one host. The choice matters. Random forest records F1/AP of 0.914/0.978 with pooled gates, 0.636/0.851 with one netlist held out, and 0.460/0.577 with a host family held out. XGBoost falls from 0.946/0.976 to 0.464/0.544 across the same comparison. Logistic regression loses AP, although its fixed-threshold F1 is not monotonic. Each family shows the same pooled-to-family direction. Feature removal, repeated model and simulator seeds, score normalization, parser-related exclusions, and a smaller sample change the size of the gap without reversing it. Aggregation also matters: a gate-weighted average is dominated by the larger ISCAS files, so the headline values give each host family one vote. Bootstrap and jackknife summaries keep the gap positive, but their folds reuse training families. We treat the five family rows as descriptive evidence rather than independent trials. Five host families are too few for a population claim, and the experiment says nothing about transfer to a new cell library or an industrial design. It supports a narrower conclusion: sibling benchmark variants can inflate apparent transfer. Benchmarks with several variants of one host circuit should report family-aware holdouts and all five family results beside pooled scores.

Mon 7 SeptMachine LearningArtificial IntelligenceComputational Engineering, Finance, and Science
The gist
Detecting hidden harmful circuits called hardware Trojans is important for computer chip security. The researchers found that when testing detection methods, including very similar circuits in both training and testing phases makes the detection seem more accurate than it really is. They showed that using stricter testing — where entire related sets of circuits are withheld — leads to much lower detection scores. This means that how test data is chosen can greatly affect evaluations of security tools.
Open → 2609.07199v1