The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean

arXiv 2609.09218v1ff09bc8a276912bdf2b82cdf5cd5ca0387c50b1bd4fd940e92b523b0c59b1510

Paper metadata

arXiv ID
2609.09218
Version
v1
Category
cond-mat.mtrl-sci, cs.SE
Authors
Yonghong Zhang, Shadi Motaali, Vu Phong Dinh, Avin Piroutiniya, Jorge E. L\'opez de Vergara, Luis de Pedro, Ricardo Correia, Isabel M. Parra, Yong Xie
Publication date
2026-09-10T04:00:00Z
Source identifier
2609.09218v1
Public record ID
record:sha256:ff09bc8a276912bdf2b82cdf5cd5ca0387c50b1bd4fd940e92b523b0c59b1510

Source license ↗

The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.

This is source-provided metadata, not an enriched summary or an impact assessment. Follow the canonical source link for the published material.

Evidence and limitations

Source ID
arxiv_research
Record identifier
ff09bc8a276912bdf2b82cdf5cd5ca0387c50b1bd4fd940e92b523b0c59b1510
Record type
Source metadata

This record may overlap with other records. Source metadata can be incomplete or change. Validate consequential decisions against the linked source and your own environment.