EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

arXiv 2609.01611v127491f64d11f7db8b3ea6436f4118d0eff92bab3e2a0b334fbb571a6643e6eb1

Paper metadata

arXiv ID
2609.01611
Version
v1
Category
cs.AI, cs.CL
Authors
Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
Publication date
2026-09-03T04:00:00Z
Source identifier
2609.01611v1
Public record ID
record:sha256:27491f64d11f7db8b3ea6436f4118d0eff92bab3e2a0b334fbb571a6643e6eb1

Source license ↗

The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.

This is source-provided metadata, not an enriched summary or an impact assessment. Follow the canonical source link for the published material.

Evidence and limitations

Source ID
arxiv_research
Record identifier
27491f64d11f7db8b3ea6436f4118d0eff92bab3e2a0b334fbb571a6643e6eb1
Record type
Source metadata

This record may overlap with other records. Source metadata can be incomplete or change. Validate consequential decisions against the linked source and your own environment.