Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation

arXiv 2607.02577•66887d56c4b7f64dde7e4b210e5d9bd5ba86feca8da4fffa1dd479d1486f06b4
AgentLTLAgents4PentestChainSWEGovMemLLM agentsRLVRattack automationautonomous penetration testingbenchmark auditdecomposed metricsevaluation reliabilityfalse promotionknowledge architecturememory governancepen-testingprocedural complianceprovenancereproducibilityrisk managementsequential bug fixingsoftware maintenancetool-callingtrace inspection

Paper metadata

arXiv ID
2607.02577
Version
Not specified by this published record
Category
Computer Science — Software Engineering (cs.SE)

The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.

Evidence and limitations

Source ID
arxiv_cs_se
Record identifier
66887d56c4b7f64dde7e4b210e5d9bd5ba86feca8da4fffa1dd479d1486f06b4
Enrichment time
2026-07-07T08:51:47Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation · Baitaphish