Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation
2026-07-07T08:51:47Z•66887d56c4b7f64dde7e4b210e5d9bd5ba86feca8da4fffa1dd479d1486f06b4
AgentLTLAgents4PentestChainSWEGovMemLLM agentsRLVRattack automationautonomous penetration testingbenchmark auditdecomposed metricsevaluation reliabilityfalse promotionknowledge architecturememory governancepen-testingprocedural complianceprovenancereproducibilityrisk managementsequential bug fixingsoftware maintenancetool-callingtrace inspection
What happened
Collection of recent papers auditing and proposing methods for tool-using LLM agents, memory governance, benchmark validity, and LLM-driven pentesting. Key findings with security relevance: (1) Tool-calling benchmarks (BFCL, τ2-Bench, LiveMCPBench, MCP-Atlas) show high evaluator misalignment and nondeterministic scoring that can misrepresent agent capability and lead to unsafe deployment decisions; (2) GovMem demonstrates that naive memory promotion from agent traces causes large false-promotion rates, arguing for dependency-aware, review-gated governance to avoid propagating erroneous or copy
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_se
- Record identifier
- 66887d56c4b7f64dde7e4b210e5d9bd5ba86feca8da4fffa1dd479d1486f06b4
- Enrichment time
- 2026-07-07T08:51:47Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.