Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation

2026-07-07T08:51:47Z66887d56c4b7f64dde7e4b210e5d9bd5ba86feca8da4fffa1dd479d1486f06b4
AgentLTLAgents4PentestChainSWEGovMemLLM agentsRLVRattack automationautonomous penetration testingbenchmark auditdecomposed metricsevaluation reliabilityfalse promotionknowledge architecturememory governancepen-testingprocedural complianceprovenancereproducibilityrisk managementsequential bug fixingsoftware maintenancetool-callingtrace inspection

What happened

Collection of recent papers auditing and proposing methods for tool-using LLM agents, memory governance, benchmark validity, and LLM-driven pentesting. Key findings with security relevance: (1) Tool-calling benchmarks (BFCL, τ2-Bench, LiveMCPBench, MCP-Atlas) show high evaluator misalignment and nondeterministic scoring that can misrepresent agent capability and lead to unsafe deployment decisions; (2) GovMem demonstrates that naive memory promotion from agent traces causes large false-promotion rates, arguing for dependency-aware, review-gated governance to avoid propagating erroneous or copy

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_se
Record identifier
66887d56c4b7f64dde7e4b210e5d9bd5ba86feca8da4fffa1dd479d1486f06b4
Enrichment time
2026-07-07T08:51:47Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.