Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
2026-03-20T08:51:52Z•3e8d2da44754c1ea97ed16c803ff8e02163980c7f400d623d604e34f4df7a37d
Claude CodeDirect Preference Optimization (DPO)GitHub CopilotGreen AILLM agent safetyLLM-assisted code reviewRust verificationSQL comment generationSafeAuditSpaceTime ProgrammingVCoT-Benchadversarial PRsautomated theorem provingbenchmark coverageconfirmation biasdebiasingdesign discussion detectionomniscient debugging and tracing tools","BenchBrowser","benchmarrepository miningrule-resistancesoftware supply-chain attackssustainable ML practicestransformer modelsverification chain-of-thoughtvulnerability detection
What happened
This collection of recent CS/SE preprints highlights multiple security-relevant findings around LLMs and developer tooling. SafeAudit introduces an LLM-driven enumerator and a non-semantic metric (“rule-resistance”) to meta-audit LLM agent tool-call safety, finding >20% residual unsafe behaviors across benchmarks and environments and showing coverage improves with more tests. A separate study on LLM-assisted security code review demonstrates strong confirmation bias: framing a change as bug-free reduces vulnerability detection by 16–93%, and adversarial pull requests can reintroduce known CVEs
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_se
- Record identifier
- 3e8d2da44754c1ea97ed16c803ff8e02163980c7f400d623d604e34f4df7a37d
- Enrichment time
- 2026-03-20T08:51:52Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.