Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming
2026-06-05T07:23:31Z•86d40349e1d128a6b24a4f7d70e646dac7a6f16d7e00bbc09ef8b0c54f58663b
Bitcoin economicsCRESS scoreCWE-89DeFiGDPR complianceLLM safetyMEVMITRE ATT&CKOS hardeningSHIELDSSigma rulesZERO-APTabliterationautomated pentestingautomated remediationbenchmarkingconfidential computingdetection-as-codehardware reverse engineeringmodel weight editprompt injectionred-teamingreproducibilitysearch-time contamination
What happened
Collection of 10 new arXiv papers (Jun 5 2026) covering safety, evaluation, and automation risks in modern AI and security systems. Highlights include: CUA-HandCrafted — a 793-episode benchmark showing frontier browser-focused CUAs resist hand-crafted prompt-injection (0/140 multi-step successes) but the same models are highly vulnerable in coding-agent tasks (up to 100% skill-injection); a study of Search-Time Contamination (STC) showing web-retrieving agents can leak benchmark metadata/answers and inflate scores (up to ~4%); deterministic synthesis from BAS findings to Sigma rules with probe
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cr
- Record identifier
- 86d40349e1d128a6b24a4f7d70e646dac7a6f16d7e00bbc09ef8b0c54f58663b
- Enrichment time
- 2026-06-05T07:23:31Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.