Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

2026-06-05T07:23:31Z86d40349e1d128a6b24a4f7d70e646dac7a6f16d7e00bbc09ef8b0c54f58663b
Bitcoin economicsCRESS scoreCWE-89DeFiGDPR complianceLLM safetyMEVMITRE ATT&CKOS hardeningSHIELDSSigma rulesZERO-APTabliterationautomated pentestingautomated remediationbenchmarkingconfidential computingdetection-as-codehardware reverse engineeringmodel weight editprompt injectionred-teamingreproducibilitysearch-time contamination

What happened

Collection of 10 new arXiv papers (Jun 5 2026) covering safety, evaluation, and automation risks in modern AI and security systems. Highlights include: CUA-HandCrafted — a 793-episode benchmark showing frontier browser-focused CUAs resist hand-crafted prompt-injection (0/140 multi-step successes) but the same models are highly vulnerable in coding-agent tasks (up to 100% skill-injection); a study of Search-Time Contamination (STC) showing web-retrieving agents can leak benchmark metadata/answers and inflate scores (up to ~4%); deterministic synthesis from BAS findings to Sigma rules with probe

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_cr
Record identifier
86d40349e1d128a6b24a4f7d70e646dac7a6f16d7e00bbc09ef8b0c54f58663b
Enrichment time
2026-06-05T07:23:31Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.

Record · Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming · Baitaphish