Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback

arXiv 2603.16928•b521aaf58e950616995b749b58b4e0b0f14dfc0bb678958f265ea46b0927a468
AI governanceAPTDNA OTPDRLLLM monitoringLLM security assessmentPAuthVLM robustnessadversarial MLadversarial attacksagent evasionautonomous defensebackdoor detectionchain-of-thoughtcode generation securitycryptographic policygradient attackslog forensicsmodel poisoningprovenance graphsruntime enforcementtask-scoped authorizationunconditional cryptographyvision-language models

Paper metadata

arXiv ID
2603.16928
Version
Not specified by this published record
Category
Computer Science — Cryptography and Security (cs.CR)

The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.

Evidence and limitations

Source ID
arxiv_cs_cr
Record identifier
b521aaf58e950616995b749b58b4e0b0f14dfc0bb678958f265ea46b0927a468
Enrichment time
2026-03-19T07:23:31Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.