PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks
2026-05-13T07:23:32Z•aaade3025ffb8c1868822033da27c36aebc7e2636b16aad77de14ec4d3517513
AgentShieldDCVDDPO attackFragBenchLLM watermarkingMT-JailBenchMerkle-DAG provenancePASAPortable Agent Memoryagent safetyauthorization-execution gapbenchmarksbenign fine-tuningcross-lingual defensescross-session detectiondeception-based detectionfragmented attacksindirect prompt injectionjailbreakinglocalizationmulti-turn jailbreaksopen-world agentsparaphrase robustnesssemantic watermarkingvulnerability detection
What happened
This collection of recent papers highlights new attack vectors, evaluation frameworks, and defenses in LLM- and agent-related security. Key contributions: PASA — a semantic (embedding-space) watermarking method robust to paraphrasing; a low-cost, few-shot “benign” DPO fine-tuning attack that reliably suppresses refusals and enables jailbreaking across multiple OpenAI models; MT-JailBench — a modular benchmark showing prompt generation and evaluation budgets strongly affect multi-turn jailbreak evaluations; a position paper identifying the Authorization‑Execution Gap (AEG) as a structural cause
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cr
- Record identifier
- aaade3025ffb8c1868822033da27c36aebc7e2636b16aad77de14ec4d3517513
- Enrichment time
- 2026-05-13T07:23:32Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.