PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks

2026-05-13T07:23:32Zaaade3025ffb8c1868822033da27c36aebc7e2636b16aad77de14ec4d3517513
AgentShieldDCVDDPO attackFragBenchLLM watermarkingMT-JailBenchMerkle-DAG provenancePASAPortable Agent Memoryagent safetyauthorization-execution gapbenchmarksbenign fine-tuningcross-lingual defensescross-session detectiondeception-based detectionfragmented attacksindirect prompt injectionjailbreakinglocalizationmulti-turn jailbreaksopen-world agentsparaphrase robustnesssemantic watermarkingvulnerability detection

What happened

This collection of recent papers highlights new attack vectors, evaluation frameworks, and defenses in LLM- and agent-related security. Key contributions: PASA — a semantic (embedding-space) watermarking method robust to paraphrasing; a low-cost, few-shot “benign” DPO fine-tuning attack that reliably suppresses refusals and enables jailbreaking across multiple OpenAI models; MT-JailBench — a modular benchmark showing prompt generation and evaluation budgets strongly affect multi-turn jailbreak evaluations; a position paper identifying the Authorization‑Execution Gap (AEG) as a structural cause

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_cr
Record identifier
aaade3025ffb8c1868822033da27c36aebc7e2636b16aad77de14ec4d3517513
Enrichment time
2026-05-13T07:23:32Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.