MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning
2026-05-27T07:23:30Z•2d840d0e623d0bd93f84f55e0d0d9686f35bf7cf6d1c1076e0afdaee3dbe29b8
DDoS detectionIDSIoT securityLLM securityOS sandboxRAGSDNSandlockadversarial attacksagent securityautonomybenchmarksdifferential privacyjailbreakmemory poisoningmetric differential privacymodel evaluationprompt injectionsafety engineeringsandboxingself-evolving agentstool hijacking
What happened
Collection of new 2026 arXiv papers exposing and addressing security risks in LLM-driven agents and related systems. Key offensive findings: MemMorph demonstrates high-success memory-poisoning attacks that bias long-term memory to hijack tool selection (up to 85.9% success with only three poisoned records); BITE shows black-box stylistic-edit attacks that inflate LLM-judge scores; Furina reveals an instability region in safety/refusal behavior enabling robust jailbreaks via uncertainty amplification. Other contributions: AgentSecBench formalizes measurement/games for instruction/retrieval/cap-
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cr
- Record identifier
- 2d840d0e623d0bd93f84f55e0d0d9686f35bf7cf6d1c1076e0afdaee3dbe29b8
- Enrichment time
- 2026-05-27T07:23:30Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.