MultiTurnPSB: Evaluating Multi-Turn Jailbreak Attacks an dClassifier-Based Defenses for Medical AI Safety
2026-06-03T07:23:31Z•7dc69dcaf613e4f7950e3cfcd1affdc8073a8755c98cf453b76383c091ba0acd
CREEPD-JudgeISPMLITL (Lies-in-the-Loop)LLM safetyOWASP-LLM-Top-10RA-ICAagent refusalbyte-native LLMconsent integritydeepfake detectionfalse positivesidentity securityinference-cost attackinput-side classifierjailbreak attacksjudge model manipulationknowledge poisoningmalware analysismulti-turn jailbreakoutput rewritingparaphrase brittlenessretrieval-augmented generationrobustness
What happened
Collection of recent AI-security research describing growing practical risks and defenses for LLMs and agentic systems. Key findings include: multi-turn jailbreaks dramatically increase unsafe outputs (GPT-4.1-mini unsafe responses rose from ~35% to ~80% by Turn 4 under live adversary; a lightweight input-side classifier cut Turn‑4 unsafe responses by 52 percentage points but incurred ~45% false alarms); retrieval-augmented systems are vulnerable to inference-cost attacks via poisoning (RA-ICA / CREEP increased token consumption up to 13.12× with >90% success); judge-driven multi-turn attacks可
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cr
- Record identifier
- 7dc69dcaf613e4f7950e3cfcd1affdc8073a8755c98cf453b76383c091ba0acd
- Enrichment time
- 2026-06-03T07:23:31Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.