BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models
2026-05-14T07:23:35Z•37ad203f0d64821522920c4755df14b92512812d3e93f9c0b038795a7ee7766f
CPythonCoT-monitoringLLM-backdoorsLuaQuickJSRoPESafeContextautonomous-drivingbackdoor-detectioncontext-assemblydecision-time-assemblyhard-brakinghidden-objectivesjailbreaksmodel-extraction-benchmarks','GNN-extraction','watermarking-defemodel-unlearningoverride-hookspersona-conditioned-attacksphysical-adversarial-camouflagered-teamingscript-runtimessemantic-fuzzingsmall-model-monitoringtrajectory-manipulationwatermark-preservation
What happened
This RSS batch (arXiv CS/CR) collects several security- and ML-safety-focused papers: BackFlush — a knowledge-free backdoor detection and elimination framework that preserves legitimate watermarks using Rotation-based Parameter Editing (RoPE) unlearning and reports ≈1% ASR with ≈99% clean accuracy; Ghost in the Context — documents policy-carriage failures introduced during decision-time context assembly and proposes SafeContext mitigations; OverrideFuzz — a semantic-aware two-phase grammar fuzzer targeting override hooks in script runtimes (CPython, Lua, QuickJS); PCAP — persona-conditioned ad
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cr
- Record identifier
- 37ad203f0d64821522920c4755df14b92512812d3e93f9c0b038795a7ee7766f
- Enrichment time
- 2026-05-14T07:23:35Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.