BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models

2026-05-14T07:23:35Z37ad203f0d64821522920c4755df14b92512812d3e93f9c0b038795a7ee7766f
CPythonCoT-monitoringLLM-backdoorsLuaQuickJSRoPESafeContextautonomous-drivingbackdoor-detectioncontext-assemblydecision-time-assemblyhard-brakinghidden-objectivesjailbreaksmodel-extraction-benchmarks','GNN-extraction','watermarking-defemodel-unlearningoverride-hookspersona-conditioned-attacksphysical-adversarial-camouflagered-teamingscript-runtimessemantic-fuzzingsmall-model-monitoringtrajectory-manipulationwatermark-preservation

What happened

This RSS batch (arXiv CS/CR) collects several security- and ML-safety-focused papers: BackFlush — a knowledge-free backdoor detection and elimination framework that preserves legitimate watermarks using Rotation-based Parameter Editing (RoPE) unlearning and reports ≈1% ASR with ≈99% clean accuracy; Ghost in the Context — documents policy-carriage failures introduced during decision-time context assembly and proposes SafeContext mitigations; OverrideFuzz — a semantic-aware two-phase grammar fuzzer targeting override hooks in script runtimes (CPython, Lua, QuickJS); PCAP — persona-conditioned ad

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_cr
Record identifier
37ad203f0d64821522920c4755df14b92512812d3e93f9c0b038795a7ee7766f
Enrichment time
2026-05-14T07:23:35Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.