Backdoor Attacks on Decentralised Post-Training
2026-04-06T07:23:32Z•ab5ae783c63840c990a787c597d5067f49ab2cc47156295503ff74184b9515e3
CAPECCWELLM ensemblesLLM safetyLLM unalignmentORAM decoupling','trusted enclave' (note: keep as plain tag)Opalagent memoryalignment attackbackdoordataset generationdual-use dataseteTAMPenvironmental attackjailbreak-tuningmalware classificationmemory poisoningmodel poisoningpipeline parallelismpost-trainingprivate memoryvulnerable code datasetweb agentsweight orthogonalizationzero-label classification
What happened
This collection of papers (arXiv 2604.*) highlights multiple emerging security risks and some corresponding mitigations in AI, ML, and cyber systems. Key threats: a novel backdoor attack on pipeline-parallel post-training where an adversary controlling an intermediate pipeline stage can inject alignment-breaking backdoors; environment-injected trajectory-based memory poisoning (eTAMP) that contaminates web-agent memories via passive observation to achieve cross-session, cross-site compromise; safety “unalignment” techniques (weight orthogonalization and jailbreak-tuning) that meaningfully re‑e
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cr
- Record identifier
- ab5ae783c63840c990a787c597d5067f49ab2cc47156295503ff74184b9515e3
- Enrichment time
- 2026-04-06T07:23:32Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.