MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models
2026-06-04T07:23:34Z•c9e518bbff91807402a6d5d356764ee30f5cdbf571ef88d65558f1e9b772cd55
activation-detectioncontextual-integritycovert-influencecredential-exfiltrationdiffusion-llmfile-type-detectionforensicsformal-verificationgnn-privacyhoneytokenshpkejailbreakmaskforgemembership-privacymemory-poisoningmerkle-logmimelensmodel-integritymodel-poisoningnotarized-receiptsparameter-attacksprivacyquery-rewritingsellostarkware
What happened
This feed collects multiple 2026 ML-security research contributions that raise urgent integrity, privacy, and supply-chain risks for deployed LLMs and agents. Key findings: MaskForge — a fully black-box, structure-aware attack — achieves ~79% average jailbreak success vs. diffusion LLMs, exposing infill-native exploit surfaces; Covert Influence demonstrates covert, human-undetectable behavioral payload transfer between models across fine-tuning, distillation, and in-context learning; credential-exfiltration work shows feasible pre-output and multi-turn attacks and proposes activation-based,_hh
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cr
- Record identifier
- c9e518bbff91807402a6d5d356764ee30f5cdbf571ef88d65558f1e9b772cd55
- Enrichment time
- 2026-06-04T07:23:34Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.