The Hidden Costs of Domain Fine-Tuning: Pii-Bearing Data Degrades Safety and Increases Leakage
2026-03-04T19:46:30Z•d71508205dd81d88559688eca10eb3ab5e51f8a7c575b6318a37e5615ed68813
LLM safetyPII leakageRAG poisoningReverse CAPTCHATEE (SGX)TalariaThreat modelingThreatFormer-IDSZK-PoPadversarial MLagentic AI skillsbehavioral biometrics (keystroke)confidential inferencedomain fine-tuningformal analysisintrusion detectioninvisible Unicodemetadata poisoningmultimodal securityprompt injectionskill marketplace securitysupply chain attackstoken-inference attackszero-day generalizationzero-knowledge proofs
What happened
A collection of recent arXiv security papers (Mar 3, 2026) highlighting multiple high-impact risks and defenses across AI, ML, and cyber operations. Key findings: domain fine-tuning with PII severely degrades refusal behavior and increases sensitive-data leakage in chat models; invisible Unicode control characters enable robust invisible instruction/prompt-injection (Reverse CAPTCHA); metadata-only multimodal poisoning (MM‑MEPA) can steer retrieval-augmented generation with high success rates; agent-skill marketplaces and supply chains are actively exploited (ClawHavoc), motivating formal, ver
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cr
- Record identifier
- d71508205dd81d88559688eca10eb3ab5e51f8a7c575b6318a37e5615ed68813
- Enrichment time
- 2026-03-04T19:46:30Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.