The Hidden Costs of Domain Fine-Tuning: Pii-Bearing Data Degrades Safety and Increases Leakage

2026-03-04T19:46:30Zd71508205dd81d88559688eca10eb3ab5e51f8a7c575b6318a37e5615ed68813
LLM safetyPII leakageRAG poisoningReverse CAPTCHATEE (SGX)TalariaThreat modelingThreatFormer-IDSZK-PoPadversarial MLagentic AI skillsbehavioral biometrics (keystroke)confidential inferencedomain fine-tuningformal analysisintrusion detectioninvisible Unicodemetadata poisoningmultimodal securityprompt injectionskill marketplace securitysupply chain attackstoken-inference attackszero-day generalizationzero-knowledge proofs

What happened

A collection of recent arXiv security papers (Mar 3, 2026) highlighting multiple high-impact risks and defenses across AI, ML, and cyber operations. Key findings: domain fine-tuning with PII severely degrades refusal behavior and increases sensitive-data leakage in chat models; invisible Unicode control characters enable robust invisible instruction/prompt-injection (Reverse CAPTCHA); metadata-only multimodal poisoning (MM‑MEPA) can steer retrieval-augmented generation with high success rates; agent-skill marketplaces and supply chains are actively exploited (ClawHavoc), motivating formal, ver

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_cr
Record identifier
d71508205dd81d88559688eca10eb3ab5e51f8a7c575b6318a37e5615ed68813
Enrichment time
2026-03-04T19:46:30Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.