Calibrated Inference for the Conditional Average Treatment Effect in the Few-Placebo Regime via Gaussian Processes
2026-05-28T07:24:00Z•788f024f0c525910011301cac3487e27fe9ba4639e041472b660a7076bc6f0e3
GRASPLLM safetyLoRAadversarial MLdataset generationdeception detectiondual-usefine-tuninggeometric signaturesmodel alignmentmulti-turn attacksprompt engineeringprompt-injectionspurious correlations
What happened
This arXiv batch includes multiple papers with security-relevant implications. The most relevant are: (1) “Evolving and Detecting Multi-Turn Deception using Geometric Signatures” — a method to generate realistic multi-turn deceptive prompt sequences via multi-objective genetic optimization and a lightweight detector using geometric embedding features (angular coverage, distance ratio, linearity plus pairwise similarities). The detector attains high recall (0.89) and test F1 in 0.74–0.86 across base, reworded and 3-turn truncations. (2) “Unsupervised Identification and Removal of Spurious Cor-
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_stat_ml
- Record identifier
- 788f024f0c525910011301cac3487e27fe9ba4639e041472b660a7076bc6f0e3
- Enrichment time
- 2026-05-28T07:24:00Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.