Stability and Robustness via Regularization: Bandit Inference via Regularized Stochastic Mirror Descent

2026-03-12T07:24:01Z65035e4a6e3239946db865d8d217d8ad95667294fbc39637bc6b791af58b587d
KSDLLM-evaluationMMDadversarial-robustnessbanditsbayesian-optimizationbias-detectiondata-synthesisdifferential-privacy (mitigation)distribution-shiftequivalence-testinggaussian-processeskernel-testslearning-ratemonitoringonline-learningoptimizationpoisoningpolicy-gradientprivacy-riskregularizationreinforcement-learningstabilitysynthetic-datatensor-clustering

What happened

Collection of ML/optimization papers with several security-relevant implications. Key points: (1) “Stability and Robustness via Regularization” shows regularized stochastic mirror-descent (regularized-EXP3) yields statistical stability and provable robustness to o(T^{1/2}) adversarial corruptions, while common algorithms (e.g., UCB) can fail under small corruption—implications for deployed adaptive sampling/bandit systems and poisoning/online-corruption attacks. (2) Policy-gradient diffusion results identify learning-rate regimes where regret becomes linear, indicating that misconfiguration or

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_stat_ml
Record identifier
65035e4a6e3239946db865d8d217d8ad95667294fbc39637bc6b791af58b587d
Enrichment time
2026-03-12T07:24:01Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.

Record · Stability and Robustness via Regularization: Bandit Inference via Regularized Stochastic Mirror Descent · Baitaphish