AIS: Adaptive Importance Sampling for Quantized RL
2026-05-15T07:24:00Z•304536c98b5585e30a9967d1725a61dbd17f1f359b9587ad0788b4d878ebd339
BF16Bayesian-inferenceDDIMFP8INT8LLM-inferenceMXFP4Multi-Scale-Dequantabstentionchain-of-thoughtconformal-predictioncovariancedequantizationdiffusion-modelsfairnesshardware-accelerationimportance-samplingmachine-learningmean-shiftonline-multiple-testingparticle-systems','training-free-sampling','MM-SOLD','kernel-ridquantizationregretreinforcement-learningsamplers
What happened
Batch of arXiv ML/statistics papers (May 15, 2026) covering advances in efficient LLM inference, quantized RL, diffusion-model sampling, sampling and Bayesian quadrature, conformal/chain-of-thought calibration, fairness in conformal prediction, online multiple testing, training-free generative sampling, and new theoretical results for kernel methods and generalization bounds. Key contributions: (1) AIS — Adaptive Importance Sampling to correct non-stationary rollout/trainer bias when using low-precision (FP8) rollouts with BF16 trainers in RL, preserving exploration while preventing training-
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_stat_ml
- Record identifier
- 304536c98b5585e30a9967d1725a61dbd17f1f359b9587ad0788b4d878ebd339
- Enrichment time
- 2026-05-15T07:24:00Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.