TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs
2026-06-11T08:52:29Z•8a8eb55bc2d651bd81b761444f1195057f25d441d53e879cf8a1095df4a6712b
AMD XDNA2AWQFP8ForeMoELLM hosting costsLLM inferenceMixture-of-ExpertsMoENPUsRL post-trainingRankGuardconcurrency-aware meteringdecentralized learningedge deploymentenergy efficiencyinfrastructure costload balancingmixed-precisionmodel poisoning defenseonline learning to rankperformance optimizationpoisoning attacksprocess mining on edge','edge computing','scheduling benchmarks'quantizationvllm-cost-meter
What happened
This batch of arXiv papers (June 11, 2026) covers systems and ML-infrastructure advances: TileFuse describes a close-to-metal mixed-precision kernel library to run AWQ-style low-bit quantized LLM inference efficiently on AMD XDNA2 NPUs, improving GEMM/GEMV throughput and energy use; a concurrency-aware cost methodology and the vllm-cost-meter quantify large utilization-driven variance in effective $/M-token for LLM hosting and expose severe underutilization penalties; ForeMoE proposes foresight-driven micro-step load balancing for MoE during RL post-training to reduce expert-imbalance overhead
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- 8a8eb55bc2d651bd81b761444f1195057f25d441d53e879cf8a1095df4a6712b
- Enrichment time
- 2026-06-11T08:52:29Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.