TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs

2026-06-11T08:52:29Z8a8eb55bc2d651bd81b761444f1195057f25d441d53e879cf8a1095df4a6712b
AMD XDNA2AWQFP8ForeMoELLM hosting costsLLM inferenceMixture-of-ExpertsMoENPUsRL post-trainingRankGuardconcurrency-aware meteringdecentralized learningedge deploymentenergy efficiencyinfrastructure costload balancingmixed-precisionmodel poisoning defenseonline learning to rankperformance optimizationpoisoning attacksprocess mining on edge','edge computing','scheduling benchmarks'quantizationvllm-cost-meter

What happened

This batch of arXiv papers (June 11, 2026) covers systems and ML-infrastructure advances: TileFuse describes a close-to-metal mixed-precision kernel library to run AWQ-style low-bit quantized LLM inference efficiently on AMD XDNA2 NPUs, improving GEMM/GEMV throughput and energy use; a concurrency-aware cost methodology and the vllm-cost-meter quantify large utilization-driven variance in effective $/M-token for LLM hosting and expose severe underutilization penalties; ForeMoE proposes foresight-driven micro-step load balancing for MoE during RL post-training to reduce expert-imbalance overhead

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_dc
Record identifier
8a8eb55bc2d651bd81b761444f1195057f25d441d53e879cf8a1095df4a6712b
Enrichment time
2026-06-11T08:52:29Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.