Memory Layouts for GPU-Data Transfer Buffering in SPH
2026-06-24T08:52:22Z•234308b9901471a3c1a84895b6c798e255ddb4b7bc1482c7fa84c1f371ee0c78
BlackwellCUDA runtimeCXLFirecrackerGPU Confidential ComputingGPU host-device transferKV-cacheLLM servingMicroVM snapshotsNVLinkRDMAconcurrencyfederated learninghash tablememory coherencememory poolingmulti-tenant isolationperformance regressionprivacysynchronization
What happened
Collection of systems and HPC papers (GPU/host transfers, GPU Confidential Computing on Blackwell, RDMA/CXL memory pooling and hash-table designs, MicroVM snapshot serving, multi-LLM serving and KV-cache disaggregation, federated learning, concurrency/synchronization primitives, and GPU-aware bioinformatics structures). Key findings: (1) GPU-CC (Blackwell/RTX Pro/B300) preserves raw compute but serializes host↔device transfers via a confidential VM–GPU bridge, causing large throughput and latency regressions for LLM serving and KV-cache restore; scheduling/workarounds recover a portion of the损
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_dc
- Record identifier
- 234308b9901471a3c1a84895b6c798e255ddb4b7bc1482c7fa84c1f371ee0c78
- Enrichment time
- 2026-06-24T08:52:22Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.