Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback
2026-07-03T08:51:56Z•7656cca52d7b4c9cd0de81de2a24ae12d69cc7824a364c5901d32dd9a26e3252
AI-governanceGPU-trainingGPUAlertIoTKaniLLMMatterPAIR-BenchPatchFusionRustSKILL.mdagent-skillsagentic-systemsbenchmarkingcode-generationdatasetformal-verificationinteroperabilitylabelled-corpusmodel-checkingmonitoringprogram-repairrisk-architectureskill-smellssoftware-engineering
What happened
This feed collects recent CS preprints on tooling, benchmarks, and practices for AI-assisted software engineering and related topics. Highlights include PAIR-Bench, a progressive/adaptive benchmark for feedback-guided code improvement; PatchFusion, a deterministic fusion method that improves pass@1 on multi-source patch pools (e.g., 426/500 on SWE-bench Verified); GPUAlert, a zero-instrumentation wrapper that classifies GPU training-job failures with 0.997 macro-F1 and ships a 474-log labelled corpus; Kani, a Rust model checker that verified thousands of harnesses, found six previously unknown
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_se
- Record identifier
- 7656cca52d7b4c9cd0de81de2a24ae12d69cc7824a364c5901d32dd9a26e3252
- Enrichment time
- 2026-07-03T08:51:56Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.