REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming
2026-07-10T08:51:47Z•f3700e1cbbdb6bb93b06e5268623b61f6682b6f1292edf16c10d7f1abb66806b
CWECodeGuard+DWARFDeepSWE','long-horizon tasks','SWE benchmarks','test verifiers'LLMsPERFOPT-BenchSecVecCoderTrajSpecagentic codingautomated program repairbenchmark safetybenchmarkingbinary-to-source alignmentbug report refinementcompiler optimizationdecompilationfunction namingperformance optimizationprovenancerepository analysisreverse engineeringsecure code generationsurvivorship biastask vectorsuncertainty-aware evaluation
What happened
This RSS batch summarizes a set of July 10, 2026 arXiv submissions focused on LLMs applied to software engineering, benchmarking, and tool correctness. Key papers: REFORGE — a provenance-tracked pipeline for constructing function-level ground truth through compilation, DWARF extraction and decompilation; it formalizes alignment uncertainty (an eight-gate confidence funnel) and shows optimization levels sharply reduce high-confidence evaluable functions and cause survivorship bias in evaluations. PERFOPT-Bench — a benchmark for agentic performance optimization (profiling, diagnosing, editing, &
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_se
- Record identifier
- f3700e1cbbdb6bb93b06e5268623b61f6682b6f1292edf16c10d7f1abb66806b
- Enrichment time
- 2026-07-10T08:51:47Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.