REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming

2026-07-10T08:51:47Zf3700e1cbbdb6bb93b06e5268623b61f6682b6f1292edf16c10d7f1abb66806b
CWECodeGuard+DWARFDeepSWE','long-horizon tasks','SWE benchmarks','test verifiers'LLMsPERFOPT-BenchSecVecCoderTrajSpecagentic codingautomated program repairbenchmark safetybenchmarkingbinary-to-source alignmentbug report refinementcompiler optimizationdecompilationfunction namingperformance optimizationprovenancerepository analysisreverse engineeringsecure code generationsurvivorship biastask vectorsuncertainty-aware evaluation

What happened

This RSS batch summarizes a set of July 10, 2026 arXiv submissions focused on LLMs applied to software engineering, benchmarking, and tool correctness. Key papers: REFORGE — a provenance-tracked pipeline for constructing function-level ground truth through compilation, DWARF extraction and decompilation; it formalizes alignment uncertainty (an eight-gate confidence funnel) and shows optimization levels sharply reduce high-confidence evaluable functions and cause survivorship bias in evaluations. PERFOPT-Bench — a benchmark for agentic performance optimization (profiling, diagnosing, editing, &

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_se
Record identifier
f3700e1cbbdb6bb93b06e5268623b61f6682b6f1292edf16c10d7f1abb66806b
Enrichment time
2026-07-10T08:51:47Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.