Feedback Over Form: Why Execution Feedback Matters More Than Pipeline Topology in 1-3B Code Generation

2026-04-27T08:51:50Zb5ba3df4e3d39bb414e07fedc3becfc33e1b24ba0e06c90ada23773bf6230306
CARLARAGai-safetyautonomous-vehiclescall-chain-analysiscode-generationdynamic-validationethics-testingexecution-feedbackgenerative-aimaintenanceruntime-checkersruntime-monitoringself-refinementsilent-failuressimulationsoftware-engineering-researchstack-overflowstatic-analysistest-generation

What happened

Collection of recent SE/ML papers exploring LLM-enabled software engineering, testing, and safety. Key contributions: (1) "Feedback Over Form" shows small (1–3B) code models gain most from execute-refine feedback loops rather than complex pipeline topology; (2) FlyCatcher automatically synthesizes stateful runtime checkers from tests using LLM synthesis + static/dynamic validation, detecting many silent failures; (3) CAT (call-chain-aware test generation) improves unit-test coverage by modeling caller–callee chains and dependencies; (4) TRACE converts NHTSA crash reports into high-fidelity CAR

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_se
Record identifier
b5ba3df4e3d39bb414e07fedc3becfc33e1b24ba0e06c90ada23773bf6230306
Enrichment time
2026-04-27T08:51:50Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.