Feedback Over Form: Why Execution Feedback Matters More Than Pipeline Topology in 1-3B Code Generation
2026-04-27T08:51:50Z•b5ba3df4e3d39bb414e07fedc3becfc33e1b24ba0e06c90ada23773bf6230306
CARLARAGai-safetyautonomous-vehiclescall-chain-analysiscode-generationdynamic-validationethics-testingexecution-feedbackgenerative-aimaintenanceruntime-checkersruntime-monitoringself-refinementsilent-failuressimulationsoftware-engineering-researchstack-overflowstatic-analysistest-generation
What happened
Collection of recent SE/ML papers exploring LLM-enabled software engineering, testing, and safety. Key contributions: (1) "Feedback Over Form" shows small (1–3B) code models gain most from execute-refine feedback loops rather than complex pipeline topology; (2) FlyCatcher automatically synthesizes stateful runtime checkers from tests using LLM synthesis + static/dynamic validation, detecting many silent failures; (3) CAT (call-chain-aware test generation) improves unit-test coverage by modeling caller–callee chains and dependencies; (4) TRACE converts NHTSA crash reports into high-fidelity CAR
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_se
- Record identifier
- b5ba3df4e3d39bb414e07fedc3becfc33e1b24ba0e06c90ada23773bf6230306
- Enrichment time
- 2026-04-27T08:51:50Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.