SWE-QA: A Dataset and Benchmark for Complex Code Understanding
2026-04-29T08:51:52Z•8111611a7b0b05c93e165340ce78244381ca85b3e71c28942381851a277ca5a0
benchmarkbug-detectionchain-of-thoughtcode-understandingcontinuous-integrationdatasetlarge-language-modelsmodel-robustnesssoftware-testingtooling-securitytransformersvulnerability-detection
What happened
This collection of recent CS papers focuses on code understanding, automated vulnerability detection, and LLM-driven software engineering workflows. Key items: SWE-QA — a 9K-question multi-hop code comprehension benchmark exposing cross-file reasoning challenges; a broad 'Programming with Data' methodology for traceable, test-driven data engineering for LLMs; a systematic literature review of transformer-based vulnerability detection that highlights data imbalance, interpretability, scalability and generalization problems; FGDM — a flow-graph, multi-agent LLM framework using Chain-of-Thought/
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_se
- Record identifier
- 8111611a7b0b05c93e165340ce78244381ca85b3e71c28942381851a277ca5a0
- Enrichment time
- 2026-04-29T08:51:52Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.