SWE-QA: A Dataset and Benchmark for Complex Code Understanding

2026-04-29T08:51:52Z8111611a7b0b05c93e165340ce78244381ca85b3e71c28942381851a277ca5a0
benchmarkbug-detectionchain-of-thoughtcode-understandingcontinuous-integrationdatasetlarge-language-modelsmodel-robustnesssoftware-testingtooling-securitytransformersvulnerability-detection

What happened

This collection of recent CS papers focuses on code understanding, automated vulnerability detection, and LLM-driven software engineering workflows. Key items: SWE-QA — a 9K-question multi-hop code comprehension benchmark exposing cross-file reasoning challenges; a broad 'Programming with Data' methodology for traceable, test-driven data engineering for LLMs; a systematic literature review of transformer-based vulnerability detection that highlights data imbalance, interpretability, scalability and generalization problems; FGDM — a flow-graph, multi-agent LLM framework using Chain-of-Thought/​

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_se
Record identifier
8111611a7b0b05c93e165340ce78244381ca85b3e71c28942381851a277ca5a0
Enrichment time
2026-04-29T08:51:52Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.