Where Do Large Language Models Fail on Competitive Programming? A Taxonomy of Failures by Algorithm Type and Difficulty Rating
arXiv 2606.05228•de383ed8e5ca9188fe7e4f1e3c2ec6715243117b81302c751947f6c5db636fa9
AWS CDKMCPPLC/industrial automationagent oversightautonomous agentsbenchmarkschain-of-thoughtcompetitive programmingdatasetsdeploymentinfrastructure-as-codelarge language modelsmutation testingreverse engineeringruntime faultssecurity validationsoftware engineering
Paper metadata
- arXiv ID
- 2606.05228
- Version
- Not specified by this published record
- Category
- Computer Science — Software Engineering (cs.SE)
The PDF link points to arxiv.org. Baitaphish does not expose a private stored PDF.
Evidence and limitations
- Source ID
- arxiv_cs_se
- Record identifier
- de383ed8e5ca9188fe7e4f1e3c2ec6715243117b81302c751947f6c5db636fa9
- Enrichment time
- 2026-06-05T08:51:51Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.