Where Do Large Language Models Fail on Competitive Programming? A Taxonomy of Failures by Algorithm Type and Difficulty Rating
2026-06-05T08:51:51Z•de383ed8e5ca9188fe7e4f1e3c2ec6715243117b81302c751947f6c5db636fa9
AWS CDKMCPPLC/industrial automationagent oversightautonomous agentsbenchmarkschain-of-thoughtcompetitive programmingdatasetsdeploymentinfrastructure-as-codelarge language modelsmutation testingreverse engineeringruntime faultssecurity validationsoftware engineering
What happened
This is a collection of recent AI/SE research (arXiv 05 Jun 2026) presenting benchmarks, taxonomies, and datasets that reveal reliability, safety, and deployment gaps in LLM-driven software and agent ecosystems. Key items: an empirical failure taxonomy for LLMs on competitive programming showing CoT can worsen algorithmic correctness; DeployBench and SWE-InfraBench exposing fragile artifact deployment and IaC editing failures; a first empirical taxonomy of runtime faults in MCP (Model Context Protocol) servers that explicitly calls out security-validation and integration fault classes; REStack
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_se
- Record identifier
- de383ed8e5ca9188fe7e4f1e3c2ec6715243117b81302c751947f6c5db636fa9
- Enrichment time
- 2026-06-05T08:51:51Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.