ABTest: Behavior-Driven Testing for AI Coding Agents
2026-04-07T08:51:53Z•fad5b17e14ee0c0969cbbe1f13d0979491e09260476730b0ac0facdfa80c5114
LLM-trustRAGai-coding-agentsandroid-testingautonomous-debuggingbehavioral-testingcausality-miningconnected-vehiclecontinuous-integrationdatasetfuzzingmemory-safetymerge-conflictsmobile-ad-detectionreinforcement-learning-from-human-feedbackscaffold-taxonomysoftware-architecturesoftware-supply-chainstatic-dynamic-analysisstory-point-estimation
What happened
This collection of recent software-engineering and security-focused papers centers on robustness, tooling, and evaluation of AI-assisted development and complex distributed systems. Key contributions include ABTest, a behavior-driven fuzzing framework that turns real failure reports into repository-grounded tests to find anomalies in coding agents (Claude Code, Codex CLI, Gemini CLI); AgenticFlict, a large dataset (142K+ PRs) documenting merge conflicts in AI-generated pull requests; TRACE, a framework for eliciting artifact-level trust traces to measure LLM trust allocation across conflicting
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_se
- Record identifier
- fad5b17e14ee0c0969cbbe1f13d0979491e09260476730b0ac0facdfa80c5114
- Enrichment time
- 2026-04-07T08:51:53Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.