Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs
2026-05-20T07:23:48Z•01007a99bccf45cf64ce73a9b3b277e36e892fdc32003073488f818d311093fe
AI-riskDAOLLMPLACES-datasetSME-securityT2I-safetyTRAILSXAIZero-Trustautomated-gradingblockchaindisclosure-designexplainabilityhallucinationinsurabilityinsurancelocalizationred-teamingrobustness-auditssynthetic-mediatraffic-datasettranscription-failurevision-LLM
What happened
This collection of recent arXiv papers covers security- and governance-adjacent research on AI systems, datasets, and architectures. Highlights include: an empirical evaluation of vision-capable LLM graders for handwritten mathematics showing high rubric-level accuracy but most errors (up to 87% in the best model) stem from image transcription failures, hallucinations, and equivalent-expression handling; a study of synthetic-media disclosure design that identifies tensions (normativity vs. neutrality, proactivity vs. precision) and the use of analogies (e.g., nutrition labels) in policy design
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cy
- Record identifier
- 01007a99bccf45cf64ce73a9b3b277e36e892fdc32003073488f818d311093fe
- Enrichment time
- 2026-05-20T07:23:48Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.