Benchmarking Open-Weight Foundation Models for Global AI Technical Governance

2026-06-26T07:23:49Z195a6f0d90eedb0031b4078b1794add6cdab75ed68086b7eb1a378364743c407
AI adoption indexAI governanceGAID v2LLM evaluationcopyright & AI musiceducation & AIfailure-mode analysisgeographic biasgovernance inversionhallucinationinformation disorderlegal riskmental-health riskmisinformation analysisopen-weight modelspersona simulationrecommender systemsregulatory riskresponse taxonomyvoice-cloning

What happened

Collection of new AI governance/measurement papers (arXiv, 26 Jun 2026) addressing systematic evaluation and societal risks of generative models. Key contributions: (1) Benchmarking open-weight frontier LLMs against the Global AI Dataset v2 (GAID v2) to quantify geographic bias using ~2,990 country-metric-year observations across 227 countries and a five-category response taxonomy (verified accuracy, hallucinated fabrication, honest refusal, qualitative hedging, misattribution); (2) Demonstration that LLM patient simulations (eating‑disorder personas) are overly stable but systematically over‑

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_cy
Record identifier
195a6f0d90eedb0031b4078b1794add6cdab75ed68086b7eb1a378364743c407
Enrichment time
2026-06-26T07:23:49Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.