Benchmarking Open-Weight Foundation Models for Global AI Technical Governance
2026-06-26T07:23:49Z•195a6f0d90eedb0031b4078b1794add6cdab75ed68086b7eb1a378364743c407
AI adoption indexAI governanceGAID v2LLM evaluationcopyright & AI musiceducation & AIfailure-mode analysisgeographic biasgovernance inversionhallucinationinformation disorderlegal riskmental-health riskmisinformation analysisopen-weight modelspersona simulationrecommender systemsregulatory riskresponse taxonomyvoice-cloning
What happened
Collection of new AI governance/measurement papers (arXiv, 26 Jun 2026) addressing systematic evaluation and societal risks of generative models. Key contributions: (1) Benchmarking open-weight frontier LLMs against the Global AI Dataset v2 (GAID v2) to quantify geographic bias using ~2,990 country-metric-year observations across 227 countries and a five-category response taxonomy (verified accuracy, hallucinated fabrication, honest refusal, qualitative hedging, misattribution); (2) Demonstration that LLM patient simulations (eating‑disorder personas) are overly stable but systematically over‑
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_cy
- Record identifier
- 195a6f0d90eedb0031b4078b1794add6cdab75ed68086b7eb1a378364743c407
- Enrichment time
- 2026-06-26T07:23:49Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.