OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

2026-07-16T08:52:18Zcfd0b64daa81e607e783231c1afed555b9901bb8ad44e146b6f90ceff6fc008d
AI-insuranceAI-riskLLM-auditingOracle-Databaseagent-memoryagentic-AIbenchmarkingchain-of-thoughtdata-provenancegraph-neural-networkshuman-preferencesinterventional-auditsmachine-unlearningmodel-evaluationmolecular-predictionmulti-agent-systemsneuro-symbolic-AIprivacyprobabilistic-reasoningrecord-level-provenancerobotics-deploymentsafe-RLself-improving-agentssoftware-package

What happened

This collection is an arXiv feed (multiple new CS/AI papers) with several practical and evaluative contributions across provenance, agentic systems, grounding audits, neuro-symbolic reasoning, and safety. Key highlights: OriginBlame presents record- and token-level provenance to convert revocation requests into precise forget sets (reducing dataset-level over-deletion from 101x to 1.3x and improving unlearning by ~42% on a 1.7B model with modest throughput overhead). SPINE offers an agentic multi-agent deployment/debugging framework that raised teleoperation success (75%→100%) and reduced time

Why it matters

A reviewed impact interpretation has not been published for this record.

Evidence and limitations

Source ID
arxiv_cs_ai
Record identifier
cfd0b64daa81e607e783231c1afed555b9901bb8ad44e146b6f90ceff6fc008d
Enrichment time
2026-07-16T08:52:18Z
AI-assisted enrichment
Yes

This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.