OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
2026-07-16T08:52:18Z•cfd0b64daa81e607e783231c1afed555b9901bb8ad44e146b6f90ceff6fc008d
AI-insuranceAI-riskLLM-auditingOracle-Databaseagent-memoryagentic-AIbenchmarkingchain-of-thoughtdata-provenancegraph-neural-networkshuman-preferencesinterventional-auditsmachine-unlearningmodel-evaluationmolecular-predictionmulti-agent-systemsneuro-symbolic-AIprivacyprobabilistic-reasoningrecord-level-provenancerobotics-deploymentsafe-RLself-improving-agentssoftware-package
What happened
This collection is an arXiv feed (multiple new CS/AI papers) with several practical and evaluative contributions across provenance, agentic systems, grounding audits, neuro-symbolic reasoning, and safety. Key highlights: OriginBlame presents record- and token-level provenance to convert revocation requests into precise forget sets (reducing dataset-level over-deletion from 101x to 1.3x and improving unlearning by ~42% on a 1.7B model with modest throughput overhead). SPINE offers an agentic multi-agent deployment/debugging framework that raised teleoperation success (75%→100%) and reduced time
Why it matters
A reviewed impact interpretation has not been published for this record.
Evidence and limitations
- Source ID
- arxiv_cs_ai
- Record identifier
- cfd0b64daa81e607e783231c1afed555b9901bb8ad44e146b6f90ceff6fc008d
- Enrichment time
- 2026-07-16T08:52:18Z
- AI-assisted enrichment
- Yes
This record may overlap with other records. Its enrichment can be incomplete or wrong, and machine assistance was used. Validate consequential decisions against the linked source and your own environment.