A publication count is not an accuracy score. Evidence reference checks and semantic support evaluation answer different questions. See how Research works.

Snapshot updated .

Continuous Eval V1 has no reconciled public results yet. The legacy inventory below does not measure the new GPT-6 pipeline.

Legacy RI1

Methodology: legacy-ri1. Current published inventory; no processing-window denominator available.

Comparable model-profile telemetry is unavailable for this inventory.

Current accepted public versions; other populations unavailable
Outcome / stageSource versions
Source versions discoveredUnavailable
Source eligibleUnavailable
AcquiredUnavailable
NormalizedUnavailable
Classification admittedUnavailable
Synthesis attemptedUnavailable
Semantically evaluatedUnavailable
Publication eligibleUnavailable
Published10
QuarantinedUnavailable
AbstainedUnavailable
Capacity deferredUnavailable
Budget deferredUnavailable
Infrastructure failedUnavailable
Evaluation measurements with their own denominators
DimensionObserved / assessed
Evidence sufficiencyUnavailable
Claim coverageUnavailable
Semantic supportUnavailable
Scope correctnessUnavailable
Wrong lensUnavailable
Quantitative validationUnavailable
Limitations preservedUnavailable
Citation and provenance integrityUnavailable
Normalization dependencyUnavailable
Material omissionUnavailable
Section fitUnavailable
Publication usefulnessUnavailable

Observed failure categories

No comparable failure-category data is available in this snapshot.

Usage and latency

Comparable model usage and cost are unavailable.

Unavailable means the retained data cannot establish that measurement. Cohorts are not combined into a headline pass rate. Counts describe automated processing outcomes and do not establish human scientific approval.