Scaling laws for reward model overoptimization

Published 2022-10-19T07:00:00Z1a032a768300cb63fd31ac90bb4a36fab5ad3d67c5e8fd5b2714045bbe68ad64

Source metadata

Publication date
2022-10-19T07:00:00Z
Source identifier
https://openai.com/index/scaling-laws-for-reward-model-overoptimization
Public record ID
record:sha256:1a032a768300cb63fd31ac90bb4a36fab5ad3d67c5e8fd5b2714045bbe68ad64

This is source-provided metadata, not an enriched summary or an impact assessment. Follow the canonical source link for the published material.

Evidence and limitations

Source ID
openai_research
Record identifier
1a032a768300cb63fd31ac90bb4a36fab5ad3d67c5e8fd5b2714045bbe68ad64
Record type
Source metadata

This record may overlap with other records. Source metadata can be incomplete or change. Validate consequential decisions against the linked source and your own environment.

Record · Scaling laws for reward model overoptimization · Baitaphish