Deliberative alignment: reasoning enables safer language models

Published 2024-12-20T10:00:00Z9cf289b414fdf955889f2f6da89cd07f3e9687741a433deba67d3e5322dc8330

Source metadata

Publication date
2024-12-20T10:00:00Z
Source identifier
https://openai.com/index/deliberative-alignment
Public record ID
record:sha256:9cf289b414fdf955889f2f6da89cd07f3e9687741a433deba67d3e5322dc8330

This is source-provided metadata, not an enriched summary or an impact assessment. Follow the canonical source link for the published material.

Evidence and limitations

Source ID
openai_security
Record identifier
9cf289b414fdf955889f2f6da89cd07f3e9687741a433deba67d3e5322dc8330
Record type
Source metadata

This record may overlap with other records. Source metadata can be incomplete or change. Validate consequential decisions against the linked source and your own environment.

Record · Deliberative alignment: reasoning enables safer language models · Baitaphish