Equivalence between policy gradients and soft Q-learning
Published 2017-04-21T07:00:00Z•0ec1aa4d25c6af09dde1efad23e4142b543a621671636338947442477e238914
Source metadata
- Publication date
- 2017-04-21T07:00:00Z
- Source identifier
- https://openai.com/index/equivalence-between-policy-gradients-and-soft-q-learning
- Public record ID
- record:sha256:0ec1aa4d25c6af09dde1efad23e4142b543a621671636338947442477e238914
This is source-provided metadata, not an enriched summary or an impact assessment. Follow the canonical source link for the published material.
Evidence and limitations
- Source ID
- openai_research
- Record identifier
- 0ec1aa4d25c6af09dde1efad23e4142b543a621671636338947442477e238914
- Record type
- Source metadata
This record may overlap with other records. Source metadata can be incomplete or change. Validate consequential decisions against the linked source and your own environment.