Skip to content

RLHF / Reinforcement Learning Overview ​

Translation pending

This page has not been translated yet. Please read the Chinese version.

The Classic Three-Stage Pipeline ​

A Unified Optimization Objective ​

Two Kinds of Rewards: Human Preference vs. Verifiable Rewards ​

Algorithm Evolution and Selection ​

Subtopic Navigation ​

References ​