Skip to content
AI Landscape

About

Scalable RLHF framework supporting 70B+ PPO full tuning, iterative DPO, and LoRA

See something off? Suggest an edit

Related in Fine-tuning & RLHF