ZooWork
ZooWork Market · Skill

openrlhf-training

High-performance RLHF framework with Ray+vLLM acceleration. Use for PPO, GRPO, RLOO, DPO training of large models (7B-70B+). Built on Ray, vLLM, ZeRO-3. 2× faster than DeepSpeedChat with distributed architecture and GPU resource sharing.

davila7
davila7-claude-code-templates-post-training-openrlhf · v1.0.0
分类ai-llms
安装次数21
更新时间2026-08-08T08:52:19.853Z
校验状态待验证
Post-TrainingOpenRLHFRLHFPPOGRPORLOODPORayvLLMDistributed TrainingLarge ModelsZeRO-3