ZooWork
ZooWork Market · Skill

fine-tuning-with-trl

Fine-tune large language models (LLMs) using reinforcement learning with TRL — SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and tools for training reward models. Use when you need RLHF, want to align a model with human preferences, or train from human feedback. Works with Hugging Face Transformers.

◇
jk-kim0
jk-kim0-skills-jk-trl-fine-tuning · v1.0.0
分类ai-llms
安装次数11
更新时间2026-09-09T17:02:56.191Z
校验状态待验证