ZooWork Market · Skill
fine-tuning-with-trl
Fine-tune large language models (LLMs) using reinforcement learning with TRL — SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and tools for training reward models. Use when you need RLHF, want to align a model with human preferences, or train from human feedback. Works with Hugging Face Transformers.