ZooWork
ZooWork Market · Skill

constitutional-ai

Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system.

davila7
davila7-claude-code-templates-safety-alignment-constitutional-ai · v1.0.0
分类ai-llms
安装次数67
更新时间2026-08-02T04:00:31.544Z
校验状态待验证
Safety AlignmentConstitutional AIRLAIFSelf-CritiqueHarmlessnessAnthropicAI SafetyRL From AI FeedbackClaude