English
RLHF is a training technique that uses human ratings of a model's responses to teach it which outputs people prefer, rather than relying only on the original training data. It is a key step in making chatbots more helpful, safe, and aligned with human expectations.
RLHF trains a model against human preferences: people rank competing outputs, those rankings train a reward model, and the language model is then optimised against that reward. It is the step that turned raw text predictors into assistants that decline, hedge, and follow instructions. It also means a model's manners are the encoded judgements of a particular group of annotators, which is a design decision rather than a neutral fact.
RLHF is why the assistant refuses some requests outright.
العربية
RLHF أسلوب تدريب يستخدم تقييمات بشرية لردود النموذج لتعليمه أي المخرجات يفضّلها الناس، بدلاً من الاعتماد فقط على بيانات التدريب الأصلية. وهو خطوة أساسية في جعل روبوتات المحادثة أكثر فائدة وأماناً وتوافقاً مع توقعات البشر.
يدرّب هذا الأسلوب النموذج على تفضيلاتٍ بشرية: إذ يرتّب أشخاصٌ مخرجاتٍ متنافسة، فتُدرَّب تلك الترتيبات نموذجَ مكافأة، ثم يُحسَّن النموذج اللغوي في مواجهة تلك المكافأة. وهو الخطوة التي حوّلت متنبئات النص الخام إلى مساعدين يعتذرون ويتحفظون ويتبعون التعليمات. ويعني هذا أيضاً أن آداب النموذج هي أحكامٌ مُرمَّزة لمجموعةٍ بعينها من المُعلِّمين، وهو قرار تصميمٍ لا حقيقةٌ محايدة.
التعلم بالتعزيز من تغذيةٍ راجعةٍ بشرية هو سبب رفض المساعد بعض الطلبات رفضاً صريحاً.
يُعرف أيضاً بـ
- RLHF
- human feedback training
- التعلم بالتعزيز من ملاحظات بشرية
- آر إل إتش إف
- RLHF
يُكتب أيضاً هكذا، ولا نوصي به
- التعلم المعزز بالتغذية الراجعة البشرية
أول ظهور
2022 · شاع مع عمل «إنستركت جي بي تي» في أوبن إيه آي.
المصدر