All terms

Reinforcement Learning from Human Feedback

Training

التعلم بالتعزيز من تغذية راجعة بشرية

RLHFPronounced at-taʿallum bi-t-taʿzīz min taghdhiya rājiʿa basharīyyaExpert

English

RLHF is a training technique that uses human ratings of a model's responses to teach it which outputs people prefer, rather than relying only on the original training data. It is a key step in making chatbots more helpful, safe, and aligned with human expectations.

RLHF trains a model against human preferences: people rank competing outputs, those rankings train a reward model, and the language model is then optimised against that reward. It is the step that turned raw text predictors into assistants that decline, hedge, and follow instructions. It also means a model's manners are the encoded judgements of a particular group of annotators, which is a design decision rather than a neutral fact.

RLHF is why the assistant refuses some requests outright.

العربية

RLHF أسلوب تدريب يستخدم تقييمات بشرية لردود النموذج لتعليمه أي المخرجات يفضّلها الناس، بدلاً من الاعتماد فقط على بيانات التدريب الأصلية. وهو خطوة أساسية في جعل روبوتات المحادثة أكثر فائدة وأماناً وتوافقاً مع توقعات البشر.

يدرّب هذا الأسلوب النموذج على تفضيلاتٍ بشرية: إذ يرتّب أشخاصٌ مخرجاتٍ متنافسة، فتُدرَّب تلك الترتيبات نموذجَ مكافأة، ثم يُحسَّن النموذج اللغوي في مواجهة تلك المكافأة. وهو الخطوة التي حوّلت متنبئات النص الخام إلى مساعدين يعتذرون ويتحفظون ويتبعون التعليمات. ويعني هذا أيضاً أن آداب النموذج هي أحكامٌ مُرمَّزة لمجموعةٍ بعينها من المُعلِّمين، وهو قرار تصميمٍ لا حقيقةٌ محايدة.

التعلم بالتعزيز من تغذيةٍ راجعةٍ بشرية هو سبب رفض المساعد بعض الطلبات رفضاً صريحاً.

Also known as

  • RLHF
  • human feedback training
  • التعلم بالتعزيز من ملاحظات بشرية
  • آر إل إتش إف
  • RLHF

Also seen as, but not recommended

  • التعلم المعزز بالتغذية الراجعة البشرية

First use

2022 · Popularised by OpenAI's InstructGPT work.

Source

Related terms