كل المصطلحات

LLM-as-Judge

التقييم

النموذج حَكَماً

النطق an-namūdhaj ḥakamanمتقدم

English

Using one language model to score another's output. It scales evaluation far beyond what human raters can cover, and inherits the judging model's own biases while doing so.

Judging is cheaper than generating, so a judge model can evaluate thousands of responses for the cost of a small human study. The known failure modes are systematic rather than random: judges prefer longer answers, prefer their own family's phrasing, and are poor at detecting factual errors in domains where they are themselves weak, which includes, for most current models, non-English fluency. It is a useful instrument that should not be the only one.

LLM-as-judge scored 4,000 responses overnight.

العربية

استعمال نموذجٍ لغوي لتقييم مخرجات نموذجٍ آخر. يوسّع التقييم إلى ما يتجاوز طاقة المُقيِّمين البشر بكثير، ويرث في الوقت نفسه تحيزات النموذج الحاكم.

الحكم أرخص من التوليد، فيستطيع نموذجٌ حاكم تقييم آلاف الإجابات بكلفة دراسةٍ بشريةٍ صغيرة. وأنماط إخفاقه المعروفة منهجيةٌ لا عشوائية: إذ يفضّل الإجاباتِ الأطول، ويفضّل صياغة عائلته النموذجية، ويضعف في كشف الأخطاء الواقعية في المجالات التي هو نفسه ضعيفٌ فيها، ومنها، في معظم النماذج الحالية، الطلاقةُ بغير الإنجليزية. فهو أداةٌ نافعة لا ينبغي أن تكون الوحيدة.

قيّم النموذجُ حَكَماً أربعةَ آلاف إجابةٍ في ليلة.

يُعرف أيضاً بـ

  • model-based evaluation
  • AI judge
  • التقييم بالنموذج
  • المحكّم الآلي

مصطلحات ذات صلة