English
The software that runs a model against a suite of tests reproducibly, fixed prompts, fixed scoring, fixed conditions, so that two results can legitimately be compared.
Scores are only comparable when everything except the model is held constant, and a surprising amount can vary: prompt phrasing, the number of examples, whether answers are parsed leniently, even temperature. A harness fixes those choices in code. Most published disagreements about which model is better dissolve once both sides run the same harness.
Both models were scored on the same evaluation harness.
العربية
البرمجية التي تُشغّل النموذج على مجموعة اختباراتٍ تشغيلاً قابلاً للتكرار، بتوجيهاتٍ ثابتة وتسجيلٍ ثابت وظروفٍ ثابتة، حتى تصح المقارنة بين نتيجتين.
لا تصح مقارنة النتائج إلا حين يثبت كل شيءٍ عدا النموذج، والمتغيرات أكثر مما يُظن: صياغة التوجيه، وعدد الأمثلة، وأيُّ تساهلٍ في تحليل الإجابات، وحتى درجة الحرارة. والمنصة تثبّت تلك الخيارات في الشيفرة. ومعظم الخلافات المنشورة حول أيُّ النماذج أفضل تتلاشى متى شغّل الطرفان المنصة نفسها.
قُيِّم النموذجان على منصة التقييم نفسها.
يُعرف أيضاً بـ
- eval framework
- evaluation suite
- إطار التقييم
- حزمة الاختبارات
