كل المصطلحات

Benchmark

التقييم

اختبار معياري

النطق ikhtibār miʿyārīمتوسط

English

A benchmark is a standardized test or dataset used to measure and compare how well different AI models perform on a specific kind of task, such as math problems or reading comprehension. Published benchmark scores let researchers and buyers compare models on a common basis.

A benchmark is a fixed test set used to compare models on a common footing. Its value depends entirely on the test staying unseen, and public benchmarks leak into training corpora over time, so a rising score can mean a better model or a more contaminated test. Benchmarks also measure what is easy to score, which is why they systematically under-represent qualities like judgement, tone, and correctness in low-resource languages.

It leads the benchmark but performs worse on our own evaluations.

العربية

الاختبار المعياري مجموعة اختبارات أو بيانات موحّدة تُستخدم لقياس أداء نماذج الذكاء الاصطناعي المختلفة ومقارنتها في نوع محدد من المهام، مثل المسائل الرياضية أو فهم النصوص. تتيح نتائج الاختبارات المعيارية المنشورة للباحثين والمشترين مقارنة النماذج على أساس مشترك.

الاختبار المعياري مجموعة اختبارٍ ثابتة تُستعمل لمقارنة النماذج على أرضيةٍ واحدة. وقيمته متوقفةٌ كلياً على بقاء الاختبار غير مرئي، والاختبارات العامة تتسرب إلى مدونات التدريب مع الوقت، فالنتيجة الصاعدة قد تعني نموذجاً أفضل أو اختباراً أشد تلوثاً. كما تقيس الاختبارات ما يسهل تسجيله، ولهذا تُمثِّل تمثيلاً ناقصاً ومنهجياً صفاتٍ كالحكم والنبرة والصواب في اللغات شحيحة الموارد.

يتصدر الاختبار المعياري لكنه أسوأ أداءً في تقييماتنا الخاصة.

يُعرف أيضاً بـ

  • benchmarks
  • benchmark test
  • معيار قياس
  • اختبارات معيارية

يُكتب أيضاً هكذا، ولا نوصي به

  • المعيار المرجعي
  • البنشمارك

مصطلحات ذات صلة