English
Knowledge distillation trains a smaller student model to mimic the outputs of a larger, more capable teacher model, transferring much of its behavior at a fraction of the size and cost. This produces compact models that are cheaper and faster to run while keeping most of the original quality.
Knowledge distillation trains a small model to reproduce the behaviour of a large one, transferring capability into a cheaper package. The student learns from the teacher's full output distribution, not merely its final answers, which carries more information than the original labels did. It is how most small, fast production models are actually made, and it is contested legally when the teacher belongs to someone else.
The 3-billion-parameter model was distilled from a much larger one.
العربية
التقطير المعرفي يدرّب نموذجاً تلميذاً أصغر ليحاكي مخرجات نموذج معلّم أكبر وأكثر قدرة، فينقل إليه جزءاً كبيراً من سلوكه بحجم وتكلفة أقل بكثير. يُنتج هذا نماذج مدمجة أرخص وأسرع في التشغيل مع الحفاظ على معظم الجودة الأصلية.
يدرّب التقطير المعرفي نموذجاً صغيراً على استنساخ سلوك نموذجٍ كبير، فينقل القدرة إلى وعاءٍ أرخص. ويتعلم الطالب من توزيع مخرجات المعلّم كاملاً لا من إجاباته النهائية فحسب، وفي ذلك معلوماتٌ أكثر مما حملته التسميات الأصلية. وهكذا تُصنع فعلاً معظم نماذج الإنتاج الصغيرة السريعة، وهو محل نزاعٍ قانوني حين يكون المعلّم ملكاً لغير المُقطِّر.
قُطِّر النموذج ذو الثلاثة مليارات معلمة من نموذجٍ أكبر منه بكثير.
يُعرف أيضاً بـ
- model distillation
- student model
- تقطير النموذج
- الديستيليشن
يُكتب أيضاً هكذا، ولا نوصي به
- تقطير المعرفة
- التقطير النموذجي
