English
An architecture that holds many specialised sub-networks but activates only a few per token, so total capacity grows without proportionally increasing the cost of each request.
A router picks which experts handle each token, so a model with hundreds of billions of total parameters may use only tens of billions on any given pass. This decouples capacity from inference cost, which is why most frontier models are now sparse rather than dense. The cost is elsewhere: all experts must still be held in memory, and the routing itself can be unstable during training.
The mixture-of-experts model activates 2 of 64 experts per token.
العربية
معماريةٌ تضم شبكاتٍ فرعية متخصصة كثيرة لكنها لا تُفعِّل إلا قليلاً منها لكل توكن، فتنمو السعة الكلية دون أن ترتفع كلفة كل طلبٍ بالنسبة نفسها.
يختار موجّهٌ أيَّ الخبراء يعالج كل توكن، فقد يستعمل نموذجٌ بمئات المليارات من المعاملات الكلية عشراتِ المليارات فقط في أي تمريرة. وهذا يفصل السعة عن كلفة الاستدلال، ولهذا صارت معظم النماذج المتقدمة متفرقةً لا كثيفة. والكلفة في موضعٍ آخر: إذ يجب إبقاء الخبراء جميعاً في الذاكرة، وقد يكون التوجيه نفسه غير مستقرٍّ أثناء التدريب.
يُفعِّل نموذج خليط الخبراء خبيرين من أربعةٍ وستين لكل توكن.
يُعرف أيضاً بـ
- MoE
- sparse model
- مزيج الخبراء
- النموذج المتفرق
