English
Empirical relationships showing model performance improves predictably as parameters, data and compute grow together, the reasoning that justified spending billions on larger training runs.
Scaling laws made model building an investment case rather than a gamble: given a compute budget, they predict roughly what loss you will reach. They also specify the ratio between data and parameters, and the finding that most models were badly under-trained on data reshaped the field. They are empirical regularities, not physical laws, they describe what has held so far, and say nothing about where they stop.
Scaling laws suggested the model needed four times more data.
العربية
علاقاتٌ تجريبية تُظهر أن أداء النموذج يتحسن تحسناً قابلاً للتنبؤ كلما نمت المعاملات والبيانات والحوسبة معاً، وهي الحجة التي بُرِّر بها إنفاق المليارات على جولات تدريبٍ أكبر.
جعلت قوانين التوسع بناءَ النماذج حجةً استثمارية لا مقامرة: إذ تتنبأ، عند ميزانية حوسبةٍ معطاة، بالخسارة التي ستبلغها تقريباً. وهي تحدد كذلك النسبة بين البيانات والمعاملات، وقد أعاد اكتشافُ أن معظم النماذج كانت ناقصة التدريب على البيانات تشكيلَ الحقل. وهي انتظاماتٌ تجريبية لا قوانين فيزيائية، تصف ما صحّ حتى الآن، ولا تقول شيئاً عن موضع توقفه.
أشارت قوانين التوسع إلى أن النموذج يحتاج بياناتٍ أكثر بأربعة أضعاف.
Also known as
- compute-optimal scaling
- Chinchilla scaling
- قوانين القياس
- علاقات التوسع
First use
2020 · Kaplan et al., then revised by DeepMind's Chinchilla work in 2022.
Source