English
Training data generated by a model rather than collected from the world, used to cover gaps, to avoid privacy constraints, or where real examples are scarce.
Synthetic data is now standard practice rather than a last resort, and it is particularly relevant for languages and dialects where natural corpora are thin, including much of Arabic. The known risk is compounding error: a model trained on another model's output inherits its mistakes and its blind spots, with no external signal correcting them, which degrades quality across generations if unchecked.
Synthetic data filled the gap in Gulf dialect coverage.
العربية
بياناتُ تدريبٍ يولّدها نموذجٌ بدل أن تُجمع من العالم، تُستعمل لسد الثغرات، أو لتفادي قيود الخصوصية، أو حيث تندر الأمثلة الحقيقية.
صارت البيانات التخليقية ممارسةً معيارية لا ملاذاً أخيراً، وهي وثيقة الصلة خاصةً باللغات واللهجات التي تشحّ فيها المدونات الطبيعية، ومنها كثيرٌ من العربية. والخطر المعروف تراكمُ الخطأ: فالنموذج المدرَّب على مخرجات نموذجٍ آخر يرث أخطاءه ونقاطه العمياء، بلا إشارةٍ خارجية تصححها، وهو ما يُدهور الجودة عبر الأجيال إن تُرك بلا ضبط.
سدّت البيانات التخليقية الثغرة في تغطية اللهجة الخليجية.
يُعرف أيضاً بـ
- generated data
- model-generated data
- البيانات المولّدة
- البيانات الاصطناعية
