English
How Arabic text is cut into tokens, and why the same meaning costs more tokens in Arabic than in English, making Arabic requests measurably more expensive and effectively shrinking the context window.
Tokenizers are trained on corpora that are mostly English, so they learn efficient pieces for English and inefficient ones for everything else. Arabic words, already dense with morphology, get cut into more fragments, commonly two to three times as many tokens as an English sentence of equivalent meaning. Two consequences follow directly: the same request costs more, and a 100,000-token context window holds noticeably less Arabic than English. This is a quiet, structural disadvantage rather than a quality judgement about the language.
Arabic tokenization inflated the prompt from 900 to 2,400 tokens.
العربية
كيفية تقطيع النص العربي إلى توكنات، ولماذا يكلّف المعنى نفسه توكناتٍ أكثر بالعربية منه بالإنجليزية، فتصير الطلبات العربية أعلى كلفةً قياساً وتتقلص نافذة السياق فعلياً.
تُدرَّب المُقطِّعات على مدوناتٍ إنجليزيةٍ في معظمها، فتتعلم قطعاً كفؤةً للإنجليزية وأخرى غير كفؤةٍ لسواها. والكلمة العربية، وهي أصلاً كثيفة صرفياً، تُقطَّع إلى شظايا أكثر، بما يبلغ عادةً ضعفي أو ثلاثة أضعاف توكنات جملةٍ إنجليزيةٍ تؤدي المعنى نفسه. وتتبع ذلك نتيجتان مباشرتان: الطلب نفسه أعلى كلفة، ونافذةُ سياقٍ من مئة ألف توكن تسع من العربية أقل بوضوحٍ مما تسع من الإنجليزية. وهذا ضعفٌ بنيويٌّ صامت لا حكمٌ على اللغة.
ضخّم تقطيع النص العربي التوجيهَ من 900 توكن إلى 2400.
يُعرف أيضاً بـ
- tokenizer fertility
- subword segmentation
- خصوبة المقطّع
- التقطيع إلى وحدات فرعية
