English
A transformer is a neural network architecture that processes sequences using self-attention rather than recurrence. It underlies nearly all modern large language models.
A transformer processes a whole sequence at once and lets every position attend to every other, instead of walking through the sequence step by step as recurrent networks did. That single change is what made training on very large corpora practical: the work parallelises across a sequence rather than being forced into strict order. Nearly every model discussed in AI news since 2018, language, image, audio, or multimodal, is a transformer or a descendant of one.
The model is a 12-layer transformer trained on 300 billion tokens.
العربية
المحوّل معمارية شبكة عصبية تعالج السلاسل النصية عبر آلية الانتباه الذاتي بدلاً من التكرار. تقوم عليها كل النماذج اللغوية الكبيرة الحديثة تقريباً.
يعالج المحوِّل السلسلة كاملةً دفعةً واحدة، ويسمح لكل موضع فيها بالانتباه إلى كل موضع آخر، بدلاً من المرور عليها خطوةً بخطوة كما كانت تفعل الشبكات التكرارية. هذا التغيير وحده هو ما جعل التدريب على مدونات ضخمة أمراً عملياً: إذ صار العمل قابلاً للتوزيع المتوازي بدل أن يُجبر على ترتيب صارم. وتكاد كل النماذج التي تتناولها أخبار الذكاء الاصطناعي منذ 2018، اللغوية أو الصورية أو الصوتية أو متعددة الوسائط، أن تكون محوِّلاً أو أحد أحفاده.
النموذج مُحوِّلٌ من اثنتي عشرة طبقة دُرِّب على 300 مليار توكن.
Also known as
- transformer architecture
- المحول
- معمارية المحول
Also seen as, but not recommended
- الترانسفورمر
- المُبدِّل
Why this form: «الترانسفورمر» is a transliteration that carries no meaning in Arabic, and «المُبدِّل» means a switch or swapper. «المُحوِّل» keeps the sense of transformation the architecture actually performs.
First use
2017 · Introduced in "Attention Is All You Need" (Vaswani et al., Google).
Source