English
Multimodal describes a model that can understand or generate more than one type of content, such as text, images, audio, or video, within a single system. This lets it, for example, answer a question about a photo or describe the content of an audio clip.
A multimodal model handles more than one kind of input or output, text and images, or text and audio and video together, inside a single system rather than by bolting separate models together. The gain is that relationships across modalities can be learned instead of engineered: the model can be asked about a chart, not merely handed a transcription of it.
The multimodal model reads the invoice image directly.
العربية
متعدد الوسائط يصف نموذجاً قادراً على فهم أو توليد أكثر من نوع واحد من المحتوى، مثل النص والصور والصوت والفيديو، ضمن نظام واحد. يتيح له هذا مثلاً الإجابة عن سؤال حول صورة أو وصف محتوى مقطع صوتي.
النموذج متعدد الوسائط يتعامل مع أكثر من نوعٍ من المدخلات أو المخرجات، نصاً وصوراً، أو نصاً وصوتاً وفيديو معاً، داخل نظامٍ واحد لا بربط نماذج منفصلة بعضها ببعض. والمكسب أن العلاقات العابرة للوسائط تُتعلَّم بدل أن تُهندَس: فيمكن أن يُسأل النموذج عن رسمٍ بياني لا أن يُسلَّم نصاً منقولاً عنه.
يقرأ النموذج متعدد الوسائط صورة الفاتورة مباشرةً.
Also known as
- multimodal AI
- multi-modal
- الملتيموديل
Also seen as, but not recommended
- متعدد الأنماط
- متعدد الوسائل
Why this form: «الأنماط» means patterns or modes in a different sense; «الوسائط» is the media/modalities being combined.
