كل المصطلحات

Quantization

البنية التحتية

التكميم

النطق at-takmīmمتقدم

English

Quantization reduces the numeric precision of a model's weights, for example from 16-bit to 4-bit numbers, to shrink its memory footprint and speed up inference. This usually causes a small drop in accuracy in exchange for running on cheaper or smaller hardware.

Quantization stores a model's weights at lower numeric precision, 8-bit or 4-bit instead of 16, so it needs less memory and runs faster. Quality degrades, but far less than the size reduction suggests, which is why quantized models are what actually run on laptops and phones. It is the main technique that puts a large model on hardware that could not otherwise hold it.

The 4-bit quantized build runs on a single consumer GPU.

العربية

التكميم يُقلّل دقة الأرقام المستخدمة في تمثيل أوزان النموذج، مثلاً من 16 بت إلى 4 بت، لتقليل حجمه في الذاكرة وتسريع الاستدلال. يسبب هذا عادة انخفاضاً طفيفاً في الدقة مقابل إمكانية التشغيل على عتاد أرخص أو أصغر.

يخزّن التكميم أوزان النموذج بدقةٍ عدديةٍ أقل، ثمانية بتات أو أربعة بدل ستة عشر، فيحتاج ذاكرةً أقل ويعمل أسرع. وتتراجع الجودة، لكن بأقل كثيراً مما يوحي به تقلّص الحجم، ولهذا كانت النماذج المُكمَّمة هي ما يعمل فعلاً على الحواسيب المحمولة والهواتف. وهو التقنية الرئيسة التي تضع نموذجاً كبيراً على عتادٍ ما كان ليسعه لولاها.

يعمل الإصدار المُكمَّم بأربعة بتات على وحدة معالجةٍ رسوميةٍ استهلاكيةٍ واحدة.

يُعرف أيضاً بـ

  • model quantization
  • quantized model
  • تكميم النموذج
  • الكوانتيزيشن
  • ضغط الأوزان

يُكتب أيضاً هكذا، ولا نوصي به

  • التكميّة
  • التقطيع الكمي

مصطلحات ذات صلة