English
Deliberately attacking your own system to find failures before someone else does, for AI, probing a model for harmful output, leaks, and jailbreaks under adversarial conditions.
Ordinary testing asks whether a system does what it should; red teaming asks what it can be made to do. For models this means adversarial prompting at scale, often mixing automated attack generation with human creativity, because the failures that matter are usually the ones nobody thought to specify. It is increasingly a regulatory expectation rather than a voluntary practice.
External red teaming preceded the model's release.
العربية
مهاجمة نظامك عمداً لاكتشاف إخفاقاته قبل أن يكتشفها غيرك، وفي الذكاء الاصطناعي: سبر النموذج بحثاً عن مخرجاتٍ ضارة وتسريباتٍ وكسرٍ للقيود في ظروفٍ عدائية.
يسأل الاختبار المعتاد أيفعل النظامُ ما ينبغي، أما الفريق الأحمر فيسأل ما الذي يمكن أن يُجبَر على فعله. ويعني هذا مع النماذج توجيهاً عدائياً واسع النطاق، يمزج غالباً توليد الهجمات آلياً بإبداع البشر، لأن الإخفاقات ذات الشأن هي عادةً ما لم يخطر لأحدٍ أن يحدده. وهو يصير توقعاً تنظيمياً أكثر منه ممارسةً طوعية.
سبق إصدارَ النموذج اختبارٌ عدائيٌّ من فريقٍ أحمر خارجي.
يُعرف أيضاً بـ
- adversarial testing
- الاختبار العدائي
- فرق المهاجمة
