English
Checks placed around a model that constrain what it may receive or emit, filters, validators, allowed-action lists, enforced outside the model rather than requested inside the prompt.
A guardrail is deterministic where a prompt instruction is merely persuasive. Asking a model not to reveal a system prompt is a request it may decline to honour under adversarial pressure; refusing to send the response when it contains the system prompt is a rule it cannot argue with. The distinction matters most for agents, where an unchecked action has effects that a bad sentence does not.
A guardrail blocks the agent from sending mail to addresses outside the domain.
العربية
فحوصٌ تُوضع حول النموذج تقيّد ما يجوز أن يتلقاه أو يُصدره، من مرشحاتٍ ومدققاتٍ وقوائم أفعالٍ مسموحة، تُطبَّق خارج النموذج لا تُطلَب داخل التوجيه.
حاجز الأمان حتميٌّ حيث تكون تعليمة التوجيه إقناعيةً فحسب. فأن تطلب من النموذج ألا يكشف توجيه النظام طلبٌ قد لا يستجيب له تحت ضغطٍ عدائي، أما أن ترفض إرسال الرد حين يتضمن توجيه النظام فقاعدةٌ لا يجادلها. وللتمييز أهميةٌ قصوى مع الوكلاء، إذ للفعل غير المفحوص آثارٌ ليست للجملة السيئة.
يمنع حاجز أمانٍ الوكيلَ من إرسال البريد إلى عناوين خارج النطاق.
Also known as
- safety rails
- output filters
- الضوابط
- المرشحات الأمنية
