English
A prompt crafted to make a model produce output its training was meant to prevent, usually by reframing the request as fiction, translation, or a hypothetical.
Refusal behaviour is trained, not enforced, so it is a tendency rather than a boundary. Jailbreaks exploit that: the same request that is refused when asked plainly may succeed when embedded in a story or a role. Every published defence has been circumvented, which is why serious deployments do not rely on the model's own refusal as their only control.
The jailbreak worked by asking for the answer as a poem.
العربية
توجيهٌ مصنوعٌ ليجعل النموذج ينتج مخرجاً قُصد بتدريبه منعُه، وذلك غالباً بإعادة تأطير الطلب بوصفه خيالاً أو ترجمةً أو افتراضاً.
سلوك الرفض مُدرَّبٌ لا مُنفَّذ، فهو ميلٌ لا حد. وكسر القيود يستغل ذلك: إذ إن الطلب الذي يُرفض حين يُطرح صريحاً قد ينجح حين يُدرَج في قصةٍ أو دور. وقد جرى الالتفاف على كل دفاعٍ نُشر، ولهذا لا تعتمد عمليات النشر الجادة على رفض النموذج نفسه ضابطاً وحيداً.
نجح كسر القيود بطلب الجواب في صورة قصيدة.
يُعرف أيضاً بـ
- prompt attack
- تجاوز القيود
- الاختراق بالتوجيه
