AI Jailbreak & Roleplay-Persona

Detects overt jailbreak / roleplay-persona framings in AI/LLM input (OWASP LLM01): persona-override framings (DAN, developer mode), restriction-removal roleplay, and "for research/educational purposes" pretexts. Curated from public jailbreak taxonomies; no novel jailbreaks authored.

Type
keyword_list
Confidence
low
Confidence justification
Low by design. The library classifies this as trainable because regex/keyword detection collapses under paraphrase and multi-turn (Crescendo) attacks. Two tiers: a 65 phrase-only seed floor, and a 75 tier when the regex co-occurs with a persona keyword (Evidence_personas). Both remain low-confidence overall: this seed catches only overt persona framings and is explicitly necessary-not-sufficient; pair with a trainable classifier.
Jurisdictions
global
Regulations
OWASP LLM Top 10 2025, NIST AI RMF GenAI Profile
Frameworks
ISO 27001
Data categories
emerging, security
Risk rating
7

Pattern

(?i)\b(?:DAN|do\s+anything\s+now|developer\s+mode|jailbreak|for\s+(?:research|educational)\s+purposes\s+only|pretend\s+you\s+(?:are|have\s+no)|roleplay\s+as|act\s+as\s+an\s+unrestricted)\b

Corroborative evidence keywords

persona, roleplay, restrictions, mode, [object Object], artificial intelligence, [object Object], large language model, Copilot, chatbot, assistant, agent, prompt, system prompt, tool call, completion, model

Proximity: 300 characters

Should match

Should not match

Known false positives

Collections