jailbreak_attempt_detector

Detects potential LLM jailbreak attempts by analyzing user input against NIST AI Risk Management Framework adversarial patterns. Designed for persona risk assessment, this tool evaluates text for common jailbreak techniques such as prompt injection, role-playing, or obfuscation. Inputs include th...

SERVERMcp Knowledge SOURCEhttps://mcp.gapup.io
Low RISK CLASS
Category Read
Parameters 41 required
Recommended Allowedsee the rule below
Registry record Grade F, identity unverified Pull the record →

This record as markdown: /tools/io-github-getgapup-mcp-knowledge/jailbreak-attempt-detector.md

What jailbreak_attempt_detector does on Mcp Knowledge

AI agents call jailbreak_attempt_detector to retrieve information from Mcp Knowledge without modifying anything. It is typically the context-gathering step in research, monitoring, and reporting workflows, before the agent takes action elsewhere.

ParameterTypeRequiredDescription
async boolean If true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client ti
context string Optional conversation context for better pattern matching
message string Yes User input text to analyze for jailbreak attempts
threshold number Confidence threshold for flagging attempts

Parameters from the server's own tool schema.

Why jailbreak_attempt_detector is rated Low

This tool analyzes and evaluates text input, returning a risk assessment. It is purely a read/query operation — no data is created, modified, deleted, or executed. The medium severity reflects that misuse or misconfiguration could lead to false positives/negatives affecting moderation decisions, but the tool itself has no destructive or write side effects.

From the tool's definition Detects potential LLM jailbreak attempts by analyzing user input... returning a risk assessment with confidence scores and pattern matches

Questions about jailbreak_attempt_detector

What does the jailbreak_attempt_detector tool do? +

Detects potential LLM jailbreak attempts by analyzing user input against NIST AI Risk Management Framework adversarial patterns. Designed for persona risk assessment, this tool evaluates text for common jailbreak techniques such as prompt injection, role-playing, or obfuscation. Inputs include the user message and optional context, returning a risk assessment with confidence scores and pattern matches. Ideal for real-time moderation in chat applications or API gateways. It is categorised as a Read tool in the Mcp Knowledge MCP Server, which means it retrieves data without modifying state.

What parameters does jailbreak_attempt_detector accept? +

jailbreak_attempt_detector accepts 4 parameters: async, context, message, threshold. Required: message. The full parameter table on this page comes from the server's own tool schema.

How do I enforce a policy on jailbreak_attempt_detector? +

Register the Mcp Knowledge MCP server in PolicyLayer and add a rule for jailbreak_attempt_detector: allow, deny, rate-limit, or require approval. Point your MCP client at the PolicyLayer proxy URL and the rule is enforced on every call, before it reaches Mcp Knowledge. Nothing to install.

What risk level is jailbreak_attempt_detector? +

jailbreak_attempt_detector is a Read tool with low risk. Read-only tools are generally safe to allow by default.

Can I rate-limit jailbreak_attempt_detector? +

Yes. Add a rate_limit block to the jailbreak_attempt_detector rule in your PolicyLayer policy. For example, setting max: 10 and window: 60 limits the tool to 10 calls per minute. Rate limits are tracked per agent session and reset automatically.

How do I block jailbreak_attempt_detector completely? +

Set action: deny in the PolicyLayer policy for jailbreak_attempt_detector. The AI agent will receive a policy violation error and cannot call the tool. You can also include a reason field to explain why the tool is blocked.

What MCP server provides jailbreak_attempt_detector? +

jailbreak_attempt_detector is provided by the Mcp Knowledge MCP server (https://mcp.gapup.io). PolicyLayer sits as a proxy in front of this server to enforce policies before tool calls reach the server.

More on Mcp Knowledge, and thousands of servers like it.

Across the catalogue

// THE MCP REGISTRY

PolicyLayer tracks 44,603 MCP servers and 515,000+ tools.

Every server has a live record: who publishes it, whether it answers without auth, its risk grade, every tool classified, the recommended policy. This page is one line of Mcp Knowledge's. Pull the full record:

Teams ship this data inside their own products. See what a licence covers →

// GET IN TOUCH

Have a question or want to learn more? Send us a message.

Message sent.

We'll get back to you soon.