ugc_moderation_classifier
Multi-language UGC content moderation for marketplaces, social platforms and comment systems. Detects policy violations in text content across 9 policies and 12 languages without external API calls. Policies checked: • hate — hate speech, slurs, dehumanization (50+ terms × 12 languages) • sexual ...
This record as markdown: /tools/io-github-getgapup-mcp-knowledge/ugc-moderation-classifier.md
What ugc_moderation_classifier does on Mcp Knowledge
AI agents call ugc_moderation_classifier to retrieve information from Mcp Knowledge without modifying anything. It is typically the context-gathering step in research, monitoring, and reporting workflows, before the agent takes action elsewhere.
| Parameter | Type | Required | Description |
|---|---|---|---|
lang | string | — | Language override. If omitted, language is auto-detected. |
async | boolean | — | If true, returns a job_id immediately (<200ms) instead of waiting for the result. Poll the result with job_result(job_id). Use for slow tools to avoid client ti |
content | string | Yes | Text content to moderate (comment, review, post, chat message). |
policies | array | — | Policies to check. Default: all 9 policies. |
content_type | string | — | Type of content. Affects recommended_action heuristic. Default: comment. |
Parameters from the server's own tool schema.
Why ugc_moderation_classifier is rated Low
This tool analyzes/classifies text content for policy violations and returns a result. It reads and evaluates input text but does not modify data, execute commands, delete anything, or involve financial transactions. The 'without external API calls' note confirms it is a local analysis operation. Misuse risk is low — the worst outcome is a misclassification of content, not a destructive or financial action.
From the tool's definition Detects policy violations in text content across 9 policies and 12 languages without external API calls
Risk signalsAccepts raw HTML/template content (content)
Attacks that exploit this kind of access
The rule that runs ugc_moderation_classifier safely
PolicyLayer is an MCP gateway: it sits between your AI agents and Mcp Knowledge, and checks every tool call against a rule you set before the call runs. Nothing changes on the server itself. For ugc_moderation_classifier, this is the rule to start with:
ugc_moderation_classifier is read-only, so it stays allowed. Everything else on the server is denied unless you say otherwise.
The button opens the PolicyLayer dashboard: create your workspace, connect Mcp Knowledge, apply this rule, and every ugc_moderation_classifier call is checked against it from then on.
Questions about ugc_moderation_classifier
Multi-language UGC content moderation for marketplaces, social platforms and comment systems. Detects policy violations in text content across 9 policies and 12 languages without external API calls. Policies checked: • hate — hate speech, slurs, dehumanization (50+ terms × 12 languages) • sexual — explicit sexual content, pornography references, nudity solicitation • violence — threats, weapon references, graphic violence • self_harm — suicidal ideation, self-injury, eating disorder promotion • harassment — doxxing, stalking, cyberbullying, blackmail • scam — phishing, investment fraud, romance scam, lottery fraud • spam — bots, keyword stuffing, excessive caps, emoji storms, suspicious URLs • copyright — piracy, leaked content, serial keys, streaming fraud • minor_safety — grooming signals, CSAM references, minor + adult content combos Languages: en / fr / de / es / it / pt / nl / zh / ja / ko / ar / ru (auto-detected) Output includes severity (low/medium/high/severe), confidence (0-100), matched patterns, excerpt, recommended action, age appropriateness (adult/teen/child), and signals. No API key required. Stateless — no content is stored or logged. It is categorised as a Read tool in the Mcp Knowledge MCP Server, which means it retrieves data without modifying state.
ugc_moderation_classifier accepts 5 parameters: lang, async, content, policies, content_type. Required: content. The full parameter table on this page comes from the server's own tool schema.
Register the Mcp Knowledge MCP server in PolicyLayer and add a rule for ugc_moderation_classifier: allow, deny, rate-limit, or require approval. Point your MCP client at the PolicyLayer proxy URL and the rule is enforced on every call, before it reaches Mcp Knowledge. Nothing to install.
ugc_moderation_classifier is a Read tool with low risk. Read-only tools are generally safe to allow by default.
Yes. Add a rate_limit block to the ugc_moderation_classifier rule in your PolicyLayer policy. For example, setting max: 10 and window: 60 limits the tool to 10 calls per minute. Rate limits are tracked per agent session and reset automatically.
Set action: deny in the PolicyLayer policy for ugc_moderation_classifier. The AI agent will receive a policy violation error and cannot call the tool. You can also include a reason field to explain why the tool is blocked.
ugc_moderation_classifier is provided by the Mcp Knowledge MCP server (https://mcp.gapup.io). PolicyLayer sits as a proxy in front of this server to enforce policies before tool calls reach the server.
More on Mcp Knowledge, and thousands of servers like it.
This server
Across the catalogue