MCP server for ClawGuard Shield — scan AI agent inputs for prompt injection threats
ClawGuard Shield demonstrates solid definition quality with well-named tools following the verb_noun pattern, comprehensive descriptions (195-286 chars average), and clear input parameter documentation. All 5 tools have explicit schemas with typed parameters and descriptions. However, output schemas lack formal documentation, the code infers responses but does not document expected return fields. Error handling is present but could be more actionable (e.g., distinguishing retryable vs. fatal errors). The tools follow good separation of concerns and accept natural identifiers (source parameter in scan_text/scan_batch). Missing are: enum constraints on the 'severity' field that tools return, pagination support (not applicable here, but batch tool caps at 10 items without documenting why), and explicit permission/scope declarations. Overall tooling is production-grade; polish needed on output schema documentation and error guidance.
List all available ClawGuard detection patterns. Returns all 42+ security detection patterns organized by category: - prompt_injection: Override attempts, system tag spoofing - jailbreak: DAN, roleplay, hypothetical bypasses - data_exfiltration: Markdown image leaks, URL injection - social_engineering: Authority claims, credential phishing Each pattern includes its name, severity level, and description. No API key required.
Get API usage statistics for your ClawGuard Shield account. Shows your current tier (free/pro/enterprise), daily request limits, today's usage count, remaining quota, and rate limit status. Requires a valid API key.
Check if the ClawGuard Shield API is healthy and responding. No API key required. Returns the service status, API version, number of active detection patterns, and response time. Use this to verify connectivity before running scans.
Scan multiple texts for security threats. Scans each text individually and returns all results. Useful for checking multiple user inputs, chat messages, or document sections in one call.
Scan text for prompt injection and security threats. Analyzes the provided text using ClawGuard Shield's 42+ detection patterns to identify prompt injection attacks, jailbreak attempts, data exfiltration, social engineering, and other AI security threats. Returns a scan result with: - is_clean: whether the text is safe - risk_score: threat level from 0 (safe) to 10 (critical) - severity: NONE, LOW, MEDIUM, HIGH, or CRITICAL - findings: list of detected threats with pattern names and descriptions - scan_id: unique identifier for this scan
Output schemas are not formally documented. Tools return rich dictionaries (is_clean, risk_score, severity, findings, scan_id) but the server.py code does not declare the structure in a schema. LLMs cannot plan downstream operations or extract fields reliably without seeing the exact response shape.
Error handling does not categorize errors as retryable vs. fatal. scan_text and scan_batch catch httpx.HTTPStatusError and httpx.ConnectError but return generic error dicts without guidance. An LLM sees {'error': 'Cannot connect...'} and does not know if it should retry, ask the user, or give up.
Severity field (NONE, LOW, MEDIUM, HIGH, CRITICAL) in scan_text/scan_batch responses is not declared as an enum constraint. LLMs can infer valid values from the description but cannot validate input. If the API returns an unexpected severity, the LLM has no schema-level protection.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 59 | - | v1 |
scan_batch documents a '10 items max' constraint in the description but does not enforce it at the parameter level (no maxItems in schema). The code checks len(texts) > 10 at runtime and returns an error array, but LLMs cannot see the constraint in the schema and may attempt larger batches, wasting tokens.
No permission or scope declarations. Tools like get_usage require an API key, but the server does not advertise permission scopes (e.g., 'read:usage', 'read:scans'). This prevents least-privilege agent configuration and unclear audit trails.