Multi-layer security detection system for MCP tools with configurable models. Provides guardrail scanning using local rule-based detection, LearnableShield ML model, and configurable LLM-based adjudication.
MCP-Guard is a security-focused tool server, but has critical deficiencies in definition quality. Only 1 of 7 tools (scan) has a substantive description. Parameters lack descriptions entirely across most tools. Schemas are minimally documented. The server targets a specific domain (guardrail scanning) but fails to expose this domain through clear, LLM-friendly tool definitions. Tool naming is reasonable (verb_noun pattern), but parameter documentation and error handling guidance are nearly absent.
Disable a specific model for use in detection
Enable a specific model for use in detection
Get current active model and list of available models
List all available models and their status
Main guardrail scanning endpoint for detecting security issues in MCP tool descriptions through multi-layer detection (local rules, LearnableShield, and model-based analysis)
Switch the active model for detection
Test the current active model with sample text
6 of 7 tools have descriptions under 20 characters or generic placeholder text ('Get current active model...', 'Enable a model...', 'Disable a model...'). LLMs cannot distinguish when to use enable_model vs switch_model or decide between list_models and get_active_model.
Parameter descriptions are entirely missing for 6 of 7 tools. E.g., switch_model takes 'model_name' with no description of what model names are valid, what the parameter controls, or consequences of switching. test_model's 'text' param has description but no guidance on expected format or length.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 8 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 34 | - | v1 |
No output schema documentation visible. Tools like list_models, get_active_model, and scan do not document what fields LLMs should expect in responses. Without documented return structures, agents cannot extract IDs, status codes, or next-step parameters.
Unclear tool composition and overlap. 'get_active_model' and 'list_models' both return model information, distinction not documented. 'enable_model' and 'switch_model' have overlapping semantics (activate vs enable). LLMs will struggle to choose correctly.
'scan' is the primary tool but has no documented constraints on input size, timeout behavior, or what 'issues' field contains. Tool descriptions state it performs 'multi-layer security detection' but do not clarify what detection stages are possible, what confidence thresholds apply, or how to interpret the 'allowed' boolean.
No error handling guidance. Tools do not document what errors can occur, whether they are retryable, or what the LLM should do next. E.g., if switch_model fails because a model is not available, should the agent call list_models() first? No guidance provided.
No schema format specifications. test_model's 'text' parameter has a default in Chinese ('这是一个测试文本') but no documented max length, character restrictions, or format requirements. scan's 'tool_input_schema' param is typed as 'object' with no indication of JSON Schema structure expected.
Ambiguous parameter semantics for 'scan'. The 'tool_input_schema' parameter description says 'Input schema of the tool for exfiltration detection', unclear whether this is the schema OF the tool being scanned or schema FOR exfiltration detection logic. No guidance on what format is expected.
Mutation clarity lacking. write-risk tools (switch_model, enable_model, disable_model) have no documentation of side effects, reversibility, or retry safety. Agents cannot determine if retrying is safe or if they risk duplicate state changes.