MCP security scanner and CI gate. Test, secure, and monitor MCP servers with attack simulation, schema drift detection, health scoring, and SARIF before agents depend on them.
mcp-observatory exhibits solid tool definition quality with 13 well-named, verb-action-oriented tools covering a coherent scanning/monitoring domain. All tools have descriptions ranging 100-300+ characters with clear use cases. However, schema visibility is limited, while the provided sample shows rich input schemas (command, args, deep, security parameters with descriptions), only 3 tools (scan, check_server, score_server) have complete visible schemas in the evaluation materials. The remaining 10 tools (record, replay, verify, watch, lock_verify, get_history, ci_report, and others) lack explicit schema definitions in the source snippets, forcing conservative scoring despite strong naming and description evidence. Error handling is present (errorReporting=true) and output documentation is inferred from tool descriptions, but full response schemas are not visible. Composition is excellent, each tool solves one problem; no conflation. Tool chains work well (check_server → diff_runs → get_last_run). All parameter names match expected LLM-friendly patterns (command, args, deep, security, cassette, etc.). No secrets exposed as parameters.
Use this to test a specific MCP server before installing or after updating it. Launches the server by command, checks all capabilities, and saves a run artifact for future comparison. Example: check_server({ command: 'npx -y @modelcontextprotocol/server-everything' }). Use deep=true to invoke tools, security=true to analyze schemas for vulnerabilities.
Generates a CI-friendly report (JSON or SARIF) from a check run artifact, suitable for GitHub code scanning, GitLab SAST, or other CI systems.
Compares two previously recorded check runs to detect schema changes, missing tools, parameter mutations, or response shape differences. Use after updating a server to see exactly what changed.
Retrieves historical check runs for a given server, allowing trend analysis of health scores and check status over time.
Retrieves the most recent recorded check run for a given server command. Useful for comparing current server state against the last baseline.
Verifies an mcp-lock.json lockfile by checking that all pinned servers are still available, accessible, and have not been modified since they were locked.
10 of 13 tools lack visible input schemas in source. Only scan, check_server, and score_server have documented input parameter schemas with types and descriptions. record, replay, verify, watch, lock_verify, get_history, diff_runs, get_last_run, suggest_servers, and ci_report are inferred from package.json or markdown but have no explicit TypeScript schema definitions visible. This limits confidence in parameter validation and LLM invocation accuracy.
Output schemas not documented. While descriptions hint at outputs (e.g., score_server returns '0-100 score with A-F grade'), no formal response schemas are visible. LLMs cannot predict output structure or chain downstream operations reliably. Missing response documentation violates pattern:response-shaper.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
Launches an MCP server and records all requests and responses into a cassette file (VCR-style). Useful for creating regression test baselines or capturing real-world traffic for replay.
Replays a cassette (recorded request-response file) against a live or test MCP server. Compares the current responses against the recorded baseline to detect regressions.
Scans a local MCP server for security issues, schema problems, and capability gaps. Launches the server by command, checks protocol compliance, invokes tools, and analyzes security aspects like parameter validation and dangerous hints.
Use this to get a quick health grade for an MCP server. Runs all checks (capabilities, tool invocation, security) and returns a 0-100 score with A-F grade and detailed breakdown across protocol compliance, schema quality, security, reliability, and performance.
Use this when setting up a project or wondering what MCP servers to add. Scans the working directory for languages, frameworks, databases, and cloud providers, lists currently configured servers, and cross-references the MCP registry to recommend servers you're missing.
Use this after updating a server to confirm nothing broke. Connects to the live server, sends the same requests from a recorded cassette, and compares responses. Reports exactly what changed — added tools, removed parameters, different response shapes.
Monitors a server command or HTTP endpoint for availability and schema stability. Reports uptime, response times, and detects breaking changes in real time.
Enum constraints and format validation not visible. Parameters like 'format' in ci_report ('json' or 'sarif') should declare enum constraint; 'intervalMs' and 'durationMs' in watch should document numeric bounds and defaults. Current descriptions contain natural language hints instead of machine-parseable constraints, forcing LLMs to infer valid options.
Error recovery guidance incomplete. While errorReporting=true indicates error handling exists, descriptions do not provide actionable recovery hints (e.g., 'If server launch fails, check command syntax' or 'If diff shows no changes, servers may be identical'). Agents cannot self-correct without explicit recovery paths.
No pagination or result limit declarations visible. Tools like get_history accept 'limit' but no bounds are documented (e.g., 'limit: max 1000'). Tools returning lists (scan, suggest_servers) should declare caps and pagination support explicitly. This risks context window exhaustion.