Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The MCP server defines 20 tools across GitHub, PDF, sample, and security domains. While all tools have names and descriptions present, the quality is inconsistent. Most descriptions are generic and lack actionable context for LLM selection. Input schemas are present but parameters frequently lack detailed type information beyond the base type. Security-related tools return verbose responses with verbose metadata. Most critically, error handling is not documented, recovery guidance is absent, and output schemas are not specified. The tools follow basic naming conventions (verb_noun) but lack the depth and specificity expected in production-grade MCP servers. The sample and security tools in particular show signs of copy-pasting and incomplete implementation.
No output schemas documented. Tools return complex nested objects but LLMs have no specification of what fields to expect, forcing them to infer structure from examples or trial-and-error. This violates pattern:response-shaper and pattern:tool.
No error handling or recovery guidance in any tool description. When a tool fails (e.g., PDF parse error, GitHub rate limit hit, invalid domain), LLMs have no guidance on whether to retry, which tool to call next, or how to proceed. Violates pattern:recovery-guide.
Document output schemas for all tools. Specify the exact fields, types, and nesting expected in responses. For example, get_github_repository_info should specify: {owner: string, repo: string, stars: integer, description: string, url: string, ...}. This enables LLMs to plan downstream tool calls and extract the right data.
Add error handling and recovery guidance to every tool description. Pattern: 'On error (e.g., GitHub rate limit exceeded), the LLM should wait 60 seconds and retry. For missing resources, call [related discovery tool].' Use the pattern:recovery-guide pattern to structure this.
Replace string-based enums in descriptions with actual enum constraints in JSON Schema. For security_comprehensive_scan, define scan_profile as type:string, enum:["quick", "standard", "comprehensive", "compliance"], not just description text.
Add pagination and result-limiting parameters to tools returning lists. E.g., get_github_repository_branches should accept optional limit (default 20, max 100) and offset. Document the default and max in the description. Cap security scan results at 50 findings by default.
Provide deployment guidance on expected response times and timeouts. E.g., 'PDF extraction may take 10+ seconds for large files. Set HTTP timeout to 30s. If timeout occurs, try extracting specific page ranges instead of the full document.'
Distinguish similar tools via description. Clarify when to use pdf_search_text vs pdf_extract_text (search is faster for finding a phrase; extract returns all text in order). Clarify security_quick_scan vs security_comprehensive_scan (quick: <5s, returns top 5 findings; comprehensive: 30s+, exhaustive scan).
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Sample tools (sample_countwords, sample_combineanimals, sample_hello_world, add) are trivial and appear to be placeholders. They dilute the utility of the server and suggest incomplete implementation. Consider removing or dedicating them to a separate dev/test server.
Security tools lack specificity about what 'safe, non-invasive check' means and what data they expose. Descriptions do not clarify whether SSL/DNS/header analysis will leak internal IP info, DNS resolution details, or other sensitive data. LLMs cannot reason about data privacy implications. Violates pattern:security.
Parameters lack detailed constraints. E.g., 'scan_profile' in security_comprehensive_scan lists enums in description ('quick', 'standard', 'comprehensive', 'compliance') but schema does not enforce as enum. Page ranges in PDF tools use string format like '1-5' or '3,5,7' but lack format specification or validation hints. Violates pattern:constrained-input.
No pagination or result-limiting strategy documented. Large GitHub repo queries, PDF text extraction from 1000-page documents, and security scan results could return massive datasets that blow context windows. No mention of page limits, offsets, or result caps. Violates pattern:paginated-result.
Tool descriptions are generic and lack context about when to use them vs. similar tools. E.g., pdf_search_text vs. pdf_extract_text, when is searching better? What is the performance tradeoff? Is searching case-insensitive by default? LLMs cannot distinguish or select intelligently without this guidance.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. While all security/PDF tools are read-only, the absence of explicit annotations means LLMs cannot detect this without reading descriptions. Tool annotations are part of current MCP spec and improve reliability. Violates protocol:tool-annotations.
all
Add tool annotations to the MCP schema. For all tools, add readOnlyHint: true. For destructive tools (if any added later), add destructiveHint: true. This improves agent safety and LLM decision-making.
Reduce sample tools or move to a separate development server. sample_countwords, sample_combineanimals, sample_hello_world, and add are trivial and dilute credibility. Either remove them or house them in a separate debug server.
Clarify security implications in security tool descriptions. State explicitly: 'This tool performs DNS lookups and TLS certificate analysis. It does NOT execute exploits and does NOT access the target server internals. All queries are logged and may be subject to rate limits.' This prevents misuse.
Add per-parameter min/max constraints for numeric and string parameters. E.g., max_results in pdf_search_text: min 1, max 100, default 10. Page numbers: min 1, max inferred from pdf_get_page_count. These constraints should appear in the description and JSON Schema.
Document idempotency guarantees. Are PDF extractions idempotent? GitHub queries? If repeated calls with the same parameters produce the same result, state this explicitly. If there are side effects (caching, rate limit consumption), document them.
Add natural-language identifiers where applicable. E.g., security tools could accept human-readable domain names (example.com) in addition to full URLs. This matches the chat data model and reduces agent lookup overhead.
Return chaining IDs in responses. If security_batch_scan returns results, ensure each result includes a target identifier and scan_id so downstream analysis or follow-up tools can reference it without extra lookups.