Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
TailOpsMCP is a moderately functional MCP server with six Docker/package management tools. All tools have names starting with action verbs (manage_, get_, start_, stop_, inspect_) and include descriptions. Input schemas are present and properly typed with JSON Schema. However, there are significant gaps: tool descriptions lack depth and actionable context, parameter descriptions are minimal, output schemas are not documented, and error handling guidance is missing. The server follows basic MCP patterns but lacks the LLM-optimization and composition guidance that production-grade tools exhibit. No tool annotations (readOnlyHint, destructiveHint) are present despite several tools being destructive (manage_packages with install/update, start_container, stop_container). Descriptions average ~80 chars (below the 194-char baseline for A+ tools), and many parameters lack the detailed format/constraint documentation needed for robust LLM usage.
Output schemas not documented. Tools return data but LLMs have no visibility into the structure of responses (fields, types, nested objects). This forces LLMs to guess at field names and causes downstream lookup calls or parse failures.
Tool descriptions are generic and lack LLM-optimization guidance. E.g., 'Manage system packages: check updates, update all, or install specific package.' does not explain WHEN to call this vs alternatives, what prerequisites exist, or what the agent should do with the output. Should be 50-200 chars with clear action context.
Document output schemas for all tools. For get_container_list, specify: returns {containers: [{name, id, status, image, ports}], total_count}. For manage_packages, specify: returns {action, package, status, version, output}. This enables LLMs to parse and chain responses.
Expand tool descriptions to 80-150 chars with action context. E.g., 'manage_packages': 'Check for system package updates, install specific packages, or update all packages on the target system. Use before deploying to ensure security patches are applied. Returns status and version info.' Include dependency hints like 'Call get_container_list() first to see available containers before start_container().'
Add readOnlyHint annotations: apply to get_container_list, inspect_container, get_container_logs (safe to retry). Add destructiveHint to manage_packages (install/update), start_container, stop_container, these modify state and have side effects.
Enforce confirmation for destructive operations. Add dry_run parameter to manage_packages (currently missing). Make dry_run=true the default or require explicit dry_run=false to execute actual changes. Document in description: 'Set dry_run=false to execute; true simulates without modifying the system.'
Expand parameter descriptions with format/constraint details. E.g., 'target': 'Target system to query. Hostname, IP address, or SSH alias from config (default: "local"). E.g. "prod-server-1" or "192.168.1.10". Must be reachable via SSH/API.' For 'container', clarify: 'Container name or full/partial container ID (e.g., "my-app" or "abc123def456").' For 'action' in manage_packages, specify: 'Operation to perform: check (list updates), update (apply all available updates), or install (install specific package_name).'
Score history
Overall score trend
↑ 18 points across a rubric change (v1 → v2)
61/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
C
61
2026-07-28+
v2
2026-03-09
F
43
-
v1
Destructive operations (manage_packages with install/update, start_container, stop_container) lack confirmation/dry-run enforcement. dry_run parameter exists on container operations but manage_packages lacks it entirely. No tool annotations (destructiveHint) present to warn LLMs of irreversible consequences.
Parameter descriptions are minimal. E.g., 'target' is described as 'Target system to query (default: "local")' without explaining what 'system' means (hostname? IP? SSH config?), format, or how to discover available targets. Users/agents cannot infer valid values.
No error handling guidance or recovery hints. Code likely returns raw exceptions or HTTP error codes. LLMs receive no actionable feedback (e.g., 'Container not found. Try get_container_list() to see available containers.') and cannot self-correct.
No tool annotations present. Tools lack readOnlyHint, destructiveHint, or idempotentHint metadata. LLMs cannot determine which tools are safe to retry, which modify state, or which are read-only without parsing descriptions.
Parameter naming inconsistency: 'container' is used for container ID/name, but 'target' is used for system identifier. No guidance on whether 'container' accepts name or ID. Should clarify 'container_name_or_id' and expand 'target' docs to explain format (hostname, IP, etc.).
Add error handling and recovery guidance in code. When a container is not found, return: '{error: "Container not found: xyz", suggestion: "Available containers: [list]", recovery_tool: "get_container_list"}' instead of a raw 404. Similarly for manage_packages: if package is not found, suggest 'Try search_packages() or check package name spelling.'
Add idempotent operation hints. Document that get_* and inspect_* tools are idempotent (safe to call multiple times). For start_container and stop_container, clarify: 'Idempotent, starting an already-running container is a no-op; stopping an already-stopped container is a no-op.'
Reduce required parameters by adding smart defaults where safe. E.g., format parameter defaults to 'toon' (compact) for readability, good default. Keep target='local' default but document it clearly so remote operations are explicit.
Validate input early with clear error messages. E.g., if action is not in [check, update, install], return: 'Invalid action: "xyz". Must be one of: check, update, install.' instead of letting it propagate as a 500 error.
Document pagination and result limits in descriptions and code. If get_container_list can return 1000+ containers, cap output at 50 and return a next_cursor or limit parameter. Current docs don't mention limits, risking context window overflow.