Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This MCP server has moderate definition quality with some strengths but significant gaps. All 4 tools are explicitly defined in src/main.py with input schemas and descriptions. However, descriptions are brief and often lack actionable detail for LLM decision-making. Parameter descriptions are sparse, and output schemas are not documented. Tool naming follows verb_noun convention (submit_*, reset_*, update_*) which is good, but parameter relationships and error handling lack clarity. The audit_history and session state management create composition concerns (tools interdepend on global state rather than explicit IDs). Average across 4 tools: 42/100.
action parameter in update_rules is free-form string, not an enum. LLMs can hallucinate invalid actions like 'delete' or 'edit' instead of the documented 'add|remove|update|list'.
Tool descriptions are too brief (24-35 chars average). They lack WHEN to use each tool, WHAT happens on success/failure, and dependencies on other tools. This forces LLMs to reason backwards from behavior.
No output schemas documented for any tool. Callers and the LLM cannot predict what fields to expect. submit_draft returns a complex multi-line prompt; update_rules returns unknown structure. This violates the response-shaper pattern.
Convert action parameter in update_rules to an enum: action = enum(['add', 'remove', 'update', 'list']). This prevents LLM hallucination of invalid operations.
Expand tool descriptions to 50-150 characters each. Include WHEN to use (e.g. 'Call this after analyzing code violations'), WHAT happens (e.g. 'Stores audit result and advances session state'), and any prerequisites (e.g. 'Requires submit_draft to have been called first'). Reference production baseline: average A+ tool description is 194 chars (p10=34, p90=392).
Document output schemas for all tools. Example for submit_draft: 'Returns: {"type": "object", "properties": {"prompt": {"type": "string"}, "max_retries": {"type": "integer"}, "current_retry": {"type": "integer"}}}'. Document for submit_audit_result what happens after submission (success response, next steps).
Add session_id or user_id parameter to all tools (submit_draft, submit_audit_result, reset_session). Remove global SessionState dependency. This enables multi-user and multi-session scenarios and makes tool chains explicit.
Add validation logic to submit_audit_result: if score < 80, enforce passed=False. Return an error if the LLM sets passed=True and score < 80. Provide actionable error: 'Invalid submission: score (75) is < 80 but passed=True. Set passed=False or increase score >= 80.'
Add dry-run/confirmation for destructive operations. For reset_session: add a 'confirm' boolean parameter (default=False). Require confirm=True to actually reset. For submit_audit_result: return a confirmation object first, require LLM to call confirm_audit_submission() with a confirmation token.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Global session state (SessionState, RulesLoader) creates hidden dependencies. Tools do not accept session_id or user_id parameters, they implicitly operate on 'the' session. This breaks multi-user/multi-agent scenarios and makes tool composition fragile.
submit_audit_result has a critical inconsistency check missing: LLM must set passed=False if score < 80, but the tool accepts any combination of (passed, score, issues). No validation or guidance for conflicting inputs.
reset_session and submit_audit_result are destructive (irreversible state changes) but their descriptions do not warn of this. No dry-run or confirmation pattern offered. Agents could accidentally reset work.
submit_draft output is a system prompt injection disguised as tool output. It instructs the LLM 'STOP GENERATING' and 'Do not output the code yet.' This breaks the MCP abstraction, tools should return data, not control the LLM's reasoning. The prompt should be orchestrated by the client, not embedded in tool output.
Parameter 'language' in submit_draft defaults to 'python' with no validation. LLM could pass unsupported languages (Rust, Go, Cobol). No enum or validation list provided.
update_rules has undocumented parameter dependencies: rule_id is 'required for add/remove/update' but how does the LLM know it is optional for 'list'? The description says '(required for add/remove/update)' but does not state what happens if omitted for 'list'.
No error handling guidance. What happens if submit_audit_result is called before submit_draft? What if update_rules tries to remove a non-existent rule? Tools must tell the LLM what to do next (retry, call a different tool, ask the user).
submit_audit_resultreset_sessionupdate_rules
Move the system prompt from submit_draft output into a separate resource (via Resources capability) or into the client's system prompt. Tool output should be data, not instructions. If the server must guide the LLM, use a structured response like {"stage": "analysis_required", "instructions": "...", "code": "..."}.
Add enum constraint to 'language' parameter in submit_draft. Example: language = enum(['python', 'javascript', 'typescript', 'java', 'go', 'rust', 'c', 'cpp']). Or accept any string but validate against a whitelist in the implementation and return an error if unsupported.
Document parameter dependencies explicitly in each parameter description. For update_rules, rewrite as: "action: Operation to perform - one of 'add' (requires rule_id, severity, description, weight), 'remove' (requires rule_id), 'update' (requires rule_id, and optionally severity/description/weight), 'list' (requires no other params)."
Add error handling guidance to all tools. Return structured error responses: {"error": "<message>", "recovery": "<next step>", "details": "<debug info>"}. Examples: 'Error: submit_draft must be called before submit_audit_result. Call submit_draft with your code first.' or 'Error: Rule ID "rule-123" does not exist. Call update_rules with action=list to see available rules.'
Add field 'issues' output validation: ensure each issue string is non-empty and under 500 chars. Return an error if the LLM passes empty issues or duplicate issues.
Document the audit session lifecycle. Add a resource or tool description explaining: submit_draft → submit_audit_result (or repeat) → reset_session. State that retry limits exist (max_retries from rules_loader) and what happens at the limit.
Add idempotent hints to tool definitions. reset_session should be idempotent (calling it twice has the same effect as calling it once). Document this so agents know it is safe to retry.