This MCP tool lets AI agents ask for assistance or feedback from more capable models
PeerReviewMCP has two explicitly registered tools with basic Zod schemas and descriptions, but falls short of production quality. Tool names lack clear verb-noun structure ('ask-expert-for-peer-review' uses 'ask' which is vague; 'ask-expert-for-help' repeats the pattern). Descriptions exist but are generic and lack WHEN/WHY guidance. Parameters (problem, plan, issue) have type definitions (string) but descriptions are minimal and don't explain format constraints or dependencies. No output schema is documented, responses are hardcoded strings without structured fields. Error handling is minimal (only missing env vars). Security: the server invokes external LLMs via HELPER_MODEL_API_KEY, which is correctly injected via environment variables (no credential leakage in params), but no audit logging or rate limiting. Composition is weak, two tools do almost identical things (both ask an expert; only the prompt template differs), risking agent confusion. No resource chaining, pagination, or recovery guidance.
When confronting a challenging problem, seek guidance from an expert. Specify the problem and the issue you are having, as well as any other context.
Ask an expert to peer review your plan. Provide the problem you are working on, and your detailed plan as well as any other context.
Tool names are vague and lack clear verb-noun action structure. 'ask-expert-for-peer-review' and 'ask-expert-for-help' both use 'ask' + similar structure, making disambiguation difficult for LLMs. Pattern: verb_noun (e.g., 'review_plan', 'get_expert_feedback') would be clearer and more composable.
Parameter descriptions are minimal (1-6 words) and lack format guidance. 'problem' and 'issue' are semantically similar but never disambiguated. 'plan' lacks guidance on expected length, detail level, or structure. Pattern: descriptions should be 50 - 150 chars, stating WHAT the param is, WHAT format it expects (plain text, structured outline, pseudocode?), and WHY it matters.
No output schema is documented. Responses are free-form strings ('The response from the expert is: <text>'). LLMs cannot parse or extract structured data. Pattern: tools should return typed objects (e.g., {type: 'text', text: string, feedback: {strengths: string[], weaknesses: string[], suggestions: string[]}} for peer review).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 42 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Tool semantic overlap: both tools call the same external LLM with different prompts. From an LLM agent's perspective, these are nearly identical. Pattern: tools should each do one distinct thing, or they should be parameterized variants. Consider a single 'ask_expert' tool with a 'request_type: enum' param (peer_review | guidance).
Error handling is minimal. Only env var validation raises an explicit error. Missing: (a) timeout handling for external LLM calls, (b) retry guidance for transient failures, (c) fallback instructions if helper model returns empty or malformed text. Pattern: errors should guide recovery (e.g., 'API timeout. Retry with a shorter plan, or simplify the problem statement.').
No audit logging or rate limiting. Calls to external LLMs are unmonitored, no logging of request/response, latency, costs, or who called what. Pattern: tools should log callers, parameters, outcomes, and timestamps for compliance and debugging.