Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This MCP server exhibits significant quality gaps across naming, descriptions, schemas, and error handling. While tool names follow verb_noun conventions, descriptions are sparse and lack LLM-optimized guidance. Input schemas are present but incomplete, many parameters lack descriptions entirely, violating the 100% requirement for A+ tools. Parameter types are defined (good), but the schemas do not document output structures or explain when to use each tool versus alternatives. Error handling is minimal: generic HTTP status codes (422, 502) without actionable guidance. No parameter constraints (enums, ranges, patterns) are visible. The server is HTTP-based (good for protocol readiness) but definition quality is typical of community C-tier servers. Individual tool scores range from 18 (llmreview) to 45 (create_review); overall average is 38.
Tools (5)
accept_reviewwriteauthsource verified58/100
POST /api/v1/reviews/{review_id}/accept - Finalize a review by setting finalized_score and finalized_feedback, marks status as 'finalized'.
create_reviewwriteauthsource verified58/100
POST /api/v1/reviews - Create a new review with structured payload (scores, comments, metadata). Inserts into database and schedules background LLM processing.
get_reviewread onlyauthsource verified62/100
GET /api/v1/reviews/{review_id} - Retrieve a review by ID from the database.
llmreviewread only48/100
POST /llmreview - Calls svc.evaluate_and_parse(review_text=...) and returns validated JSON.
trigger_llm_jobwriteauthsource verified63/100
POST /api/v1/reviews/{review_id}/trigger - Manually trigger background LLM processing for an existing review. Useful for reprocessing or admin testing.
Parameter descriptions missing or generic for most tools. Examples: 'temperature' lacks 0 - 2 range guidance; 'max_points' and 'awarded_points' are union types without explanation of when to use each variant; 'round' parameter in create_review has no valid range.
No enum or constraint declarations. Parameters like 'type' in scores array, 'temperature', and 'round' accept free-form input. LLMs will hallucinate invalid values (e.g., temperature=5.0, round=-1).
llmreviewcreate_review
Recommendations
Document output schemas for all 5 tools. For llmreview, specify fields in returned validated JSON (e.g., {score: number, feedback: string, ...}). For create_review, specify response contains {review_id: int, status: string, created_at: ISO8601}. For get_review, specify full review object structure. For accept_review, specify confirmation object. For trigger_llm_job, specify job status or ID.
Add parameter constraints and ranges. Specify temperature as 0.0 - 2.0 for LLMs; max_attempts as 1 - 50; round as positive integer; max_points and awarded_points as numeric types (remove string union unless truly needed). Use enums for 'type' field in scores (e.g., ["Criterion", "Feedback", "Rating"]).
Enhance tool descriptions with LLM-optimized context. For create_review: 'Create a structured review with rubric scores and comments. Call this after retrieving the assignment and response IDs via discovery tools. LLM will validate and schedule background processing; use accept_review to finalize and publish.' For get_review: 'Retrieve a stored review by ID. Returns all scores, comments, status, and feedback. Use this before calling accept_review to check current state or after trigger_llm_job to see processing results.'
Add a discovery tool: list_reviews(assignment_id=string, course_id=string, page=int, limit=int=20) returning [review_id, response_id, status, created_at]. This enables agents to find reviews without blind ID guessing and supports pagination.
Add error recovery guidance in HTTP responses. For 422 validation errors in llmreview, return JSON: {error: 'ValidationError', detail: 'review_text must be non-empty and <5000 chars', next_steps: ['Verify review_text is a string', 'Reduce length if > 5000 chars', 'Check temperature is 0 - 2']}. For 404 (review not found), return: {error: 'NotFound', detail: 'review_id 999 not found', next_steps: ['Call list_reviews() to find valid IDs', 'Verify review was created before fetching']}.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Error handling is minimal. Generic HTTP status codes (422, 502) without actionable recovery guidance. Example: llmreview returns 422 for 'ValidationError' but does not tell the LLM what to do next (retry? fix input? call a different tool?).
Tool relationships undocumented. Unclear which tools must be called in sequence (create_review → accept_review?) or whether tools are idempotent. No discovery tool (e.g., list_reviews) to help agents find valid IDs.
Descriptions are too brief or generic. Baseline for A+ tools is 50 - 200 chars; most here are 35 - 196 chars but lack WHEN/WHY context. Example: get_review description does not explain when to use it vs querying the database directly or why an agent would call it.
Pagination not visible. If get_review or any list tool can return large datasets, there is no limit, offset, or cursor for pagination. Risks context window overflow.
get_review
Declare parameter mutual exclusivity and interdependencies. Example: In create_review, note that 'scores is required for structured review; if empty, call trigger_llm_job to generate scores from LLM.'
Add permission scopes to tool descriptions. Prefix each with 'Requires: read:review' or 'Requires: write:review, schedule:llm_job'. This enables agents to request minimal permissions.
For accept_review, explicitly state idempotence: 'Idempotent: calling this multiple times with the same review_id and finalized_score/feedback will succeed without side effects. Safe to retry on transient failures.'
Separate nullable fields into optional/required. For accept_review, clarify: 'finalized_feedback (required): human-approved feedback; null not allowed' or 'finalized_feedback (optional): if null, retain existing feedback.' This prevents LLMs from passing null when a value is needed.
Add SQL injection / prompt injection guards to descriptions. Note in reviews.py that response_id_of_expertiza is validated as integer before query. Sanitize assignment_name and course_name as strings (max 255 chars, alphanumeric + spaces only) to prevent SQL injection via malicious agent inputs.