Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
This MCP server has significant definition quality gaps. Tools are implemented as Flask routes but lack proper MCP structure. While 13 tools are present, many have incomplete schemas, vague descriptions, and missing parameter documentation. The server appears to be a Flask web app retrofitted as an MCP server without following agentic tool patterns. Average tool score is 38/100, reflecting pervasive issues with naming clarity, parameter descriptions, schema completeness, and error handling guidance.
Tools (13)
admin_dataread onlyauth50/100
Retrieves admin dashboard data including episodes, categories, top scorers, total users and questions
calculate_scorewriteauthsource verified45/100
calculates scores
create_temp_userwritesource verified73/100
Create temporary username and password
get_episodesread onlyauth50/100
Retrieves all episodes
get_questionsread onlyauthsource verified70/100
retrieves trivia questions
get_top_scoresread onlyauthsource verified67/100
Get top five scorers and current user's score and position
calculate_score has non-standard schema with only 'description' field and no 'properties' or 'required' array. Parameter structure completely opaque to LLMs, does not follow JSON Schema conventions. LLM cannot infer valid input structure.
update_user_account has only a generic type description, no field names specified. LLM cannot determine which account fields are updatable (email? password? role? preferences?). Requires introspection or trial-and-error.
admin_data description is too vague ('Retrieves admin dashboard data including...') and omits critical details: Does it require admin role? What is the exact structure returned? Are there pagination limits? Does it include sensitive data like password hashes?
Expand tool descriptions from current 20-60 characters to 100-200 characters. Follow the pattern: 'Retrieves [noun]. Call this when [use case]. Returns [key fields]. Prerequisites: [auth level].' Example: 'Retrieves a user account by ID. Call this to fetch the current user profile before updating. Returns username, email, score, role, created_at. Requires user authentication.'
Add explicit input/output schemas to every tool. Use proper JSON Schema with 'type', 'properties', 'required', and 'additionalProperties': false. For calculate_score, replace the vague 'Player answers keyed by question ID' with: {"type": "object", "properties": {"questionId": {"type": "string", "description": "Question ID"}, "answerText": {"type": "string", "description": "Selected answer text"}}, "required": ["questionId", "answerText"], "additionalProperties": false}
Document all enum constraints. For register_user 'role' parameter, add: 'enum': ["member", "admin"] and description: 'User role. Must be one of: member (default), admin. Cannot be changed after registration.'
Add error handling guidance to each tool description. Example for login: 'On failure, returns {"error": "Invalid credentials"}. If you receive this, try: (1) search for the user with search_users() to confirm they exist, (2) offer password reset via send_password_reset_email(), or (3) suggest account creation via register_user().'
Specify parameter format constraints. Example for episodeId: 'episodeId (string, required): UUID v4 format (e.g. 550e8400-e29b-41d4-a716-446655440000). Call get_episodes() first to discover valid IDs.' and for role: 'role (string, optional, default: member): One of member|admin. Cannot be changed after registration.'
Score history
Overall score trend
↑ 57 points across a rubric change (v1 → v2)
57/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
57
2026-07-28+
v2
2026-03-09
F
0
-
v1
source verified
75/100
Log in a user with credentials
logoutwriteauthsource verified75/100
Log out a user
new_episodewriteauth50/100
Create a new episode
new_questionwriteauth50/100
Create a new trivia question
register_userwritesource verified73/100
Register a new user with username and password
update_user_accountwriteauth48/100
Update user account information
Destructive tools (logout, delete implied in admin_data, register_user with WRITE risk) lack confirmation or dry-run patterns. An agent cannot preview consequences before calling logout or registering thousands of dummy users.
No error handling guidance in any tool description. If login fails, what should the LLM do? Retry? Search for the user? Ask the human for password reset? Error responses appear to be JSON, but no documentation tells the LLM what to expect or how to recover.
Parameter descriptions are minimal or missing detail. 'episodeId' in get_questions lacks format details (UUID? integer? string?). 'role' in register_user has no enum constraint, LLM might invent roles like 'superadmin' or 'guest' that don't exist in the system.
No output schemas documented. What does get_top_scores return? Is it {top_five: [...], current_user: {...}}? Does it include timestamps, percentiles, or raw scores? LLM cannot plan follow-up actions without knowing response structure.
Tools operating on the same resource (get_user_account, update_user_account, register_user) use inconsistent naming and unclear field mappings. Does update_user_account accept 'username' or 'user_id'? Is there a get_account vs get_user_account ambiguity?
Pagination not documented. get_episodes, get_top_scores, and admin_data could return large lists, but no limit, offset, or cursor parameters are visible. Returning 1000 episodes will blow LLM context.
Separate user lookup from update. Instead of one generic update_user_account, create update_email(), update_password(), update_profile() with explicit parameters and validation rules. This prevents the LLM from passing invalid field combinations.
Add confirmation patterns to destructive tools. For register_user, add a dry_run parameter (boolean, default false). When dry_run=true, validate input and return what would be created without persisting. Example response: {"valid": true, "preview": {"username": "user_123", "role": "member"}, "message": "Ready to register. Call again with dry_run=false to confirm."}
Document admin_data access control explicitly. Add to description: 'Requires admin role. Returns aggregated statistics only (no individual user details). Includes: episode count, category count, top 5 scorers, total user count, question count. Pagination: max 1000 records per request.'
Add idempotency keys for critical operations. Modify register_user to accept an optional 'idempotency_key' field. Document: 'If the same idempotency_key is sent twice, the second call returns the same result without creating a duplicate user. Use a UUID derived from the username to ensure safe retries.'
Remove or document the role parameter default in create_temp_user and register_user. If role always defaults to 'member', state it explicitly. If some users should be admin, require an explicit role parameter in the schema, not a hidden default.