FastAPI backend for an AI-powered test case generation and management platform with chat, file upload, and test case approval features
This MCP server has severe definition quality issues across nearly all dimensions. Tool descriptions are present but often generic and lack actionable context for LLM selection. Input schemas exist for most tools but lack proper type definitions and constraints. Critical issues include: (1) ambiguous tool names that do not clearly signal intent (e.g., 'root' vs 'health_check' are redundant; 'chat' vs 'chat_stream' lack verb clarity); (2) descriptions that are too brief (many under 50 chars, far below the 194-char baseline for production tools); (3) missing parameter descriptions in several tools; (4) no output schemas documented; (5) no error handling guidance; (6) parameters with no type constraints or validation hints; (7) tools combining multiple concerns (e.g., chat tools do not separate message composition from execution). The server is a test platform API frontend, not a well-engineered MCP tool composition.
User approval endpoint for reviewing and accepting/rejecting generated test cases
Standard (non-streaming) chat interface for user messages
Streaming chat interface that integrates with AutoGen for real-time responses
Generate test cases from an uploaded file using AI agents
Retrieve agent conversation messages for a specific test case generation batch
Retrieve all generated test cases from the database
Get detailed information about a specific uploaded file including content preview
Redundant and ambiguous tool naming: 'root' and 'health_check' both serve diagnostic purposes but lack clarity about when to call each. HTTP health checks should not be exposed as tools; they are infrastructure concerns, not user-facing operations.
Descriptions are too brief and lack actionable context. Baseline for production tools is 194 chars; most tools here are 30-50 chars. Example: 'Health check endpoint that returns the status of the API' (57 chars) does not explain WHEN to call it or what the agent should do with the response. Descriptions must state WHAT, WHEN, and WHY.
Missing or undocumented output schemas. No tool documents what fields are returned, their types, or how downstream tools should consume them. This forces LLMs to infer structure, leading to hallucination and failed chaining. Example: 'get_files' returns paginated results but schema does not specify field names, types, or total_count.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 27 | - | v1 |
Retrieve a paginated list of all uploaded files
Get statistics and metrics about test case generation
Retrieve test cases for a specific file
Health check endpoint that returns the status of the API
Root health check endpoint that returns a simple message indicating the API is running
File upload endpoint that accepts PDF, Word, Excel, and text files with content extraction
Parameters lack type definitions and constraints. Example: 'chat_stream' accepts 'model' parameter but description does not enumerate valid values, provide regex patterns, or specify defaults. Unbounded string parameters invite LLM hallucination.
No error handling guidance. Tools do not document what happens on failure, whether errors are retryable, or what the LLM should do next. Example: 'upload_file' will fail on unsupported file types, but the error response description is absent.
Chat tools ('chat' and 'chat_stream') do not declare their relationship or explain when to use each. Both accept the same inputs and likely produce overlapping outputs. LLMs will waste reasoning cycles deciding between them or call both unnecessarily.
Parameter descriptions are missing or trivial in several tools. Example: 'get_test_cases' accepts 'file_id' but description does not explain what file_id is, how to obtain it, or whether it is required. LLMs cannot infer such details.
No tool declares idempotency, mutability, or permissions. 'approve_test_case' is clearly destructive (modifies state), but this is not signaled in the tool definition. Agents cannot distinguish between safe read-only calls and irreversible mutations, risking incorrect retry logic.
Pagination strategy not documented. 'get_files' accepts 'skip' and 'limit' but does not specify max limit, default values, or whether a total count is returned. Result limits are not capped, risking context window exhaustion if an LLM requests thousands of records.