MCP server for SciTeX Hub - an AI-powered scientific document authoring platform with LaTeX writing, data analysis, visualization, and academic research tools
Critical failure across all dimensions. The server provides 17 tools with tool names inferred from a configuration file (data/users/.mcp.json), but NO actual tool implementation code is visible in the source. Tool definitions appear to be static declarations in JSON with only names and one-line descriptions. No input schemas, output schemas, parameter definitions, or error handling are evident. The Makefile and partial Python infrastructure suggest this is a complex orchestration system, but no MCP tool registration, schema definitions, or protocol implementation details are provided. This server cannot be scored reliably because the actual MCP implementation is not visible in the provided code sample.
Tool definitions appear to be inferred from static JSON configuration (data/users/.mcp.json) rather than explicitly registered with MCP protocol handlers. No CallTool, ListTools, or tool registration code visible. Combined with missing schemas, this caps all tools at 0-10.
URGENT: Provide the actual MCP server implementation code (tools.py, tool_definitions.json, or equivalent). The Makefile is infrastructure, not tool definitions. Cannot evaluate without seeing how tools are registered.
Add explicit, verb-based tool names: replace 'plt' with 'create_plot', 'stats' with 'calculate_statistics', 'scholar' with 'search_publications', 'writer' with 'edit_latex_document', 'clew' with 'organize_document_structure', 'audio' with 'transcribe_audio', 'diagram' with 'generate_diagram', 'capture' with 'capture_screenshot', 'introspect' with 'analyze_code', 'template' with 'list_templates' or 'apply_template', 'project' with 'create_project' or 'list_projects', 'dataset' with 'load_dataset' or 'analyze_dataset', 'dev' with 'run_debug_tool' or 'execute_dev_command', 'linter' with 'lint_latex_document', 'social' with 'share_document' or 'invite_collaborator', 'ui' with 'render_ui' or 'get_ui_state', 'usage' with 'get_usage_stats' or 'track_usage'.
Write comprehensive descriptions (50-200 chars) for each tool: state WHAT it does, WHEN to use it (vs similar tools), and WHAT it returns. Example: 'Create a scientific plot from data: scatter, line, bar, heatmap. Call this after load_dataset(). Returns plot_id for embedding in documents.'
Define complete input schemas for each tool with JSON Schema: specify parameter names, types (string, number, boolean, array, object), required flags, descriptions, enums for constrained values, and min/max for numeric ranges. Example: { 'type': 'object', 'properties': { 'plot_type': { 'type': 'string', 'enum': ['scatter', 'line', 'bar', 'heatmap'], 'description': 'Type of plot to generate' }, 'data_source': { 'type': 'string', 'description': 'Dataset ID (get from load_dataset)' }, 'title': { 'type': 'string', 'description': 'Plot title (optional)' } }, 'required': ['plot_type', 'data_source'] }
All 17 tool descriptions are trivial one-liners (10-40 chars). Rule violation: descriptions must be 10-1024 chars and explain WHAT, WHEN, and WHY the LLM should use the tool.
No evidence of parameter definitions, constraints, or enums for any tool. LLMs cannot determine what parameters to pass or what values are valid. Violates pattern:constrained-input.
No output schemas documented. LLMs cannot plan downstream tool calls or extract the right data from responses. Violates pattern:tool-schema and pattern:response-shaper.
Tool names do not follow verb_noun pattern. Names like 'plt', 'stats', 'clew', 'dev', 'ui', 'usage' are nouns or abbreviations, not action verbs. LLMs cannot infer intent from the name alone. Should be 'create_plot', 'analyze_statistics', 'generate_diagram', etc.
No error handling, recovery guidance, or error classification visible. If a tool fails, LLMs have no guidance on whether to retry, ask the user, or give up. Violates pattern:recovery-guide and pattern:error-classification.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible. Four tools marked WRITE risk (writer, project, dev, social) but no schema indicators to signal destructive behavior to LLMs. Violates pattern:command-tool.
No MCP protocol implementation code visible. Provided source is a Makefile for Docker orchestration and environment management, not MCP server implementation. Cannot verify CallTool, ListTools, Tool request/response handling, or protocol compliance.
Document output schemas: describe the fields returned (plot_id, plot_url, dimensions, format, error_code, error_message, etc.) so LLMs know what to expect and can chain calls correctly.
For the 4 WRITE tools (writer, project, dev, social), add destructiveHint and idempotentHint annotations in the tool schema. Explicitly note in descriptions which operations are reversible and which are not. Example: 'This tool modifies the document. Changes are saved to version control and cannot be undone with undo_document_edit().'
Add error handling with recovery guidance: instead of returning a bare error, explain what went wrong and what to try next. Example: '{ "error": "Dataset not found: ds-123", "recovery": "Call list_datasets() to see available datasets, then retry with a valid dataset_id" }'
Reduce parameter ambiguity: if a tool accepts user input, provide separate typed parameters (e.g., user_id vs user_email) rather than a single overloaded 'user' param. If parameters are mutually exclusive, state this explicitly.
Add parameter descriptions explaining format and constraints: e.g., 'plot_type: Must be one of scatter, line, bar, heatmap. Default: scatter.' or 'limit: Number of results to return. Range: 1 - 100. Default: 20.'
Consider composition: some tools appear to combine multiple concerns (e.g., 'writer' for LaTeX writing and editing). Split into separate tools: edit_latex_document, insert_section, delete_section, apply_formatting so agents can compose them granularly.
If any tool returns lists (e.g., list_datasets, search_publications), add pagination: include limit, offset/cursor parameters and return a total_count or next_cursor in responses. Cap results at 20 - 50 items by default.
Strip irrelevant metadata from responses: if an API returns 50 fields, return only the 5 - 10 fields the LLM needs for reasoning and chaining. Every extra token costs money and dilutes signal.
Validate all inputs inside the tool handler and return clear, actionable error messages. Do not let invalid data propagate to downstream services. Example: 'Invalid priority: got "high-urgent", must be one of: low, medium, high, urgent.'
If tools handle sensitive data (user documents, research datasets), add security notes: which parameters require authentication, which operations are logged, which data is encrypted, etc. Reference pattern:secret-injection.
For tools like 'social' (collaboration), gate destructive operations (delete collaborator, revoke access) behind permission checks. Verify the calling user has authority before executing.
Add idempotency keys or deduplication for write operations. E.g., create_project(project_id, ...): if project_id already exists, return the existing project instead of failing or creating a duplicate.
Test tool composition: ensure tool A's output contains the IDs and references tool B needs. Example: create_plot returns plot_id; embed_plot_in_document accepts plot_id. No downstream tool should require an extra lookup call.
For tools that interact with external APIs (scholar, social), document rate limits, timeouts, and fallback behavior. Set explicit timeouts (not infinite waits) and return timeout errors with retry guidance.