The MCP Research Server demonstrates basic tool implementation with clear verb-noun naming and reasonable parameter schemas, but falls short of production quality due to incomplete descriptions, missing output schema documentation, and lack of error handling guidance. Four tools are properly registered via the @mcp.tool() decorator, each with input schemas using proper JSON types. However, tool descriptions are generic and lack the LLM-optimization required for reliable tool selection. Parameter descriptions are minimal (often just restating the parameter name). Critical gaps include: (1) no documented output schemas for any tool, making it impossible for LLMs to plan downstream operations; (2) error responses are unstructured (print statements, bare strings) with no guidance on recovery; (3) no discussion of idempotency, side effects, or prerequisites; (4) missing parameter constraints (e.g., no min/max for max_results, no format guidance for paper_path); (5) tools perform file I/O without documented permission scopes or audit capabilities. The tools' logic suggests they interact with the filesystem and arxiv API, but these dependencies and failure modes are not exposed to the LLM via descriptions. Baselines: average tool description in production is 194 chars; these average ~110 chars and lack WHEN/WHY context. Average param description is 72 chars; these average ~40 chars and are often trivial.
Tools (4)
extract_inforead onlysource verified58/100
Search for information about a specific paper across all topic directories.
extract_text_paperwritesource verified55/100
Extract text content from a PDF file and save it as a text file.
NO OUTPUT SCHEMAS DOCUMENTED for any tool. LLMs cannot predict return types, structures, or downstream tool compatibility. All four tools return values but none document what fields/structure the LLM should expect. This forces LLMs to guess and breaks tool chaining.
Tool descriptions are generic and lack WHEN/WHY/PREREQUISITE context. Descriptions range 67 - 115 chars (below 194-char baseline). None explain when to use one tool vs another (e.g., when to call search_papers vs generate_search_prompt), mention prerequisites, document side effects, or provide failure guidance.
Document OUTPUT SCHEMA for every tool. Example for search_papers: 'Returns: {"paper_ids": ["2312.xxxxx", "2312.yyyyy"], "count": 2, "saved_to": "/path/to/topic/papers_info.json"}'. For extract_text_paper: 'Returns: {"success": true, "output_path": "/path/to/file.txt", "pages_extracted": 42} or {"success": false, "error": "PdfReadError: corrupted PDF", "recovery": "Check file integrity or try a different PDF"}'.
Expand tool descriptions to 150 - 250 chars. Add WHEN/WHY context and prerequisites. Example for search_papers: 'Search arXiv for papers matching a topic and save their metadata locally. Call this first to populate the papers database; then use extract_info to retrieve saved papers or extract_text_paper to convert PDFs to text. Requires internet access to arxiv.org and write permission to ./papers directory.'
Add min/max constraints to numeric parameters via both schema and description. Example: 'max_results: integer, minimum 1, maximum 50 (arxiv limits). Default: 5. Note: Higher values may hit rate limits; retry after 5 seconds if throttled.' Similarly, num_papers should specify 1-100 range.
Add format guidance to string parameters. Example for paper_path: 'Absolute or relative path to a PDF file (e.g., /data/papers/2312.xxxxx.pdf or ./papers/ai/paper.pdf). Must be readable and a valid PDF. Relative paths are resolved from the current working directory.' For paper_id: 'arXiv paper ID format (e.g., 2312.xxxxx). Case-sensitive. Must be previously saved via search_papers.'
Implement structured error responses. Instead of returning bare strings or None, return JSON objects: {"error": "<message>", "error_code": "<code>", "recovery": "<next step>"}. Example: extract_info should return {"error": "paper_not_found", "message": "Paper 2312.xxxxx not found in saved topics.", "recovery": "Call search_papers with a topic first, or check paper_id format."}
Score history
Overall score trend
↑ 1 points across a rubric change (v1 → v2)
42/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
42
2026-07-28+
v2
2026-03-09
F
41
-
v1
Parameter descriptions are trivial and lack format/constraint guidance. Examples: 'The topic to search for' (restates param name), 'The ID of the paper to look for' (restates param name). No min/max for numeric params (max_results, num_papers should specify 1-50 range or similar). No format guidance for string params (paper_id format? paper_path absolute vs relative?).
Error handling is unstructured and provides no recovery guidance. extract_info returns a bare string 'There's no saved information related to paper {paper_id}.' with no actionable next step. extract_text_paper returns None on failure (PdfReadError, file not found) with only a print() statement, LLM receives no error info at all. Try search_papers() first to save it.').
Side effects and idempotency not documented. search_papers creates filesystem directories and appends to JSON files, repeated calls with same input do NOT produce identical results (new entries added to papers_info.json). extract_text_paper overwrites .txt files (idempotent in that sense, but lossy).
Path traversal vulnerability in extract_text_paper. paper_path parameter accepts raw user input; while code checks os.path.isfile(), it does not prevent parent directory escape (../../../etc/passwd could slip through if the underlying filesystem allows it).
No permission scopes declared. Tools interact with filesystem and external API (arxiv) but provide no scope declarations (e.g., 'read:filesystem', 'write:filesystem', 'read:arxiv').
No audit logging or traceability. Tools perform filesystem I/O and API calls without logging who called them, with which parameters, when, or what happened.
search_papersextract_infoextract_text_paper
Document idempotency and side effects in tool descriptions. Example for search_papers: 'SIDE EFFECTS: Creates ./papers/<topic>/ directory if not exists. Appends new papers to papers_info.json; repeated calls may add duplicates if arxiv results change. Consider calling once per topic, then use extract_info for lookups. NOT idempotent.'
Sanitize file paths in extract_text_paper. Use os.path.abspath() and verify the resolved path is within an allowed directory root. Example: allowed_root = os.path.abspath('./papers'); resolved = os.path.abspath(paper_path); if not resolved.startswith(allowed_root): raise ValueError('Path outside allowed directory'). Add to description: 'Paths must be within the ./papers directory; parent directory references (..) are rejected.'
Add permission scope declarations to tool docstrings or as metadata. Example: 'Requires scopes: ["read:filesystem", "write:filesystem", "read:arxiv"]. Least-privilege agent should have only these scopes assigned.'
Implement audit logging. Log tool calls to stderr or a structured log file. Example: import logging; logger = logging.getLogger('mcp'); logger.info(f'tool=search_papers user=<agent_id> topic={topic} max_results={max_results} status=started'); ... logger.info(f'...status=completed found={len(paper_ids)} papers'). Add to description: 'All calls are logged for audit and debugging.'
Add timeout guidance for tools that call external APIs. Example for search_papers: 'May take 2 - 10 seconds depending on arxiv load. Agent timeout should be ≥15 seconds. If timeout occurs, retry after 5 seconds.' Missing from current code.
Consider batch variants or loop-optimized tools. If agents frequently call extract_text_paper on multiple PDFs, offer extract_text_papers (plural) accepting a list of paths, returning per-item success/failure. This saves tokens and latency vs N sequential calls.
Use human-friendly identifiers in parameter names and descriptions where possible. 'paper_id' is good (not 'id'); 'paper_path' is clear. Current state is acceptable, but for future tools, prefer 'user_email' over 'user_id' or support both with separate parameters.