MCP server for searching and extracting research paper metadata from arXiv
Two tools with basic schemas and generic descriptions. Both tools have input schemas visible in the source code (chatbot-basic.py and chatbot-mcp-server.py), but descriptions lack context for LLM tool selection. Parameter descriptions are minimal. No output schema documentation. No error handling guidance. No consideration of security, idempotency, or composition patterns. The server performs file system I/O with minimal input validation, creating security risks.
Search for information about a specific paper across all topic directories.
Search for papers on arXiv based on a topic and store their information.
Tool descriptions lack WHEN-to-use context. 'Search for papers on arXiv based on a topic and store their information' does not explain when to call this vs a hypothetical list_cached_papers or refine_search. LLM tool selection relies on descriptive clarity.
Parameter 'max_results' has no min/max bounds declared in the description or schema. Unbounded integers invite absurd values (e.g., max_results=999999) that could overwhelm the arXiv API or timeout. The description says 'default: 5' but does not specify valid range.
No input validation in tool implementations. Path traversal vulnerability: extract_research_info accepts any 'research_doc_id' string and uses it directly in os.listdir() loops without sanitizing. A malicious input like '../../../etc/passwd' could expose files outside the research directory.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
No output schema documented for either tool. search_research_papers returns 'List[str]' (paper IDs) but does not specify the format, validation, or structure of the IDs. extract_research_info returns a JSON string or error message, unstructured output requires the LLM to parse freeform text, which is error-prone.
No error handling guidance for LLM recovery. When extract_research_info returns 'Paper with ID ... not found in any topic directory', the LLM does not know whether to retry, ask the user for clarification, or call search_research_papers. Error messages must be actionable.
Naming ambiguity in extract_research_info: the tool searches across directories but returns raw JSON. The name does not clearly convey that it retrieves cached paper metadata, the verb 'extract' is vague (extract what, from where?). Better: 'get_cached_paper_info' or 'retrieve_paper_metadata'.
search_research_papers has a side effect (writes to local filesystem) but does not document this in the description. Descriptions for write/stateful tools must declare side effects so agents understand the operation is not idempotent and may have costs.
Parameter 'research_doc_id' in extract_research_info has no format specification. arXiv IDs have a known format (e.g., '1509.01213v1'), but the description and schema do not document this. LLMs may pass arbitrary strings, causing lookup failures. Specify the format: 'arXiv paper ID (format: YYMM.NNNNN or YYMMNNNNv<version>)'.