The server has two tools with basic structure but significant quality gaps. Tool naming follows verb-first convention (search_papers, extract_info) which is good. However, descriptions are superficial, parameter descriptions are minimal, output schemas are not documented, and error handling lacks recovery guidance. The tools operate on file-based storage with no validation, injection protection, or structured error messages. This is a typical community server with functional but underdeveloped definitions.
Tools (2)
extract_inforead onlysource verified47/100
Search for information about a specific paper across all topic directories.
search_paperswritesource verified57/100
Search for papers on arXiv based on a topic and store their information.
Output schemas not documented. Both tools return unstructured strings/lists without documenting the shape of returned data, forcing LLMs to parse and guess at structure.
Parameter descriptions are minimal or missing context. 'topic' and 'paper_id' lack guidance on format, valid ranges, or examples of what to pass. max_results has a default (5) but no min/max bounds documented.
Error handling provides no recovery guidance. extract_info returns 'There's no saved information related to paper {paper_id}' but does not suggest how to find the paper or what to try next. Errors are strings, not structured error objects with codes and actionable instructions.
Document output schemas explicitly. For search_papers, return {"paper_ids": [...], "total_found": int, "total_returned": int}. For extract_info, return {"paper_id": string, "title": string, "authors": [...], "summary": string, "published": string, "pdf_url": string} or a clear error object {"error": string, "suggestions": [...]}, not a bare string.
Add bounds and format guidance to parameters. E.g. 'max_results: integer, 1 - 100 (default: 5)' and 'topic: string, 1 - 200 characters, alphanumeric and spaces only'. Validate in the tool implementation and reject invalid input with a clear message.
Enhance error messages with recovery hints. E.g. 'Paper not found: {paper_id}. Try search_papers(topic="...") first to populate the database, or check the paper ID spelling.' Include a list of available papers if relevant.
Sanitize topic parameter against path traversal. Use os.path.normpath() and validate that the resulting path stays within PAPER_DIR. Explicitly validate that topic contains only alphanumerics, spaces, and hyphens. Reject any topic containing '/' or '..'.
Update search_papers description to state: 'Searches arXiv for papers on the given topic and STORES the results to disk (papers/{topic}/papers_info.json). Returns list of paper IDs. This tool modifies state; use extract_info to retrieve stored papers by ID.' This signals to LLMs that the tool has side effects.
Expand extract_info description: 'Retrieves metadata for a specific paper from previously stored search results. PREREQUISITE: Must call search_papers() first on a related topic to populate the database. Returns JSON with title, authors, summary, publication date, and PDF URL, or an error message with suggestions.'
Score history
Overall score trend
↓ 6 points across a rubric change (v1 → v2)
45/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
45
2026-07-28+
v2
2026-03-09
D
51
-
v1
No input validation or injection protection. Topic is used directly in os.path.join() after only a basic .lower().replace(), vulnerable to path traversal (e.g., topic='../../../etc'). Paper ID is used as dictionary key without validation. No sanitization against prompt injection.
search_papers description does not state that it modifies state (creates directories and writes files). LLMs need to know which calls are safe to retry and which have side effects.
extract_info description is vague about what 'information' it searches for. Does it return metadata, full paper content, or stored JSON? The description does not clarify the prerequisite: that search_papers must have been called first to populate the papers_info.json files.
search_papers returns a bare list of paper IDs without additional context (title, URL, etc.). Downstream, LLMs must call extract_info to get usable data, forcing a multi-step workflow. Response should include key metadata needed for chaining.
max_results parameter has a default (5) but no bounds validation. An LLM could pass max_results=999999, causing API strain and long wait times. No documented min/max limits.
search_papers
Return richer data from search_papers. Instead of just paper IDs, return {"paper_ids": [...], "papers": [{"id": ..., "title": ..., "authors": [...], "pdf_url": ...}], "total": ...} so LLMs have enough context to plan next steps without a second call.
Add a dry-run or batch limit to search_papers to prevent accidental large queries. E.g. 'max_results capped at 50 for performance; use pagination tokens for larger sets.' Or return a warning if max_results > 20.
Implement structured error codes. Return JSON objects: {"error": "paper_not_found", "paper_id": ..., "context": "No papers indexed yet", "next_action": "Call search_papers() with a topic first"} instead of plain strings.
Add logging and audit trail. Log who called which tool (if available), what parameters were passed, and what the outcome was. This supports debugging and compliance.