Server implements 2 tools with partial schema coverage and reasonable descriptions. Tool naming follows verb_noun convention (search_papers, extract_info), which is correct. Both tools have descriptions and input parameter schemas with type definitions. However, output schemas are not explicitly documented, return types are inferred from docstrings only. The descriptions are moderate length (100-150 chars) and include basic what/when/why context, meeting baseline minimums but lacking LLM-optimized detail. Parameter descriptions are present and type-annotated, but lack constraints (ranges, enums, validation rules). No output schema documentation, error handling guidance, or recovery strategies. Resources are implemented (papers://folders, papers://{topic}) with markdown rendering, showing partial feature completeness. Overall definition quality is fair, above minimal threshold but with significant gaps in schema rigor, error guidance, and composition documentation.
Search for information about a specific paper across all topic directories. Args: paper_id: The ID of the paper to look for. Returns: JSON string with paper information if found, error message if not found.
Search for papers on arXiv based on a topic and store their information. Args: topic: The topic to search for. max_results: Maximum number of results to retrieve. Returns: List of paper IDs found in the search
Output schemas not documented. Tools return List[str] and str respectively, but no structured object definition provided. LLMs cannot infer downstream tool chain requirements or data availability without explicit response schema documentation.
No error handling guidance. Both tools can fail (FileNotFoundError, JSONDecodeError, missing paper) but error responses lack actionable recovery hints. 'No saved information related to paper {paper_id}' is user-friendly but does not guide the LLM toward remediation (search first, check topic, etc.).
Parameter constraints missing. max_results has default=5 but no bounds (min/max). topic is free-form string with no validation rules. Descriptions lack enum constraints, regex patterns, or range specifications. LLMs may pass invalid values without clear self-correction path.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 46 | - | v1 |
Tool composition incomplete. search_papers returns a List[str] of paper IDs, but extract_info expects paper_id as a string. No documentation of the expected workflow or composition dependency. LLMs must infer that extract_info consumes search_papers output.
Descriptions lack LLM-specific context. Both tool descriptions state WHAT but not clearly WHEN to use each tool, what makes them distinct, or what the LLM should do if both are available. Descriptions are 100-150 chars, below the 194-char baseline for A+ tools.