A command-line research chatbot using Claude and arXiv for searching academic papers and extracting metadata
The server defines 2 tools with adequate naming (verb_noun pattern) and clear, descriptive descriptions. However, parameter descriptions are sparse or missing entirely, input schemas lack type constraints (enums, min/max), and output schemas are not formally documented. The search_papers tool performs a WRITE operation but lacks idempotent hints or dry-run support. extract_info returns unstructured JSON strings instead of typed objects, forcing LLM parsing. No error handling guidance is present, failures return raw strings without recovery hints. Logging is present but not integrated with error responses.
Retrieve stored metadata for *paper_id* across all topics.
Search arXiv for *topic* and persist metadata. Returns stored short IDs.
Parameter descriptions are minimal or absent. The 'topic' parameter lacks format/length constraints; 'max_results' has no min/max bounds stated in the description.
Input schemas lack formal constraints (enums, minLength, maxLength, minimum, maximum). The 'topic' parameter accepts any string; no validation guidance is provided in the description.
Output schemas are not formally documented. search_papers returns a List[str] of paper IDs without describing the structure. extract_info returns raw JSON string, untyped, forcing the LLM to parse manually.
search_papers is a destructive/side-effecting tool (persists metadata to disk) but lacks destructive hints, dry-run capability, or confirmation steps. Agents cannot reason about idempotency.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Error responses are unstructured strings (e.g., 'No papers database found' or 'No stored information for paper'). They provide no recovery guidance, the LLM cannot infer what to try next.
extract_info returns entire paper metadata as a JSON string. No pagination or limit is enforced. If a paper has a very long summary, token cost is not managed.
Resources (papers://folders, papers://{topic}) are present but not chainable to tools. The LLM must infer that @papers://folders lists topics and then manually construct a topic name to pass to search_papers. No dependency hints link resources to tools.