A comprehensive MCP toolkit for research productivity and lab management with tools for email, calendar, file operations, code analysis, academic literature, maps, screenshots, document editing, and more
Server has 3 tools with complete input schemas and descriptions. Naming follows verb_noun patterns (download, generate_prompt, sync). Descriptions are substantive (80-180 chars) and explain use cases. However, output schemas are entirely undocumented, critical omission for LLM planning. Parameters lack some constraints (enums for format values, timeout bounds). Error handling guidance is absent. Risk annotations present (READ_ONLY, WRITE) but not formalized as toolAnnotations per MCP spec.
Downloads an ArXiv paper by ID in various formats ('src', 'pdf', 'tex', 'bibtex') or lists its source files. Use search to find arxiv IDs by author or topic.
Generates a structured, token-counted summary of a codebase for LLM analysis. Pass output to chat tool for review. Supports include/exclude, git diffs, and formatting options.
Unified email synchronization with multiple modes.
No output schemas documented for any tool. LLMs cannot reason about what fields to expect or plan downstream calls.
Format parameter (download) and mode parameter (sync) lack enum constraints. Free-form strings invite invalid values.
Timeout_seconds parameter has no bounds (0 or positive). Unbounded numeric parameters allow absurd values (e.g. 999999).
No error recovery guidance. Tool descriptions do not explain what LLM should do on failure or what errors to expect.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 35 | - | v1 |
Risk annotations (READ_ONLY, WRITE) are metadata but not formalized as MCP toolAnnotations (readOnlyHint, destructiveHint). Agents cannot distinguish safe vs unsafe operations without parsing comments.
generate_prompt has many optional boolean parameters (line_numbers, full_directory_tree, follow_symlinks, etc.) with no guidance on typical usage. LLMs will guess default behavior.