A collection of MCP server implementations for various services including arXiv paper search/retrieval, graph generation, Kaggle dataset search, markdown-to-PDF conversion, web scraping via Playwright, and Yandex Art image generation
Collection of 7 tools with mixed quality. All tools have explicit definitions with names, descriptions, and input schemas visible in source code. Schemas are well-structured with enums and constraints. However, output schemas are completely undocumented, the code does not show what these tools return, making it impossible for LLMs to plan downstream operations. Error handling exists but is minimal and not actionable. Descriptions are adequate (100-150 chars) but lack usage context and prerequisites. No tool annotations (readOnlyHint/destructiveHint) despite clear risk profiles (arxiv tools are READ_ONLY; graph/markdown/image tools are WRITE). No composition hints or field-matching between tools.
Retrieves a specific paper from arXiv by arxiv_id
Searches arXiv papers with configurable query parameters
Generates a graph (bar, line, or pie chart) from provided data with customizable options
Generates an image from a text prompt using Yandex Art AI model
Output schemas completely undocumented. Code shows error handling returns JSON with error messages, but no tool shows what successful responses look like (field names, types, nested structures). LLMs cannot infer downstream parameters or chain tools without knowing what fields are returned.
No tool annotations despite clear risk profiles. arxiv_get and arxiv_search are READ_ONLY (safe to retry), but lack readOnlyHint. generate_graph, markdown_to_pdf, and generate_image are WRITE operations (resource-consuming, state-changing) but lack destructiveHint. This information is critical for agent safety and retry logic.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-03-09 | F | 40 | - | v1 |
Error handling is minimal and not actionable. Code returns error dict with 'error' field containing a message, but does not categorize errors as retryable vs. user-fixable vs. fatal. No recovery guidance. E.g., 'paper with arxiv_id ... was not found' gives no hint to try search_arxiv() first.
Parameter descriptions lack actionable constraints. fetch_web_pages describes 'wait_until' enum but omits what each value means (load = DOMContentLoaded? networkidle = zero pending requests?). generate_graph 'options' parameter is object without structured field descriptions, LLMs must infer allowed keys from the description prose alone.
Descriptions lack usage context and prerequisites. generate_image requires folder_id and api_key but description says 'optional if env var set', unclear to LLM when env vars are available. No guidance on when to use generate_graph vs. markdown_to_pdf or whether tools can be chained (e.g., fetch_web_pages → generate_graph → markdown_to_pdf).
No batch variants for tools agents call in loops. fetch_web_pages accepts max 5 URLs per call, if agent needs to fetch 20 pages, it makes 4 sequential calls, wasting tokens and latency. No indication whether tool is idempotent (can be safely retried).
max_results defaults and limits vary inconsistently. arxiv_search defaults to 3, max 25. kaggle_search defaults to 3, max 5. No guidance on why limits differ or how agent should interpret truncated results. Should also document whether pagination (offset/cursor) is supported.