Collection of Model Context Protocol (MCP) servers including template, arXiv, Chrome DevTools, Google Drive, code execution, Notion, OpenRouter, and Slack integrations
This is a large heterogeneous collection (41 tools across 8+ servers and frameworks) with severe quality gaps. Only ~40% of tools have adequate descriptions, parameter schemas are incomplete across most tools, and tool names show inconsistency and duplication. The arxiv-mcp and chrome-devtools-mcp subservers have the most complete schemas, but the collection as a whole lacks the coherence, documentation rigor, and error-handling guidance expected of production tools. Many tools lack actionable descriptions, some parameters are undocumented, and there is no evidence of consistent validation or recovery guidance. The presence of duplicate tools (click, type_text, navigate across multiple servers with varying descriptions and parameter sets) signals poor composition and likely confusion for LLM selection.
Send a message to an LLM and get a response
Click on an element in the page
指定したセレクタの要素をクリックする
Close a browser page
Drag an element to a new position
Emulate a device or network condition
Evaluate JavaScript in the page context
Fill a form field with text
Fill multiple form fields at once
Tool 'hello' (template) has no description and no input schema, completely undocumented. Should be removed or properly specified.
Duplicate tools across servers with inconsistent schemas and descriptions: 'click' appears in chrome-devtools-mcp (TypeScript, CSS selector focus) and agent-browser (Python, possibly different behavior); 'type_text' and 'navigate' similarly duplicated. This forces LLMs to guess which variant to call and risks silent failures.
Many tools have descriptions under 20 characters or are generic: 'List all available arXiv categories' (34 chars, acceptable), but 'New browser page' (15 chars), 'Close a browser page' (19 chars minimal), 'List console messages' (20 chars bare minimum). Descriptions lack context for WHEN to use them or what result they provide.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 28 | 1.25.3+ | v1 |
Get a specific console message
Get details of a specific network request
現在のページのテキストコンテンツを取得する
Get detailed information about a specific arXiv paper
Get the PDF URL for an arXiv paper
Handle a dialog (alert, confirm, prompt) in the page
Hover over an element in the page
Run a Lighthouse audit on the page
List all available arXiv categories
List console messages from the page
List all available models on OpenRouter
List network requests made by the page
List all open browser pages
指定したURLに移動する
Navigate to a specified URL in the browser
Create a new browser page
Analyze performance insights from a trace
Start a performance trace
Stop a performance trace and get results
Press a keyboard key
Resize the browser page viewport
現在のページのスクリーンショットを撮影する
Search for papers on arXiv
Select a browser page by ID
Take a memory snapshot for debugging
Take a screenshot of the current page
Take a DOM snapshot of the current page
Type text in the currently focused element
指定した要素にテキストを入力する
Upload a file to a file input element
Wait for a condition to be met
Parameters often lack detailed constraints and context. Example: 'selector' parameter in click/fill/hover tools described as 'CSS selector or text to identify the element', but LLMs don't know whether to pass a class selector (.button), ID selector (#submit), or plaintext ('Click here'). Expected format and examples are missing.
No error handling guidance. Tools like 'navigate_page', 'fill', 'evaluate_script' can fail (page not loaded, element not found, script error) but descriptions do not explain what errors mean or how to recover. LLMs have no way to know whether to retry, try a different selector, or ask for user help.
Tool descriptions in Japanese (navigate, click, type_text, get_page_content, screenshot) lack English descriptions. While multi-language support is valid, the audit assumes English descriptions for clarity. These tools are also less detailed than their English counterparts.
Performance/tracing tools (performance_start_trace, performance_stop_trace, performance_analyze_insight) lack guidance on output structure. What fields does a trace return? How do insights differ from raw trace data? LLMs cannot reason about downstream use.
Tools like 'emulate' (device and network emulation) lack schema documentation for what device types and network conditions are valid. Parameters are present but undocumented, LLMs must guess valid enum values.
'chat' and 'list_models' tools expose OpenRouter API surface but lack guidance on model selection criteria, rate limiting, cost implications, or failure modes. Tool descriptions do not explain when to use different models or what happens on auth failure.
Several tool schemas show 'properties:{}' (empty property sets) with no description of what the tool returns. 'new_page', 'list_pages', 'list_console_messages', 'list_network_requests', 'list_categories' lack documented output schemas, LLMs cannot plan downstream chaining.