Reduce Claude Desktop consumption by 10x — integrate external AI models (GLM-5, Gemini 3.x) via MCP for intelligent task delegation
The server defines 10 tools (5 unique tools replicated across GLM-5 and Gemini backends). Tool naming follows verb_noun pattern (ask_*, web_*, parse_*), which is good. All tools have descriptions (194-char baseline met). However, there are significant gaps: (1) NO output schemas documented for any tool, only input schemas are visible. This is a critical omission per the rubric: 'Document the output schema. LLMs need to know what fields to expect.' (2) Error handling is implicit; no tool descriptions explain recovery paths or error categories. (3) Parameter descriptions are present and reasonably detailed (72-char baseline mostly met), but lack explicit constraints on some parameters (e.g., no min/max on timeout, no regex patterns on URLs). (4) No tool annotations (readOnlyHint, destructiveHint) despite all tools being read-only, this would improve clarity. (5) Duplication: web_search and web_reader appear twice (once in GLM-5 config, once in Gemini config), creating confusion about which variant to use. Average definition quality is fair; present but incomplete.
Delegate tasks to Google Gemini 3 Flash for general analysis, synthesis, summarization, and reasoning tasks. Fast and cost-effective for most delegation needs.
Delegate to Google Gemini 3 Pro for complex reasoning, code generation, architecture design, and demanding cognitive tasks. Most capable model.
Delegate tasks to GLM-5 (Z.ai's flagship 744B parameter model). Use this for: complex reasoning, advanced analysis, system design, and demanding cognitive tasks. Best overall performance.
Delegate to GLM-5 (Z.ai's flagship 744B parameter model) with coding-optimized system prompt. Use this for: code generation, programming tasks, refactoring, debugging, and technical implementation. Optimized for software development.
Extract text from documents, images, and PDFs using Gemini's multimodal capabilities. Handles complex layouts, tables, multi-column text. Use for PDF proposals/contracts, scanned documents, invoices. Supports up to 20MB files.
Output schemas are completely undocumented. The source code shows input schemas for all tools, but no return types or output field structures are defined. This violates the critical pattern: 'Document the output schema. LLMs need to know what fields to expect so they can plan downstream tool calls and extract the right data.' Without output documentation, agents cannot predict what data is available for chaining or composition.
No error handling guidance in tool descriptions. Tools do not explain retryable vs fatal errors, or what the agent should do if a call fails (e.g., 'If URL fetch times out, retry with a shorter timeout parameter' or 'If parse fails, check URL is publicly accessible'). This forces agents to guess recovery strategies.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Extract text from documents, images, and PDFs using GLM-OCR. Handles complex layouts, tables, multi-column text. Use for: PDF proposals/contracts, business cards, competitor materials, scanned documents, invoices. Delegate when: user uploads PDF/image for text extraction, document analysis needed (>1 page), or OCR parsing required. Supports up to 50MB files or 100 pages.
Fetch and parse full content from a specific URL. Use this to read articles, blog posts, documentation, competitor pages, or any web content after finding it via web_search. Returns markdown-formatted content ready for analysis. Delegate when: user provides a URL to read, needs full article/page content, or wants to analyze specific web pages.
Fetch and parse full content from a specific URL using Gemini's URL context capability. Returns markdown-formatted content ready for analysis. Use to read articles, documentation, competitor pages.
LLM-optimized web search using Gemini with Google Search grounding. Returns structured search results with titles, URLs, and summaries. Use for competitive intelligence, market research, and real-time information.
LLM-optimized web search for competitive intelligence, market research, and real-time information. Delegate when: user asks for competitive research, market trends, company background, content sources (>3), or real-time news. Returns structured summaries ready for analysis.
Duplicate tool definitions (web_search and web_reader appear in both GLM-5 and Gemini backends). The tools have identical schemas and descriptions but are registered separately. This increases cognitive load for agents deciding which variant to invoke and violates the composition pattern: 'Avoid multiple tools that do the same thing differently. If search_users and find_users both exist, the LLM wastes reasoning cycles deciding between them.'
Missing parameter constraints on some fields. For example: 'timeout' (in web_reader) has no min/max; 'url' has no URL format validation; 'file_url' has no size limit enforcement in schema. Per the rubric: 'Specify minimum and maximum for numeric parameters... Unbounded numbers let LLMs pass absurd values.'
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite all tools being read-only. Modern MCP servers should declare these hints to help agent planners reason about side effects and retry safety.
Temperature and max_tokens parameters lack validation guidance. Descriptions say '0.0-1.0' for temperature and '4000/8192/16384' for max_tokens, but do not explain what happens if an LLM passes 1.5 or -1 for temperature.