REST API + MCP Server giving AI agents real-world superpowers — web search, content extraction, screenshots, weather, finance, email validation, and translation
The server provides 15 tools (with significant duplication across integrations) with basic schemas and descriptions, but falls short of production quality in multiple critical areas. Tool naming is inconsistent (mix of verb-noun and prefixed variants), descriptions are present but lack LLM-optimized depth, parameter documentation is incomplete, and error handling is not visible in the provided code. The duplication of tools across langchain and llamaindex integrations (tools 6-15 are near-identical copies with minor naming variations) suggests poor maintenance. No evidence of output schema documentation, input validation guidance, or recovery-oriented error handling. The STDIO transport and missing protocol features further compound the quality assessment.
Extract the main readable content from a web page URL. Returns clean text/markdown, removing ads and navigation.
Extract the main readable content from a web page URL. Returns clean text/markdown, removing ads and navigation.
Get real-time stock quotes or currency exchange rates. For stocks: provide symbol. For exchange: provide from_currency, to_currency, amount.
Get real-time stock quotes or currency exchange rates. For stocks: provide symbol (e.g. 'AAPL'). For exchange rates: provide from, to, and amount.
Capture a screenshot of a web page. Returns base64-encoded PNG image data.
Capture a screenshot of a web page. Returns base64-encoded PNG. Provide url and optional width/height.
Critical: 10 duplicate tools across llamaindex integration (tools 6-15 are near-identical copies of tools 1-5 with prefixed names). This violates the composition pattern of avoiding multiple tools that do the same thing, LLMs waste reasoning cycles deciding between search vs agent_toolbox_search.
High: Tool naming is inconsistent. Early tools use simple verbs (search, extract, weather, finance, screenshot), but later integrations add prefixes (agent_toolbox_*), making it unclear to LLMs whether they are different tools or aliases. Naming convention should be unified across all integrations.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Search the web and get titles, URLs, and snippets. Useful for finding up-to-date information on any topic.
Search the web and get titles, URLs, and snippets. Useful for finding up-to-date information on any topic.
Get current weather conditions and forecast for any location. Returns temperature, humidity, wind, and multi-day forecast.
Get current weather conditions and forecast for any location. Returns temperature, humidity, wind, and multi-day forecast.
Extract content from a web page using Readability
Get stock quotes, history, or exchange rates
Take a screenshot of a web page
Search the web using DuckDuckGo
Get weather forecast for a location
High: Parameter descriptions are minimal or missing constraint details. E.g., 'weather' tool accepts 'location', 'lat', 'lon' but does not clarify: are all three required? Can location and lat/lon be mixed? What happens if conflicting parameters are provided? 'finance' tool has 'from' and 'to' parameters with no indication they are for currency, not stock operation modes.
High: No documented output schemas visible in provided code. Tool descriptions state what they return (e.g., 'returns base64-encoded PNG', 'returns clean text/markdown') but no structured schema showing field names, types, or required fields. LLMs cannot plan downstream tool chains without knowing what data is available.
High: No error handling or recovery guidance visible. Tool descriptions do not mention: what errors can occur, are they retryable, should the LLM try an alternative tool, what information should the user provide to fix the error? E.g., 'weather' tool gives no guidance if a location is not found or if coordinates are invalid.
Medium: Parameter interdependencies are undocumented. 'finance' tool shows 'symbol', 'from_currency', 'to_currency', and 'amount' params, but does not explain: which combination is valid? Is 'symbol' required for stock quotes and forbidden for currency exchange? Does 'type' enum value (quote vs exchange) determine which params apply?
Medium: Descriptions are present but lack LLM-optimization. They range 35 - 75 chars (vs production baseline of 194 chars avg). E.g., 'weather' description is 'Get weather forecast for a location', it does not explain WHEN to use it vs a search, WHAT FORMAT is returned, or any prerequisites. Production descriptions typically include use case and return type hints.
Medium: 'screenshot' and 'agent_toolbox_screenshot' tools accept width/height with min/max bounds (320 - 1920, 240 - 1080) but do not document WHY these bounds exist or what happens if limits are exceeded. LLMs cannot reason about constraints without context.
Medium: No indication of tool idempotency. 'screenshot' and 'extract' appear safe to retry, but there is no explicit documentation. 'search' may have pagination implications on retry. Production tools must declare idempotency so agents can safely retry on ambiguous failures.