A Streamlit-based chatbot client that uses the Model Context Protocol (MCP) to connect to a Tavily web search tool via a remote MCP server. Implements a ReAct agent pattern with LangGraph for autonomous web research and reasoning.
This server exposes a single tool 'web_search' with minimal definition quality. The tool has a basic description (67 chars) that explains purpose but lacks critical details about what happens during execution, constraints on query length/format, pagination behavior, rate limits, and error recovery. The input schema is present but minimal, only a 'query' string parameter with a basic description. No output schema is documented anywhere in the source code, making it impossible for LLMs to know what fields to expect in the response. The implementation uses a remote MCP server (Tavily) via HTTP Streamable protocol, but the client-side tool wrapper provides no enrichment, validation, or error handling. The tool is READ_ONLY (good) but lacks any annotations (readOnlyHint missing). No pagination limit is enforced, and the description does not mention result limits or pagination behavior, which is critical for a search tool that could return hundreds of results. The code snippet shows tool integration via 'langchain_mcp_adapters' but does not reveal the actual Tavily MCP server implementation or schema details, only that it is proxied through an HTTP endpoint.
Web search tool provided through MCP/Tavily that returns a JSON object containing a preliminary answer summary and a list of raw sources with URLs and content excerpts.
No output schema documented. LLMs cannot predict what fields the search result contains (summary, sources, URLs, confidence scores, etc.), forcing them to reason about unstructured responses and waste tokens parsing.
Input schema lacks constraints on query length, format, or allowed characters. No minimum/maximum bounds. LLMs can pass empty strings or 10,000-char queries without validation.
Tool description (67 chars) is incomplete. Does not explain: what the response contains, when it returns a summary vs raw sources, pagination behavior, rate limits, or error conditions (e.g., when Tavily times out or hits rate limits).
No pagination guidance in description or schema. A search tool returning unlimited results can blow the context window. Baseline: 20 - 50 items with next_cursor or limit parameter.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 27 | - | v1 |
Missing error handling and recovery guidance. Description does not say: what happens if Tavily is unreachable, if query is invalid, if rate limits are hit. No actionable error messages for LLM recovery.
Tool lacks readOnlyHint annotation. Although marked as READ_ONLY in metadata, MCP tool annotation (destructiveHint=false, readOnlyHint=true, idempotentHint=true) is missing, reducing LLM discoverability and safety classification.
Parameter description is generic ('The search query string') and does not hint at format, length constraints, or how natural language queries are handled. Baseline: 'A natural-language search query (1 - 500 characters). Examples: "latest AI breakthroughs", "Python async best practices".', but WITHOUT literal examples in description, use enums/patterns instead.