A multi-agent retrieval-augmented generation (RAG) system that routes queries to specialized agents: web search (Airbnb, Tavily, Wikipedia), document search (vector store), and geography knowledge (IP geolocation, geocoding, distance calculation)
This server exhibits pervasive gaps across naming, descriptions, schemas, and error handling. Of 7 declared tools, only 4 have explicit parameter schemas visible in the source code. The native tools (document_tool, ip_to_location, get_coordinate, haversine) have inconsistent quality: document_tool has a reasonable description but the schema lacks type info for the output; ip_to_location and haversine lack proper parameter descriptions in their schema views; get_coordinate lacks formal parameter schema constraints. No tool has output schema documentation. No tool implements error guidance (e.g., 'If query returns no results, try...'), and several tools lack actionable error messages. Parameter descriptions are minimal or missing entirely. The codebase shows no evidence of secret injection patterns, rate limiting, or permission gates. This is a rapid prototype lacking production-grade tool definitions.
Tools from the Airbnb MCP server
Retrieves relevant information from the document store based on a user's query.
Retrieve the coordinates (latitude and longitude) of a given city or settlement.
Calculate the great-circle (straight line) distance between two coordinates using the Haversine formula.
Convert an IP address to its geographical location information
Web search tools from Tavily MCP server
Tools from the Wikipedia MCP server
Three tools (airbnb_mcp_tools, tavily_mcp_tools, wikipedia_mcp_tools) are proxied from external MCP servers with no schema definitions or descriptions visible in this codebase. They are inferred imports only, violating the requirement to define tools explicitly with full schemas and descriptions.
ip_to_location parameter 'ip' has a description in the HTTP Query annotation but the schema definition does not include per-parameter type constraints or validation rules. The endpoint returns raw JSON from an external API without stripping irrelevant fields, risking context bloat. No error guidance (e.g., 'If IP is invalid, the service returns a 400. Valid IPs are IPv4 or IPv6').
get_coordinate and haversine use async FastAPI endpoints but their input parameter descriptions are missing or minimal. get_coordinate accepts 'city: str' with no description of format (e.g., 'City name (e.g., Istanbul, Paris) or ISO 3166 code'). haversine accepts raw floats with no min/max bounds documented, an LLM could pass NaN, infinity, or values outside [-180, 180] for longitude.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 32 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 47 | - | v1 |
No tool documents its output schema. LLMs expect to know: What fields will I get back? What type? document_tool returns a string, fine. But ip_to_location returns raw JSON from CollectAPI (containing 'status', 'result', etc.), the LLM has no schema to guide downstream parsing. get_coordinate returns [lat, lon] as a list, an LLM cannot infer it will receive exactly two floats in that order.
Error handling is absent or minimal. ip_to_location raises HTTPException with raw API errors ('CollectAPI failed: {response.text}'). No guidance on retryability or what the LLM should do next. get_coordinate will raise an exception if geolocator.geocode() returns None (city not found), the LLM gets a 500 error with no suggestions. haversine has no validation of lat/lon bounds; invalid input silently produces wrong distances.
API keys are hardcoded or environment-sourced in the code (TAVILY_API_KEY in tools.py, API_KEY in mcp_server.py). While the second uses env vars (better), there is no evidence of secret injection at the MCP transport layer. If an agent logs or traces a call to ip_to_location, the authorization header could leak.
document_tool returns a large concatenated string of document content without pagination, limits, or structure. No indication of how many documents were returned, which are most relevant, or how to refine the query. If vector_store.as_retriever() returns 10 documents, the concatenated output could be thousands of tokens, wastes context and risks hallucination.
Tool naming: 'document_tool' is generic and vague. Better names: 'search_documents' or 'retrieve_documents'. 'ip_to_location' is clear, but 'get_coordinate' could be 'search_coordinates_by_city' to distinguish it from 'haversine'. Generic names like 'tool' force LLMs to read descriptions to disambiguate.