MCP server for comparing flight prices across booking sites and discovering best deals. Extracts flight details from screenshots or text and searches multiple booking platforms.
This server presents a mixed quality picture. While search_flights has a detailed input schema with nested objects and comprehensive field descriptions, it has structural issues: the descriptions are verbose (>1024 chars), parameter dependencies are underdocumented, and error handling guidance is missing. The two internal tools (submit_session, get_session_results) are explicitly marked [INTERNAL] but still exposed, creating confusion about intended usage. None of the tools include output schema documentation, making it impossible for LLMs to plan downstream actions. The server provides no tool annotations (readOnlyHint, destructiveHint, idempotentHint), no error classification, and no recovery guidance. Parameter naming is somewhat inconsistent: 'infantsInSeat' vs 'infantsOnLap' use camelCase while 'departureAirport' also uses camelCase, but this contrasts with snake_case conventions in the broader ecosystem. The fallback parsing logic in http-server.js reveals that the tool is fragile and depends on Gemini AI for natural language extraction, not robust for production use.
[INTERNAL] Poll for session results. Use search_flights instead for end-to-end flight search with auto-polling.
Compare flight prices across booking sites to find the best deal. Extracts flight details from user-provided screenshots or text and searches multiple booking platforms. Returns ranked results showing prices from different websites. Note: Number of adults defaults to 1, currency is auto-detected from price symbols (€ → EUR, $ → USD), and location is set automatically.
[INTERNAL] Create a price discovery session. Use search_flights instead for end-to-end flight search.
Output schemas not documented. LLMs cannot predict what search_flights, submit_session, or get_session_results return, making it impossible to plan chains of tool calls or extract IDs for downstream operations.
Internal tools exposed as public. submit_session and get_session_results are marked [INTERNAL] in descriptions but are registered as callable tools. This violates the single-responsibility principle and creates confusion about which tool to use (search_flights vs submit_session + get_session_results polling).
Tool descriptions exceed recommended length. search_flights description is ~500 chars; includes implementation details (Gemini extraction, auto-detection) not relevant to LLM decision-making. Rule: keep descriptions 10 - 1024 chars, ideal 50 - 200.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 39 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 8 | - | v1 |
Parameter dependencies undocumented. The schema shows 'departureTime' and 'arrivalTime' have hints like 'ASK the user if not provided', but no formal constraint or description of when these are required vs optional. 'plusDays' depends on the arrival/departure date pair but this dependency is not stated.
No error handling or recovery guidance in any tool. Tool descriptions do not explain failure modes (e.g., 'If flight not found on any site, returns empty list') or what the LLM should do if a call fails (e.g., 'Retry with different dates or airports').
No tool annotations present. search_flights is READ_ONLY (correct risk classification), but it lacks tool annotations (readOnlyHint/destructiveHint/idempotentHint) in the MCP protocol layer. Modern MCP servers should declare these via tool metadata.
Inconsistent field naming conventions. Parameters use camelCase (infantsInSeat, departureAirport) rather than snake_case, which deviates from MCP SDK conventions and may confuse LLMs trained on snake_case tool interfaces.