Local MCP server with custom tools for CarBench, MedAgentBench, and TAU2 datasets
This MCP server exposes 13 tools across multiple domains (CarBench, tau2 airlines, MedAgentBench). Critical issues: (1) Tool naming lacks consistency and clarity, many names are either overly generic (get_vehicle_ctx, report_error_statistics) or lack action verbs (compute_reservation_price, convert_route_distance_and_time). (2) Descriptions vary widely in quality: some are acceptable (e.g., get_routes_from_start_to_destination has 195 chars with clear purpose), but others are severely underdeveloped (get_vehicle_ctx has only 34 chars with no context). (3) Input schemas are present for all tools but parameter documentation is incomplete, many parameters lack descriptions or have minimal clarity. For example, compute_reservation_price accepts 'flights' and 'passengers' as arrays of untyped objects with no schema details. (4) Output schemas are not explicitly documented in any tool definition. (5) Error handling is largely absent, no recovery guidance, categorization, or actionable error messages visible in tool implementations. (6) The server mixes internal debugging tools (save_state, load_state, report_error_statistics, get_tool_execution_errors_during_runtime) with domain-specific tools, creating a confused interface. These tools are marked 'disclose_to_model: False' but still clutter the public API. (7) Security concerns: file I/O in save_state/load_state accept user-supplied paths with no validation or sanitization, risking path traversal attacks.
Compute the total price of a reservation without booking it.
Compute the difference between two times. If time2 is later than time1, the result is positive; otherwise, it is negative.
Helper Tool: converts distance (in kilometer) into time (minutes) needed along specific route and vice versa.
Fetch the current system time
Search for patients in the FHIR server based on various criteria.
Search for patients in the FHIR server based on various criteria with additional safeguards for purpose validation and response filtering.
Internal debugging tools exposed in public API. save_state, load_state, get_vehicle_ctx, report_error_statistics, and get_tool_execution_errors_during_runtime are marked 'disclose_to_model: False' in metadata but still appear as callable tools. This clutters the agent interface and violates separation of concerns.
Critical security vulnerability: file I/O in save_state and load_state accepts user-supplied file paths with no validation or sanitization. An attacker-controlled agent could perform path traversal attacks (e.g., path='../../../../etc/passwd') to read or write arbitrary files on the server.
Duplicate tools with overlapping functionality and no clear distinction. get_patient and get_patient_extended both search for patients, LLMs cannot easily choose between them. The 'extended' suffix is vague; the difference (purpose validation, response filtering) should be explicit in separate, non-overlapping tool names or consolidated into a single tool with optional parameters.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 20 | - | v1 |
Routes information: gets the fastest route (plus alternative routes if existent) for the car between start and destination. Each route information includes name_via, distance in km, duration in hours and minutes, arrival time, road types (highway, urban, country roads, includes toll roads), and an route alias (first, second, third; additionaly fastest, shortest). Routes can be requested between locations or between a location and a point of interest.
Retrieves tool execution errors that occurred during runtime
Retrieves the current vehicle context state
Loads vehicle context and fixed context state from a file
Reports error statistics for tool execution failures
Saves the current vehicle context and fixed context state to a file
Points of Interest Search: Searches for points of interest of a specified category that are geometrically close to a given route. Calculates the detour distance and time for up to 3 POIs with the smallest detour. Information includes name, position, detour, opening hours, and phone number.
Tool descriptions are inconsistent in length and quality. Baseline is 194 chars (p10=34, p90=392). Examples: get_vehicle_ctx (34 chars, far below p10), compute_time_difference (153 chars, below baseline). Short descriptions lack context on when or why to call the tool, forcing LLMs to infer intent.
Input parameters lack complete documentation. compute_reservation_price declares 'flights' and 'passengers' as arrays of untyped objects {'type':'object'} with no schema or field descriptions. LLMs cannot know what fields to include in each object. This violates pattern:constrained-input and forces hallucinated payloads.
Union types in parameter schemas create ambiguity. get_patient and get_patient_extended declare patient_id, family, given, name, gender, address, and telecom as {'type': ['string', 'object']}. It is unclear what an 'object' value means here or when an LLM should pass one vs a string. This violates schema best practices and increases error likelihood.
Tool naming lacks consistency and clarity. Example: convert_route_distance_and_time contains 'and', signaling multiple responsibilities (should split into two tools). report_error_statistics lacks a clear action verb. get_vehicle_ctx is vague ('ctx' is ambiguous). Several tools (compute_reservation_price, compute_time_difference) use 'compute' instead of standard verbs like 'calculate' or domain-specific verbs.
No output schemas documented for any tool. LLMs cannot plan downstream tool calls or extract the right data without knowing what fields to expect in responses. This is a critical gap across all 13 tools.
No error handling or recovery guidance visible in tool implementations. Errors are not categorized as retryable, user-fixable, or fatal. No error messages guide the LLM toward recovery. Example: if a patient search returns no results, the tool should suggest calling a discovery tool or clarifying search criteria.
No parameter constraints for numeric inputs. compute_reservation_price accepts total_baggages and nonfree_baggages as integers with no min/max bounds. convert_route_distance_and_time accepts route_id, time_minutes, and distance_km with no constraints. Unbounded inputs let LLMs pass absurd values.
Example values embedded in descriptions (pattern:tool-description violation). compute_reservation_price includes example user_id 'sara_doe_496' in description, LLMs tend to reuse these literally. compute_time_difference includes ISO 8601 example timestamps, similar risk. Best practice: move examples to parameter constraints (enums, patterns) or separate documentation.