MCP server that turns natural language into API calls via RAG over OpenAPI schemas
YellowPages has 2 tools with reasonable naming and structure, but critical gaps in parameter descriptions, output schema documentation, and error handling guidance. The tools are explicitly registered in src/mcp/server.py using @mcp.tool() decorators with docstrings. However, the input schemas lack detailed parameter descriptions (path_params, query_params, body_data are documented only generically as 'object'). Output schemas are not documented at all, callers have no visibility into what fields the API responses contain. Error handling is minimal; no recovery guidance is provided. The design pattern (discover → execute) is sound, but implementation lacks the rigor expected for production tooling.
Discover which API operations match a natural-language query. Uses RAG over the loaded OpenAPI schema; returns a list of operation entries (operation name, method, url, parameters with name/type/required). Use this to find tool IDs, then call execute_operation with the chosen operation_name and params. No execution happens here.
Execute one API operation by name. Use operation_name from discover_operations; provide path_params, query_params, and optionally body_data as required by the schema. Returns the API response body or an error string.
Output schemas undocumented. discover_operations returns list[dict] and execute_operation returns str, but the structure and fields of these responses are not documented. Callers cannot infer what fields to extract from the discover response or how to parse execute response.
Path/query/body parameters undescribed. execute_operation accepts path_params, query_params, and body_data as bare dicts with minimal guidance on what they should contain, what fields are required, or what formats are expected. This forces the LLM to guess based on the operation schema, increasing errors.
No error recovery guidance. Both tools lack descriptions of how errors are returned, what error conditions might occur, or how to handle them. For example, execute_operation 'Returns the API response body or an error string', but what does an error string look like? How should the LLM interpret it?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
discover_operations default parameter (k=5) not justified. No description explains why 5 is the default, when to increase it, or what happens if k exceeds available operations. A parameter description should cover expected range and rationale.
No input validation or constraint documentation. execute_operation accepts operation_name as a string with no enum of valid names or guidance on how to obtain them. Callers must already have the operation_name from discover_operations, but this dependency is mentioned in prose only, not enforced.
Tool composition assumes LLM state. The pattern 'call discover_operations, then execute_operation with the operation_name' is described in docstrings, but the tools do not enforce or validate this flow. An LLM could pass an invalid operation_name to execute_operation, leading to a cryptic API error.