MCP server for semantic search over resumes using vector embeddings and PostgreSQL
Single tool with clear naming and present schema, but descriptions lack depth and lack critical LLM-optimization guidance. The tool name 'search_resumes' correctly uses verb_noun convention and is action-oriented. The input schema is present with types (string, integer) and basic descriptions. However, descriptions are functional but minimal (72 chars for tool, 10-40 chars per param), falling below the 50-200 char ideal for LLM optimization. No output schema is documented, the return type is inferred as list[str] from code but not formally described in the tool definition. Parameter descriptions lack dependency hints, constraint details, and usage guidance. Missing: error handling guidance, output schema structure, pagination advice for potentially large result sets, and examples of when/why to use this tool over alternatives.
Perform semantic search over resumes. - query: natural language search string - collection_name: which collection/tenant to search - top_k: number of results to return
Tool description lacks LLM-optimization depth. At 72 characters, it is below the 194-char baseline for production tools. Descriptions should explain WHAT (semantic search), WHEN (looking for candidates matching criteria), and WHAT IT RETURNS (matched resume content).
Parameter descriptions are minimal (10-40 chars each). 'which collection/tenant to search' and 'natural language search string' lack constraint details, examples of valid formats, or usage context. LLMs cannot infer whether collection_name is a UUID, slugified name, or numeric ID.
Output schema is not documented in the tool definition. Code shows return type as list[str] (resume content), but there is no formal schema describing the structure of each returned document (e.g., does each string contain metadata like source file, score, author?). LLMs cannot plan downstream tool calls without knowing what fields are returned.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
No error handling guidance. If the collection does not exist, if the embedding API fails, or if no results match the query, the tool provides no recovery instructions. LLMs need actionable error messages.
No pagination or result limits documented. The tool accepts top_k parameter (good) but the tool description does not clarify the default (5) or maximum advisable limit. If an LLM passes top_k=1000, it could blow the context window or timeout the embedding service.
Parameter 'collection_name' lacks guidance on valid values. Is it an enum? A regex pattern? Should the LLM know valid collection names in advance, or is discovery required? This ambiguity forces LLMs to guess or attempt invalid values.
No mention of dependency on external services (Google Generative AI embedding API, PostgreSQL/pgvector). If the embedding API is rate-limited or unreachable, the tool silently fails. Tool description should mention prerequisites and potential failure modes.