MCP server for AI agents to search Terraform AWS modules with hybrid search engine combining semantic similarity, BM25 text relevance, and keyword matching. FastMCP-based, CPU-only inference.
TFModSearch has three well-named, read-only tools with clear descriptions and proper input schemas. Naming follows verb_noun conventions (modules_list, search_modules, get_module). Descriptions are substantive (119 - 287 chars), exceeding the 10-char minimum and fitting the productive range (10 - 1024 chars). Input parameters have types and descriptions. However, output schemas are not documented in the source code, the server does not explicitly declare what fields each tool returns, forcing LLMs to infer structure. Error handling is not visible in the provided code sample. No security-relevant parameter validation is evident (e.g., sanitization, enum constraints). Tool composition is sound, each tool has one responsibility, and together they form a search-and-retrieve workflow. Parameter naming is consistent (e.g., 'module_identifier' is specific). No secrets are exposed as parameters. The server is read-only (no destructive operations), reducing risk. Overall, this is solid foundational work marred by the absence of documented output schemas and error recovery guidance.
Retrieve and orient on a chosen module. Accepts module name, relative doc path, or submodule address. Returns compact orientation head with description, module info, version pin, agent notes, gotchas, key features, and use cases, plus full section inventory. For modules with submodules, inlines submodule inventory with pinnable sources.
Returns the full catalog (names, paths, descriptions, keywords, module_id, latest_version) of indexed Terraform modules
Find the right module by functionality, technology, or exact module name. Returns top-ranked matches (default 3, up to 10) with name, path, keywords, description, relevance score, module_id, latest_version, and confidence verdict.
Output schemas not documented. Tools return structured data (name, path, keywords, description, relevance score, module_id, latest_version, confidence verdict, etc.) but these fields are described only in natural language in the tool descriptions. LLMs cannot reliably parse unstructured responses or plan downstream operations without a declared output schema.
No explicit error handling or recovery guidance visible. The code sample does not show what error responses look like, what conditions trigger errors, or what the LLM should do next (e.g., 'Module not found, try search_modules with a broader query'). Pattern requires error responses to guide the agent to recovery.
search_modules 'top_k' parameter lacks validation constraints. Description says 'default 3, max 10' but no enum or bounded integer constraint is declared. LLMs may pass top_k=1000, forcing server-side rejection. Constraints should be in schema, not just description text.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 39 | - | v1 |
get_module 'sections' parameter is declared as an array of strings but valid values are not enumerated. Description mentions examples ('inputs', 'examples', 'all') but no enum constraint is present. This invites hallucinated section names from the LLM.
modules_list input is declared as empty object ({}), no parameters. This is correct, but the description does not hint at the output size. If the catalog is large (162 modules implied by docs), does the tool paginate? Return a count? Truncate? This should be clarified to prevent LLM surprises.