A custom MCP-like server that dynamically loads and executes registered functions, with semantic search capabilities powered by embeddings and vector storage (Qdrant or in-memory).
This server has 4 tools with basic schemas and descriptions, but falls significantly short of production quality. Tool naming lacks clarity and verb-noun conventions. Descriptions are present but minimal (10-60 chars), well below the 50-200 char ideal for LLM optimization. Parameter descriptions are mostly absent, tools like `googleSearch` and `extractPageContent` have only the top-level input documented, with no guidance on expected formats or constraints. Output schemas are entirely undocumented; LLMs cannot plan downstream chains without knowing what fields to expect. Error handling is minimal, no recovery guidance, categorization, or actionable messages visible in source. The server lacks output field documentation (critical per pattern:tool and mxe:strip-api-responses), making tool composition difficult. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite all tools being READ_ONLY. Security: credentials are likely leaked (google-sr and firecrawl dependencies suggest API keys in env, but no verification of secret injection pattern in tool layer). Overall, this reads like an early prototype rather than a production tool system.
Adds two numbers together
Extracts the main content from a given webpage URL
Performs a Google search and returns relevant results
Returns the number of CPU cores available on the server.
Output schemas completely undocumented. No tool returns a schema describing what fields LLMs can expect. LLMs cannot plan downstream tool chains or extract needed data without guessing output structure.
Parameter descriptions missing or minimal. Only top-level input documented; no guidance on expected formats, ranges, enums, or constraints. LLMs must guess valid values.
Tool naming inconsistency and non-standard format. All tools use camelCase (e.g., 'serverCPUCount', 'googleSearch') instead of snake_case verb_noun convention (e.g., 'get_server_cpu_count', 'search_google'). LLMs have harder time parsing intent from non-standard names.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 38 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 14 | - | v1 |
No error handling or recovery guidance documented. No indication of which errors are retryable, user-fixable, or fatal. Tool handlers catch and return error.message but LLM receives no guidance on what to do next.
Tool annotations missing. All tools are READ_ONLY (per risk assessment) but no readOnlyHint or idempotentHint applied. LLMs cannot reason about retry safety or side effects without annotations.
Potential credential exposure. Dependencies on 'google-sr' and '@mendable/firecrawl-js' suggest external API keys. No visible server-side secret injection pattern in tool definitions, credentials may leak into logs or LLM context.
No input validation or constrained parameters. Parameters like `url` and `query` accept free-form strings with no length limits, pattern validation, or enum constraints. LLMs can pass invalid, malicious, or absurd values.
No pagination or result limits documented. 'googleSearch' returns 'relevant results' with no specification of count, order, or pagination mechanism. Large result sets could exhaust context window. No mention of rate limiting.