A Model Context Protocol server that dynamically loads tools from YAML configuration, supports script execution, task management, and multi-language code analysis with Neo4j graph database integration
This configuration-based server exposes 5 tools with highly variable quality. Three tools (data_frame_service, kv_store, agent) have comprehensive schemas and reasonable descriptions, but lack critical error guidance and output documentation. The get_session_image and time_tool have minimal descriptions (under 50 chars) that fail to guide LLM selection. Naming is action-verb based but some tools conflate multiple concerns (data_frame_service combines load and execute; kv_store bundles 8 operations). Parameter descriptions are present but inconsistent in depth. No tool explicitly documents output schemas, error recovery paths, or when to use one tool over similar alternatives. Overall, the server shows basic competence in schema structure but fails on the LLM-optimization and composition patterns that enable reliable agent behavior.
Invoke specialized AI agents for codebase exploration, implementation finding, structure analysis, and more
Data Frame Service - Load data from files/URLs and execute pandas operations. Supports loading from local files or URLs, storing with IDs for later operations. Operations: load_data, execute. Use execute with any pandas expression like df.head(), df.query('age > 30'), etc. Use this tool to load data once and perform multiple pandas operations efficiently.
Get an image for a given session_id and image_name.
In-memory key-value store with TTL support. Supports operations: set, get, delete, exists, keys, list, clear, ttl. Default TTL is 1 day (86400 seconds).
Returns a time string. If 'time_point' is missing or 'now', returns the current time. If 'delta' is provided, shifts the time_point by the delta (e.g., '5m', '2d', '3h', '42s'). Always respects the timezone (default 'UTC').
No tool documents its output schema. LLMs cannot plan downstream calls without knowing what fields to expect.
get_session_image description is 39 characters, below the 50-character minimum. Lacks context on when to use, what happens on failure, and what format is returned.
No tool provides error recovery guidance. Tools lack documentation of failure modes and what the LLM should do next (retry, ask user, try alternative tool).
data_frame_service and kv_store conflate multiple operations under a single tool (load+execute, and 8 KV operations). Each tool should do one thing; split into separate tools (load_dataframe, execute_pandas_expression; set_cache, get_cache, etc.).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 32 | - | v1 |
data_frame_service parameter 'csv_options' is typed as 'object' with no property schema. LLMs cannot determine valid CSV options (sep, header, encoding, etc.).
time_tool delta parameter format is vague ('5m', '2d', '3h', '42s'). No regex pattern or examples of invalid input. LLMs may pass '5 minutes', 'five minutes', or other variations that fail.
agent tool's context.model parameter is undocumented. Is 'gpt-4' a valid value? Does it cost more? When should an LLM choose 'haiku' vs 'gpt-4'?
kv_store and data_frame_service use operation enums, which requires LLMs to understand dispatch logic. Clearer to split into separate tools (get_cache, set_cache, delete_cache) so names convey intent directly.