Multi-module MCP server: Coder, Researcher, Secretary, and Agent modules for autonomous code writing, research, documentation, and task execution
This server exhibits significant gaps in definition quality. While 14 tools are defined with basic schemas and descriptions, most fall short of production standards. Tool names lack consistent action verbs (agent_exec_command, agent_tail_logs, agent_processes use 'agent_' prefix rather than verb-noun convention). Descriptions are generally adequate (ranging 40-250 chars) but several lack specificity about when to use a tool vs. alternatives. Parameter schemas are present but descriptions are often generic, e.g., 'Module name for log filtering' does not explain valid values or format expectations. Output schemas are not documented in the visible source code. Error handling is minimal: there are no documented recovery paths, error classifications, or actionable messages. No tool provides confirmation steps for destructive operations. Three tools (agent_exec_command, coder_simple_task) perform state mutations but lack dry-run capability or clear idempotency guarantees. The web search and research tools (web_search, deep_research, arxiv_search, fact_check) accept generic query strings with no constraints documented, inviting hallucinated or malformed queries. Overall, this reads as a functional but not polished tool suite; it would not pass a principal engineer's code review for production use.
Analyze a codebase with direct file reads + AST/grep heuristics. No secretary, no network: walks repo_root with include_patterns, optionally narrows to focus matches, and summarizes Python files via ast (functions/classes/lines/syntax errors).
Run a guarded, non-interactive shell command via the runner.
Summarize pending background jobs/tasks (read-only).
Snapshot daemon statuses plus host resource stats (read-only).
Own heuristic static review (never writes).
Return capped, redacted log tails via the runner.
Analyse a file from the codebase.
No output schemas documented for any tool. LLMs cannot predict result structure, forcing them to guess field names, types, and nesting, leading to parsing errors and wasted context.
Destructive operations (agent_exec_command with allow_write=true, coder_simple_task) lack dry-run modes or confirmation steps. Agents can accidentally execute harmful commands or rewrite files without reversibility.
Tool naming inconsistency: 'agent_*' prefix (agent_exec_command, agent_processes) does not follow verb_noun convention (run_command, get_processes). British spelling 'analyse_file' is inconsistent with American convention used elsewhere.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | 2024-11-05+ | v1 |
Search arXiv for academic papers.
Generate a comprehensive report of the codebase structure.
Execute a simple single-pass CODE WRITING task. This tool runs the AI code CLI in quick mode for fast code writing. The CLI has full responsibility for reading/writing files. Returns ONLY a concise summary - NO source code is returned.
Perform deep research on a topic using multiple queries.
Fact-check claims against web sources.
Search for files in the codebase matching patterns.
Search the web for information.
Parameter descriptions lack actionable constraints. 'focus' (agent_analyze, agent_review) is vague; 'queries' (deep_research) has no element type; 'search_provider' (web_search) lists example but not enum; 'max_results' tools have no min/max bounds. LLMs will pass invalid values.
No error handling or recovery guidance documented. Tools do not explain how to recover from failures (network timeout, invalid input, permission denied), leaving LLMs unable to decide whether to retry or escalate.
No pagination or result limits documented for tools that could return large datasets (codebase_report, file_search, agent_analyze, agent_processes). LLM may receive results that exceed context window, causing token exhaustion and hallucination.
Parameter descriptions are often one-liners that do not explain format, range, or valid values. Example: 'Session identifier' for session_id (agent_tail_logs) does not explain where to get a session or what format it takes.
Tool descriptions do not explain when to select one tool over a similar alternative. web_search vs. deep_research; agent_analyze vs. agent_review, no guidance on which to use in which scenario.