AI Agent using Serena MCP, Sequential Thinking MCP, and OpenRouter for semantic code analysis and intelligent task execution
Apollo Agent exposes 22 tools with broadly consistent naming and schemas, but suffers from critical gaps in descriptions, incomplete parameter documentation, and missing error handling guidance. Descriptions average ~80 characters (below the production baseline of 194 chars), many parameter descriptions are generic or missing entirely, and output schemas are not documented. The tool set mixes disparate concerns (GitHub API, NPM/PyPI registry, Wikipedia, file system, shell execution, task management) without clear composition chains. 11 tools have read-only risk, 6 have write/destructive risk, but there is no tool annotation metadata (readOnlyHint, destructiveHint) to guide agent decision-making. Terminal tools (pwd, ls, cd, read_file, write_file, install_package, execute_shell_command) expose shell execution with minimal input validation and no explicit security guidance. Overall, the server implements basic MCP tool registration but lacks the LLM-centric optimization, detailed error recovery, and security patterns expected in production-grade agent toolkits.
Change directory
Execute a shell command with optional sudo support. Use this for system operations that may require elevated privileges.
Get detailed information about an NPM package
Get detailed information about a PyPI package
Get detailed information about a specific GitHub repository
Get the contents of a file or directory in a GitHub repository
Get the full content of a Wikipedia article
Get a summary of a Wikipedia article
Descriptions lack LLM-optimized context. Most tool descriptions (e.g., execute_shell_command, pwd, ls, cd) are under 50 characters and do not explain WHEN to use the tool, WHAT it returns, or any prerequisites. Average desc length is ~80 chars vs production baseline of 194 chars. LLMs cannot distinguish between similar tools or determine selection priority without richer guidance.
Parameter descriptions are missing or overly generic. Terminal tools (pwd, ls, cd, read_file, write_file) lack detailed parameter descriptions. The 'session' parameter is referenced across multiple tools but never defined, what is a 'session ID'? How is it created? How does it persist across calls? This ambiguity forces LLMs to guess.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Install a package using a package manager
List directory contents
Mark the current task as completed and move to the next task. Call this when you have finished working on the current task.
Print working directory
Read the contents of a file
Search for code across GitHub repositories
Search for packages in the NPM registry
Search for a package in a package manager registry
Search for packages in the PyPI registry
Search for GitHub repositories matching a query with sorting and pagination options
Search for articles on Wikipedia
Skip a task that is not applicable or cannot be completed. Provide a reason.
Start working on a specific task. Use this to explicitly mark a task as in-progress.
Write content to a file
No output schemas documented. Tools like search_repositories, get_npm_package, search_code return complex objects (arrays, nested fields, pagination info), but the spec provides no schema defining what fields are returned, their types, or required vs optional. Agents must infer structure from examples, leading to incorrect field access and context waste.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). All 22 tools are missing MCP tool annotations. The 'Risk' metadata (READ_ONLY vs WRITE vs DESTRUCTIVE) is provided in the spec but not exposed in the tool definitions. Agents cannot automatically reason about idempotence, state mutations, or safety without this metadata, forcing manual reasoning about which tools are safe to retry.
execute_shell_command lacks input validation and security guidance. The tool accepts arbitrary shell commands with optional sudo. There is no documented sanitization against command injection, path traversal, or rate limiting. The 'requiresSudo' parameter is optional and auto-detected, but sudoPassword is not a parameter (good for security), however, the tool description does not explain how sudo elevation is handled or what permissions are required.
No error handling guidance. Tools lack explicit recovery instructions. For example, if search_repositories returns 0 results, should the agent retry with different keywords? If GitHub API rate-limits the call, what should the agent do? Generic error messages ('GitHub API error: 403') do not tell the LLM what to do next.
Parameter constraints not enforced. Many parameters (e.g., 'sort' in search_repositories, 'packageManager' in install_package) accept free-form strings when they should be enums. 'sort' is documented as 'Sort field (default: stars)' but what other values are valid? 'packageManager' lists npm, pip, apt, brew as options in the description but not as an enum constraint. LLMs will hallucinate invalid values.
Tool composition is unclear. The server mixes 22 disparate concerns: GitHub/code search, package registries, Wikipedia, file system ops, shell execution, and task management. There is no documented tool chain or workflow. How does 'search_code' flow into 'get_repository_contents'? When should an agent call 'mark_task_complete' vs 'skip_task'? Agents waste reasoning cycles deciding on tool sequencing.
Terminal tool session management is opaque. The 'session' parameter is used across terminal tools to maintain state (current directory, session context). However, no tool creates a session, and the initial session ID is never specified. Agents cannot bootstrap a session, they must guess or pass an undefined session ID and hope it works.
No pagination support documented. search_repositories and search_code accept 'perPage' and return results, but no 'page' or 'offset' parameter is visible. How does the agent paginate through results? Without pagination metadata in responses (total count, next_cursor), agents cannot discover whether more results exist.