Agentic Debugging System that integrates with Git and Filesystem MCP servers. Investigates software bugs through multiple competing hypotheses, launching a mother agent that analyzes errors, generates diverse hypotheses about potential causes, and spawns isolated scenario agents to test each hypothesis in separate git branches.
The Deebo server has 3 tools with significant quality gaps. The `start` tool has a well-structured schema with 5 parameters (all properly typed and described), earning a solid schema score. However, two critical issues lower the overall: (1) Tool naming violates verb_noun convention, `start` and `check` are generic action verbs that don't clearly convey intent (should be `start_debug_session`, `get_debug_status`); (2) The `check` tool's description is extremely verbose (374 characters) and reads like internal documentation rather than LLM-optimized guidance, violating the 10-1024 character guideline and best practice of 50-200 chars; (3) The `read_deebo_guide` tool has an empty input schema (declared as `{}` in the code), which violates the requirement that all tools must declare schema structure. The `check` tool description does provide recovery hints (useful for error handling), but this is overshadowed by verbosity and lack of clarity. Composition and chaining are unclear, the server exposes raw session IDs and hypothesis/scenario abstractions that agents must reason about without clear guidance on workflow.
Retrieves the current status of a debugging session, providing a detailed pulse report. For in-progress sessions, the pulse includes the mother agent's current stage in the OODA loop, running scenario agents with their hypotheses, and any preliminary findings. For completed sessions, the pulse contains the final solution with a comprehensive explanation, relevant code changes, and outcome summaries from all scenario agents that contributed to the solution. Use this tool to monitor ongoing progress or retrieve the final validated fix. In a short paragraph, Include the Mother agents status only if it's crashed or failing otherwise just skip over it and use the last activity and the last log message to summarize in one sentence what the mother agent did. Then describe scenario agents activity and hypotheses briefly.
Reads and returns the Deebo guide documentation for users to learn about the system.
Begins an autonomous debugging session that investigates software bugs through multiple competing hypotheses. This tool launches a mother agent that analyzes errors, generates diverse hypotheses about potential causes, and spawns isolated scenario agents to test each hypothesis in separate git branches. The mother agent coordinates the investigation, evaluates scenario reports, and synthesizes a validated solution when sufficient evidence is found.
Tool naming does not follow verb_noun pattern and uses generic verbs ('start', 'check') that lack semantic clarity. LLMs cannot infer intended action from names alone. Should be 'start_debug_session', 'get_debug_status', 'read_debug_guide'.
The 'check' tool description is 374 characters (exceeds 1024 char guideline upper bound in verbose style) and reads as internal system documentation rather than LLM-optimized guidance. Best practice is 50-200 chars. Contains implementation details ('OODA loop', 'scenario agents') that confuse rather than clarify when to call the tool.
The 'read_deebo_guide' tool declares an empty input schema (`{}` in code). The tool takes no parameters, but this must be explicitly declared with proper JSON Schema structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Output schemas are not documented in tool definitions. The 'start' tool returns a sessionId (implied, not explicit), and 'check' tool returns structured session state including mother agent status and scenario reports, but no formal output schema is visible. LLMs cannot plan downstream tool calls without knowing return field types.
Parameter descriptions for 'start' tool are adequate but lack usage guidance. The 'context' parameter description does not explain when to provide it or what format is expected (code snippet, error log, test output?). The 'language' parameter description is vague ('e.g. typescript, python'), should declare this as an enum of supported languages.
No error handling guidance in tool descriptions. What happens if a sessionId is invalid in 'check'? What if the repo path does not exist in 'start'? Descriptions do not indicate retryable vs. fatal errors, forcing LLMs to guess error recovery strategy.
Tool composition unclear. The 'start' tool spawns background agents and returns a sessionId, but the description does not explicitly state that agents run asynchronously. This forces LLMs to infer whether they should immediately call 'check' or wait. Guidance needed: 'Returns a sessionId immediately; call check() to monitor progress.'