An MCP server providing debugger control tools using GDB/pwndbg via the pygdbmi library
This GDB MCP server has 15 tools with basic structure but significant gaps in definition quality. All tools have names starting with action verbs (execute, set_*, get_*, list_*, delete_*, toggle_), which is positive. However, most tool descriptions are extremely brief (10-40 chars), falling well below the 50-200 char LLM-optimized baseline. Parameter descriptions are minimal or absent. Input schemas are present but lack type information for several parameters (e.g., 'location' and 'condition' in set_breakpoint have no type constraints). Output schemas are completely undocumented, the code returns Dict[str, Any] with no structured schema definition visible. Error handling exists via @catch_errors decorator but provides only generic error dicts without recovery guidance. The server uses stateless HTTP (FastMCP) which is protocol-current, but tool definitions fall short of production quality.
Delete a breakpoint by number.
Disassemble the specified address.
Execute arbitrary GDB/pwndbg command.
Run until the current function returns.
Get debugging context (registers, stack, disassembly, code, backtrace).
Read memory at the specified address.
Get current debugging session information.
Tool descriptions are critically short (10-40 characters), far below the LLM-optimized 50-200 char baseline. Examples: 'Run until the current function returns' (finish, 44 chars), 'Get current debugging session information' (get_session_info, 41 chars). Descriptions lack context on WHEN to call the tool, WHAT prerequisites are needed, and WHAT is returned. This violates pattern:tool-description and pattern:command-tool.
Output schemas are completely undocumented. All tools return Dict[str, Any] with no schema definition visible. The code shows generic error responses like {'success': False, 'error': str(e), 'type': type(e).__name__} but success cases are opaque. LLMs cannot plan downstream calls or extract required fields without documented schemas. This violates pattern:tool (100% of A+ tools have documented return types).
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 49 | - | v1 |
List all breakpoints.
Run the loaded binary. Requires at least one enabled breakpoint to be set before running.
Set a breakpoint at the specified location.
Load a binary file for debugging.
Use `set args poc_file` to set the proof-of-concept (PoC) file for the loaded binary.
Execute stepping commands (continue, next, step, nexti, stepi).
Connect to a remote GDB server for debugging.
Toggle a breakpoint's state.
Parameter descriptions are missing or trivial for most tools. Examples: 'location' in set_breakpoint has description 'Address or symbol for breakpoint' (40 chars) but lacks format guidance (hex prefix? symbol names only?). 'condition' has no format guidance for GDB conditional syntax. 'context_type' in get_context says 'Type of context (all, regs, stack, disasm, code, backtrace)' but does not explain what each type returns. This violates pattern:tool-description (100% of A+ params have descriptions) and pattern:constrained-input (enums should be explicit, not in description).
Enum parameters are documented as free-form strings in descriptions rather than declared as JSON Schema enums. 'context_type' accepts 'all|regs|stack|disasm|code|backtrace' but the schema does not enforce this, an LLM could pass 'registers' or 'memory' and get a generic error. 'step_control' command expects 'c|n|s|ni|si' but is a bare string. This violates pattern:constrained-input (enums prevent hallucinated values and are self-documenting).
Error handling is generic and does not guide LLM recovery. The @catch_errors decorator returns {'success': False, 'error': str(e), 'type': type(e).__name__} for all failures. This does not tell the LLM whether the error is retryable, user-fixable, or fatal. Example: if set_file() fails with 'File not found', the error should suggest 'Try checking the file path or calling list_files() first.' Current code offers no guidance. This violates pattern:recovery-guide and pattern:error-classification.
Session-based state management creates implicit dependencies. Tools like execute(), run(), step_control() require a prior call to set_file() or target_remote() to populate session_dict[context.session]. The error message 'Please call set_file or target_remote first to initialize a debugging session' is returned at runtime, not documented in tool descriptions. This violates pattern:tool-description (dependencies should be explicit) and pattern:dependency-guidance (should state 'call search_* first').
No validation of input parameters. The execute() tool accepts any GDB command string and passes it to pygdbmi without sanitization. This is a command injection risk, a malicious or misguided LLM prompt could pass 'quit' to terminate GDB, or 'shell rm -rf /' to execute shell commands. The tool should validate commands against a whitelist or at minimum reject dangerous patterns. This violates pattern:tool-gateway (treat agent input as untrusted) and pattern:secret-injection (GDB commands might leak secrets).
Missing tool annotations. The FastMCP framework supports readOnlyHint, destructiveHint, and idempotentHint decorators to classify tools, but none are used. This means LLMs cannot infer which tools are safe to call multiple times (idempotent) or which ones modify state (destructive). For example, set_breakpoint() and delete_breakpoint() should be marked @destructiveHint, and get_memory() should be marked @readOnlyHint. This violates protocol:tool-annotations (current MCP spec reward).
The 'run' tool description states 'Requires at least one enabled breakpoint to be set before running' but does not explain why or what happens if the constraint is violated. The tool should return a clear error like 'Cannot run without breakpoints. Call set_breakpoint() first.' to guide the LLM. Also, the description should explain the difference between 'start=False' (run freely) and 'start=True' (stop at entry). This violates pattern:command-tool (must document state mutations) and pattern:recovery-guide (errors must guide recovery).