MCP server for analyzing .NET memory dumps using SOS/ClrMD, with support for crash analysis, performance profiling, security analysis, and AI-powered diagnostics
The server defines 12 tools with moderate quality. Most tools have descriptions and input schemas, but quality is uneven. Key strengths: structured schemas present for most tools, detailed parameter documentation in some tools (e.g., report, analyze). Key weaknesses: tool names lack clear action verbs (session, dump, exec are generic), several tools bundle multiple responsibilities (session action='create|list|close|restore|debugger_info'), error handling guidance is absent, output schemas are not documented. The exec tool is particularly problematic, it accepts arbitrary debugger commands with no constraint or guidance on safe/unsafe values. Naming conventions do not consistently follow verb_noun patterns. Several tools use 'action' parameter as a router, which forces LLMs to understand enum values rather than invoking distinct tools.
Analyze a dump: crash | ai | performance | cpu | allocations | gc | contention | security. For security capabilities: kind=security, action=capabilities.
Performs automated crash analysis on the currently open dump. Automatically analyzes crash type, extracts exception information, analyzes call stacks, and provides recommendations. Output is structured JSON.
Performs AI-powered deep crash analysis using a server-driven sampling loop. Returns canonical JSON report document enriched with an analysis.aiAnalysis section.
Compares two memory dumps to identify differences in memory, threads, and modules. Performs comprehensive comparison including heap/memory comparison (detects memory leaks), thread state comparison (detects deadlocks), and module comparison (detects version changes).
Compares heap/memory allocations between two dumps to identify memory leaks (growing allocations), memory pressure (high memory usage), and type growth patterns.
Compares loaded modules between two dumps to identify newly loaded modules (plugins, updates), unloaded modules, module version changes, and module base address changes (ASLR).
Generic tool names lacking clear action verbs. 'session', 'dump', 'exec' do not follow verb_noun pattern. 'session' could mean get_session, create_session, or manage_session, LLMs must guess the intent.
Multiple tools use 'action' router parameter instead of distinct tool names. 'session' with action='create|list|close|restore|debugger_info' should be 5 separate tools: create_session, list_sessions, close_session, restore_session, get_debugger_info. This forces LLMs to understand enum values and violates single-responsibility principle.
exec tool has minimal constraints. Description 'Execute a raw debugger command (last resort)' is vague. Parameter 'command' accepts arbitrary string with no enum or pattern constraint. No guidance on which commands are safe/unsafe, what happens on failure, or recovery steps. This violates error-handling and security patterns.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Compares thread states between two dumps to identify new threads (created between dumps), terminated threads, thread state changes (e.g., new waiting threads), and potential deadlock situations.
Downloads Datadog.Trace symbols from Azure Pipelines or GitHub for a specific commit. Downloads symbols for native tracer symbols (.debug files), native profiler symbols (.debug files), and managed symbols (.pdb files). Symbols are automatically loaded into LLDB for improved stack traces.
Manage dumps: open | close
Execute a raw debugger command (last resort).
Generate reports: full | summary | index | get (formats: json | markdown | html). Returns report content or section JSON.
Manage sessions: create | list | close | restore | debugger_info
Output schemas are not documented. Tools return structured JSON (e.g., report, analyzeCrashWithAi) but response structure is not declared in tool definitions. LLMs cannot know which fields to expect, making downstream chaining and extraction error-prone.
Error handling guidance is absent. No tool documents what errors can occur, which are retryable, or what the LLM should do on failure.
Parameter 'command' in exec tool has no format constraint, length limit, or validation examples. LLMs cannot know what constitutes valid input. Vulnerable to injection and misuse.
analyze tool overloads 'kind' parameter with multiple unrelated actions: crash, ai, performance, cpu, allocations, gc, contention, security. Some kinds require sessionId/userId; others do not. Conditional requirements are documented but should be separate tools (analyze_crash_kind, analyze_performance_kind, etc.) or at least explicit enum with clear docs on when each param is required.
Boolean parameters (includeWatches, includeSecurity, loadIntoDebugger, forceVersion) lack explicit default values in descriptions. Defaults are mentioned in descriptions but not in the schema, creating ambiguity on agent behavior.