Monte Carlo MCP Server - Exposes Monte Carlo simulation tools to Claude Code via Model Context Protocol
Four tools are explicitly registered with schemas and descriptions. Tool names start with action verbs (validate_, test_, run_) which is correct. Descriptions are adequate in length (100-200 chars) but lack specificity on WHEN to use each tool relative to others. Parameter descriptions exist but lack critical constraints like ranges, enum values, and format specifications. Schemas include types but are permissive (additionalProperties: true) which invites hallucinated parameters. No input validation guidance, error handling instructions, or recovery paths documented. Output schemas are not documented at all, LLMs cannot know what fields to expect from results.
Runs comprehensive business scenario Monte Carlo simulation. Use this for: - Revenue forecasting with uncertainty - Profitability analysis - Business case validation - Strategic planning scenarios Returns: Expected profit, probability of success, percentile outcomes, ROI analysis, and sensitivity.
Performs tornado diagram sensitivity analysis on simulation results. Use this to: - Identify key drivers of uncertainty - Create tornado diagrams - Prioritize risk mitigation efforts Returns: Tornado diagram data and ranked key drivers.
Tests robustness of Claude's reasoning by stress-testing critical assumptions. Use this when: - Need to find breaking points where recommendation changes - Testing sensitivity to extreme scenarios - Validating stability of conclusions Returns: Robustness score, breaking points, and stability analysis.
Validates confidence in Claude's recommendation using Monte Carlo simulation. Use this when: - Making business decisions with uncertain variables - Claude wants to quantify confidence in a recommendation - Analyzing probability of success given assumptions Returns: Confidence level, expected outcome, sensitivity analysis, and key risk factors.
No output schemas documented. Tool descriptions state what is RETURNED ('Confidence level, expected outcome, sensitivity analysis') but do not define the JSON structure LLMs expect. This forces LLMs to infer field names and types, leading to parsing errors and wasted tokens.
Parameter 'assumptions' and 'stress_test_ranges' use additionalProperties: true without type constraints. This allows LLMs to pass arbitrary keys and values. The description mentions expected format (e.g., 'distribution' and 'params' keys) but this is not enforced in the schema, inviting malformed input.
No input validation guidance or error handling instructions. Tool descriptions do not explain what happens on invalid input (e.g., malformed distribution, out-of-range num_simulations, invalid comparison operator). No recovery suggestions for LLMs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 34 | - | v1 |
Parameter 'comparison' in validate_reasoning_confidence has enum constraint in schema (>=, >, <=, <) but description does not clearly list these options. Descriptions should state the enum explicitly: 'Must be one of: >=, >, <=, <' for clarity.
Missing numeric constraints. Parameters like 'num_simulations' (default 10000), 'num_scenarios' (default 1000), 'time_horizon', and 'num_scenarios' lack min/max bounds. LLMs could pass absurd values (e.g., num_simulations=1000000, time_horizon=999999) causing timeouts or memory exhaustion.
Parameter 'outcome_data' in run_sensitivity_analysis is described as 'Data structure from base simulation results' but its schema is just {"type": "object"}. No documentation of expected fields, structure, or format. This is effectively undescribed.
No guidance on distribution parameter format. Description for 'critical_assumptions' mentions 'name, distribution, and params' but does not document which distributions are supported (e.g., normal, uniform, lognormal, beta) or expected param structure (mean/std vs low/high vs alpha/beta). LLMs must guess.
No state-mutation clarity. None of the tool descriptions explicitly state whether they modify state (create records, write files) or are read-only. Tool annotations (readOnlyHint) are absent from all tools. Agents cannot distinguish safe from destructive operations.