MCP server for StatsPAI — validation-tiered causal inference and econometrics workflows. Exposes 880+ registered statistical functions as MCP tools, including causal inference estimators (DID, IV, RD), panel methods, diagnostics, and post-estimation workflows.
StatsPAI provides 14 econometric/causal-inference tools with generally strong domain-specific descriptions (avg 180 chars, well above baseline 72 for params). Tool names follow verb_noun convention (pipeline_*, audit, preflight, sensitivity_*). However, critical gaps exist: (1) input schemas lack complete type definitions for complex parameters (e.g., 'data' and 'result' params accept 'object' with no nested schema), (2) several tools accept opaque object references ('result_id', 'result') without clear serialization guidance, (3) no output schemas documented for any tool, (4) error handling is implicit (no recovery guidance in descriptions), (5) te_summary and te_rank accept 'result' objects directly rather than IDs, breaking the agent composition pattern. Strengths: descriptions are research-grade and specific (e.g., pipeline_did explains 'preflight → did/CS estimator → audit → honest-DID sensitivity → bacon decomposition → brief'), parameter constraints use enums where appropriate (method, audience, output format), and required vs optional params are clearly marked.
Reviewer-grade audit on a result. Returns the literature checklist (parallel-trends test, honest-DID, Bacon decomposition, placebo, balance, …) with status per item and the concrete suggest_function to call to fill any missing high-importance check.
Reviewer-grade audit on a previously-fitted result. Pass the result_id returned by an earlier tool call (with as_handle=true). Returns the same checklist sp.audit() produces — every robustness check the literature expects for the design, with status='present|missing|run' and concrete suggested_function names for the missing ones.
Return the one-line agent-friendly brief for a fitted result. Uses sp.brief(). Useful when an agent wants to summarise a chained workflow without paying for the full JSON payload again.
Rambachan-Roth (2023) honest CIs on a fitted DID / event-study result. Auto-extracts betas + sigma + pre/post-period counts from the result; the LLM never ferries arrays.
Natural-language interpretation of a fitted result. When the connected MCP client advertised sampling, this REUSES the agent's own model (no API key) to explain the estimate, its uncertainty, and what the design does / does not identify — optionally focused by a `question` and tuned for an `audience`. With no sampling available it falls back to a deterministic structured brief: it NEVER fabricates a narrative. Every claim is grounded in the result's own numbers — the model is told not to invent estimates. Pass the result_id from an earlier as_handle=true call.
Object parameters lack nested schema definitions. 'data' (pandas DataFrame), 'result' (frontier/causal result), and 'result_id' (opaque handle) are typed as 'object' with no properties, nested types, or serialization guidance. Agents cannot infer what fields/methods these objects expose.
No output schemas documented for any tool. Agents cannot plan downstream calls or extract required fields (e.g., what does pipeline_did return? Is it a dict with 'report' and 'result_id' keys?). This violates the response-shaper pattern.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
End-to-end DID workflow: preflight → did/CS estimator → audit → honest-DID sensitivity → bacon decomposition → brief. Returns one markdown report + the primary result_id. Use this when the user pastes a DID dataset and asks 'is the effect real?' — the pipeline runs every diagnostic the literature expects.
End-to-end IV workflow: ivreg → first-stage F (effective + Olea-Pflueger) → Anderson-Rubin CI → e-value. Returns one markdown report + result_id.
End-to-end RD workflow: rdrobust → rdplot (PNG image) → rddensity (McCrary) → rdsensitivity (bandwidth). Returns one markdown report + result_id + an image content block.
Run pre-fit identification checks for a chosen method on a DataFrame. Verdict in {PASS, WARN, FAIL}. ALWAYS call this before fitting on an unfamiliar dataset to surface design problems (overlap, cohort sizes, IV first-stage F, running-variable density at the cutoff).
Pairwise correlation matrix with significance stars. Equivalent to Stata's pwcorr var1 var2 var3, star(0.05). Returns formatted correlation matrix with optional LaTeX/HTML output.
Run sp.sensitivity / sp.evalue / sp.oster_bounds / sp.sensemakr on a cached result. Pass method='evalue' (default) for the omitted-confounder-strength bound, 'oster' for delta/R-max, 'cinelli_hazlett' for OVB bounds.
Return efficiency scores sorted descending, with rank column. If with_ci=True, calls FrontierResult.efficiency_ci for parametric-bootstrap bounds.
Return a small descriptive DataFrame of TE (technical efficiency) scores with summary statistics (count, mean, std, quartiles, min, max, fractions above/below thresholds).
Winsorize variables at specified percentiles. Equivalent to Stata's winsor2 var1 var2, cuts(1 99). Replaces values below the lower percentile and above the upper percentile with those percentile values.
te_summary and te_rank accept 'result' objects as direct parameters rather than result_id handles. This breaks agent composition: agents cannot pass results between tools without serializing Python objects, which is not portable across MCP boundaries.
Error handling is implicit. No tool description explains what errors can occur, whether they are retryable, or what the agent should do next (e.g., 'If result_id is not found, refit the model by calling pipeline_did again'). Violates recovery-guide pattern.
te_summary description is vague ('Return a small descriptive DataFrame of TE scores'). Does not explain when to call it vs te_rank, what 'TE' means in context, or what the output structure is. Violates tool-description pattern (should be 50-200 chars, specific, and actionable).