MCP server providing statistical analysis, decision support, logical reasoning, and research verification tools
The server demonstrates solid foundational quality with 12 well-named, analytically coherent tools. All tools have descriptions (avg ~180 chars, within baseline range of 34-392), and input schemas are present with type constraints and parameter descriptions. However, critical gaps exist: (1) No tool schemas document OUTPUT structure, LLMs cannot reason about what fields to expect in responses, breaking downstream chaining. (2) No error handling guidance in any tool description, agents don't know how to recover from failures. (3) Two tools (perspective_shifter, verify_research) require external API keys and network calls but lack explicit failure mode documentation. (4) Tool descriptions lack disambiguation context, when to use analyze_dataset vs advanced_statistical_analysis is implied but not explicit. (5) No per-tool annotations (readOnlyHint, idempotentHint) despite all tools being read-only operations. The server is solidly functional for exploratory analysis but falls short of production-grade agent-facing tool design.
Transform a numeric series for downstream modeling: min-max normalization, z-score standardization, missing-value handling, or IQR outlier detection. Returns a markdown report with the transform's parameters and a preview of the resulting values. Use analyze_dataset to describe data without changing it.
Fit a regression model (linear, polynomial, logistic, or multivariate) predicting a named dependent variable from named predictor columns. Returns a markdown report with fitted coefficients, performance metrics, and interpretation. Use this when you have a designated outcome to predict; for association strength without a model use advanced_statistical_analysis, and to score existing predictions use ml_model_evaluation.
Compute per-column descriptive statistics, or Pearson correlation for every numeric column pair, over a table of records. Returns a markdown report (mean/median/std/variance/min/max per column, or r plus a weak/moderate/strong label per pair); non-numeric columns are ignored. Use analyze_dataset for a single numeric series; for correlation significance (p-values) use hypothesis_testing; to fit a predictive model use advanced_regression_analysis.
Summarize a single numeric series with descriptive statistics. Returns a markdown report: 'summary' gives count/min/max/mean/sum; 'stats' adds median, quartiles, standard deviation, variance, and coefficient of variation. Accepts a number[] or an array of objects (the first numeric property is used). For multi-column tables or cross-variable correlation use advanced_statistical_analysis; to transform values use advanced_data_preprocessing.
NO OUTPUT SCHEMAS DOCUMENTED for any of the 12 tools. Tool descriptions state what is returned (e.g., 'markdown report', 'decision matrix', 'coefficients and metrics') but do NOT specify the JSON object structure, field names, or types. LLMs cannot reason about how to extract data from responses or chain results to downstream tools.
NO ERROR HANDLING GUIDANCE in any tool description. Tools do not explain what failures can occur, whether they are retryable, how LLMs should recover, or what actionable next steps exist. Example: perspective_shifter and verify_research depend on EXA_API_KEY and network; if Exa is down, the LLM is given no recovery path.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 33 | - | v1 |
Generate a chart specification (Vega-Lite) plus rendering instructions for a dataset — it describes a chart, it does not render an image. Supports scatter, line, bar, histogram, box, heatmap, pie, violin, and correlation plots. Returns a markdown report with the data-point count, the spec, and usage guidance.
Rank options against weighted criteria with a weighted-sum decision matrix. Returns a markdown report: ranked options, a per-option breakdown (score × weight contribution, strengths, weaknesses), and a recommendation. Weights are normalized to sum to 1; omit them for equal weighting.
Run a statistical hypothesis test and report the p-value with a reject / fail-to-reject decision at the chosen alpha. Supports independent (Welch) and paired t-tests, Pearson-correlation significance, chi-square independence, and one-way ANOVA. Returns a markdown report with the test statistic, p-value, and conclusion. Use this when you need significance; for descriptive correlation without inference use advanced_statistical_analysis.
Assess a natural-language argument for structure, validity, strength, and fallacies. Returns a markdown analysis; 'comprehensive' (default) runs all four plus optional improvement recommendations. Use this for overall argument quality; to only flag and name fallacies use logical_fallacy_detector.
Detect and name logical fallacies in text via pattern matching, each with a confidence score, description, and before/after examples. Returns a markdown report grouped by category with an overall severity assessment. Use this to flag specific fallacies; for a full argument assessment use logical_argument_analyzer.
Score an existing model's predictions against actual values. Classification returns accuracy/precision/recall/F1 from a binary (0/1) confusion matrix; regression returns MSE/MAE/RMSE/R². Returns a markdown report of the requested metrics plus sample count. This scores supplied predictions; to fit a model from raw data use advanced_regression_analysis.
Generate alternative viewpoints on a problem — by stakeholder or discipline — grounded in web research via Exa. Returns a markdown report with key facts and actionable insights per perspective. Requires EXA_API_KEY and ENABLE_RESEARCH_INTEGRATION=true and makes live network calls; it fails without them. To cross-check factual claims instead of generating viewpoints, use verify_research.
Cross-check a factual claim across multiple web sources using Exa, compute consistency, and report verification status. Returns a markdown report with source-by-source quotes, a per-source consistency assessment, and an overall 'verified' / 'inconclusive' / 'contradicted' label based on agreement. Requires EXA_API_KEY and ENABLE_RESEARCH_INTEGRATION=true; makes live network calls. To generate alternative viewpoints instead, use perspective_shifter.
NO TOOL ANNOTATIONS (readOnlyHint, idempotentHint, destructiveHint). All 12 tools are read-only and idempotent, they should declare this via tool annotation metadata so clients know they are safe to call without side effects and can be retried freely. No evidence of annotation support in code.
AMBIGUOUS PARAMETER BEHAVIOR in several tools. Examples: (1) decision_analysis weights parameter says 'omit them for equal weighting' but doesn't specify behavior if partial weights provided. (2) logical_argument_analyzer includeRecommendations flag effect is tied to 'comprehensive' type but undocumented. (3) advanced_data_preprocessing missing-value handling strategy is not explained. (4) ml_model_evaluation evaluationMetrics parameter accepts arbitrary strings with no enum constraint.
MISSING PARAMETER BOUNDS AND CONSTRAINTS. Examples: (1) perspective_shifter numberOfPerspectives has no max; LLM could request 1000 perspectives and hang. (2) hypothesis_testing alpha parameter has no min/max documented. (3) data_visualization_generator title parameter is unconstrained. (4) Several tools accept array parameters with no documented size limits (e.g., logical_fallacy_detector categories).
INSUFFICIENT DISAMBIGUATION between similar tools. Examples: (1) analyze_dataset vs advanced_statistical_analysis, descriptions mention the distinction, but LLMs often conflate 'single series' vs 'per-column stats'. (2) logical_argument_analyzer vs logical_fallacy_detector, both operate on argument text; which to call first is not obvious. (3) decision_analysis vs ml_model_evaluation, both produce rankings; use cases not clearly separated. Better: add explicit 'When to use THIS tool vs [related tool]' section to descriptions.
TWO TOOLS HAVE EXTERNAL DEPENDENCIES NOT CLEARLY SURFACED IN REGISTRATION. perspective_shifter and verify_research require EXA_API_KEY environment variable and ENABLE_RESEARCH_INTEGRATION=true flag. Descriptions mention this (which is good), but the tool registration should include a 'requires' or 'prerequisites' field in the schema/metadata so clients can discover and warn about missing dependencies at registration time, not at call time.