A suite of three MCP servers for quantitative research: research-server (Panel operations, RAG lookups, statistical analysis), report-server (standardized research report generation), and coding-server (Python code execution)
This server exposes 20 quantitative research tools with explicit schemas and descriptions, but has significant quality gaps that prevent higher scoring. Tools follow a consistent naming pattern (Panel_*), but descriptions lack actionable guidance for when/why to use them. Parameter descriptions are minimal (typically 1-2 sentences). Output schemas are undocumented, critical for an agent working with complex financial data structures. Error handling is absent in the visible code. The execute_python tool (irreversible, high-risk) lacks confirmation patterns or dry-run support. Most tools appear to be simple wrappers around a data Panel API without clear recovery paths when operations fail. Naming is consistent but generic; 'Panel_binary_op' with op='add'|'sub'|'mul' etc. is readable but not as clear as 'Panel_add', 'Panel_multiply', etc. would be. Parameter descriptions are present but sparse, they state WHAT the parameter is, not WHY the LLM would choose different values or how it affects the result.
Apply a binary operation between two Panel operands (element-wise) or between a Panel and a scalar/boolean.
Coalesce multiple characteristic Panels by combining their values with priority ordering.
Digitize (bin) a Panel into discrete categories based on quantile breakpoints.
Create a boolean mask Panel indicating whether each element is contained within a given list of values.
Shift (lag/lead) a cached Panel by relabeling its date index by whole months.
Lookup Panel identifiers and descriptions using a natural-language query or an ID-like string. Searches stock characteristics Panels (signals/factors/attributes) or benchmark portfolio return Panels (index/benchmark return series).
Perform matrix multiplication between two Panels.
Output schemas completely undocumented. No tool documents what fields the returned Panel object contains, what shape the data takes, or what IDs/references are returned for chaining.
execute_python is irreversible and high-risk (Risk: IRREVERSIBLE) but lacks confirmation, dry-run, or sandbox constraints. No error guidance for sandboxing, timeouts, or malicious code.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 49 | - | v1 |
Compute portfolio returns from a weights Panel and a returns Panel.
Compute portfolio weights based on a characteristic Panel and optional constraints.
Compute regression residuals from a dependent variable Panel against one or more factor Panels.
Resample a Panel to a coarser time frequency (e.g., monthly to quarterly or annual).
Restrict a Panel by keeping only samples that satisfy the optional criteria.
Apply a rolling window operation to a Panel over time.
Saves the markdown report text as a PDF.
Shift a Panel along the time or cross-section axis.
Standardize a Panel to zero mean and unit variance within cross-sections.
Returns prompt string, including computed summary tables, for generating a report to evaluate a Panel containing stock characteristics for predicting stock returns.
Apply a unary operation to a Panel (element-wise) or to a scalar/boolean resolved from panel_id.
Winsorize a Panel by clipping extreme values at specified quantile thresholds.
Execute a Python code string and returns its standard output as a string
Parameter descriptions are minimal and generic (1 - 2 sentences max). They state WHAT the parameter is but not WHEN or WHY the LLM should use different values, or what valid ranges/options lead to. E.g., 'Number of months to shift' doesn't explain sign convention (positive=lead vs lag) until you read the description text.
Generic operation names: Panel_binary_op and Panel_unary_op accept op strings ('add', 'sub', 'mul', 'eq', 'and', 'or') instead of separate tools. This forces the LLM to construct operation strings and loses semantic clarity. Should be: Panel_add, Panel_subtract, Panel_multiply, Panel_equal, Panel_and, Panel_or.
No error handling or recovery guidance visible in code. Errors likely return raw exceptions without actionable context. E.g., if Panel_lookup fails to find a characteristic, there's no suggestion of available alternatives.
Tools returning Panel IDs (panel_id as string) but no documentation of how these IDs map to actual data or persist. Agents cannot reason about when a panel_id is valid for a follow-up call vs. stale.
Panel_portfolio_weights and Panel_portfolio_returns depend on multiple Panel inputs but descriptions do not state preconditions (e.g., must returns_panel_id and weights_panel_id be same frequency/shape? What if they don't align?)
Panel_lookup accepts both 'query' (natural language) and panel_type ('characteristics'|'benchmarks') but no guidance on precedence: does the query override panel_type? If query contains 'benchmark', does that apply even if panel_type='characteristics'?