Local MCP server: PineScript v6 → C++ + Docker backtest runner with parameter sweeps + Binance OHLCV exporter. Fully local (transpiler bundled in the engine image), no API key.
The pineforge-backtest-mcp server has well-structured tool definitions with clear, detailed descriptions and comprehensive input schemas. However, there are meaningful gaps in output schema documentation, error handling guidance, and some parameter descriptions lack actionable constraints. Tools follow verb_noun naming conventions well (backtest, binance_klines_csv, check_pine_feature, list_pine_coverage, pull_engine_image, check_engine_image, engine_info). Descriptions are substantive (100-200+ characters) and explain purpose clearly. Input schemas use proper JSON Schema with types. The main weaknesses: (1) output schemas are not documented in the code shown, LLMs cannot plan downstream operations without knowing what fields to expect; (2) error handling lacks recovery guidance; (3) some parameters (e.g., 'image', 'runtime') have descriptions that are generic or assume prior knowledge; (4) no input validation or constraint details visible in the descriptions (e.g., concurrency 1-8 range is mentioned but max_combinations and page limits lack specifics).
Run a single backtest of a PineScript v6 strategy against historical OHLCV data with optional parameter inputs/overrides.
Run a parameter sweep backtest across multiple combinations of inputs/overrides, with optional concurrency control and result sorting.
Fetch historical OHLCV data from Binance Spot or Futures API and export as CSV for backtesting. Supports timeframe selection and date range pagination.
Check the local vs remote freshness of a pineforge-release Docker image. Compares local manifest digest against the registry without downloading layers. Only available in docker mode (host Docker daemon).
Check the coverage status of a PineScript v6 language feature in the pineforge transpiler/runtime (supported, partial, unsupported, or via_transpiler).
Get runtime information about the pineforge engine (version, mode, and capabilities). In docker mode, returns the default image URI; in local mode, returns 'local in-process'.
Output schemas are not documented. Tools like 'backtest', 'backtest_grid', 'binance_klines_csv' return complex objects (reports with trades, equity curves, CSV data, coverage indices) but LLMs cannot see the response structure to plan chaining or extraction. Without documented return types, agents cannot reliably extract subsequent parameters needed for follow-up calls.
Error handling and recovery guidance are absent. None of the tool descriptions explain what an LLM should do if a call fails (e.g., 'backtest fails if ohlcv_csv_path is invalid or not accessible'). No guidance on which errors are retryable vs. user-fixable vs. fatal.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 60 | 2026-07-28+ | v2 |
List all covered PineScript v6 features and their support status (supported, partial, unsupported, via_transpiler) across the pineforge transpiler/runtime.
Pull (download) the specified pineforge-release Docker image from the registry. Only available in docker mode (host Docker daemon).
Parameter constraints are informal and inconsistently documented. 'concurrency' is described as '1-8, default 1' but 'max_combinations' lacks specifics on what 'raises error if exceeded' means. 'sort_by' enum is clear, but 'runtime' is vague ('Optional runtime configuration'). Binance 'timeframe' accepts freeform strings instead of documenting the exact set (1m, 5m, 1h, 1d, 1w) in the description or as an enum.
Parameter 'image' in backtest and backtest_grid tools lacks actionable documentation. 'Optional Docker image URI for the pineforge-release engine' assumes users know what a valid URI looks like. Should document format (e.g., 'ghcr.io/<registry>/<image>:<tag>') and clarify what happens if the image is not found or not locally available.
The 'runtime' parameter (present in backtest, backtest_grid) is underdocumented. Description says 'Optional runtime configuration (input_tf, script_tf, bar_magnifier, magnifier_samples, magnifier_dist)' but does not explain what each field does, expected types (strings? numbers?), or constraints (value ranges, validation rules). LLMs cannot reliably construct valid runtime objects.
tools returning lists or potentially large results (check_pine_feature searching a coverage index, list_pine_coverage returning all features) have no documented pagination or result limits. If the coverage index is large, an LLM could receive excessive output that bloats context and increases token cost.
Tool 'binance_klines_csv' accepts freeform 'market' enum ('spot' or 'futures') but does not document that the two APIs have different endpoints, rate limits, or available pairs. An LLM might choose the wrong market without explicit guidance on when to use each.