MIR Algorithm Evaluation and AI-driven Improvement Platform. An MCP server for evaluating music information retrieval algorithms, running grid searches, managing code versions, and improving detector algorithms using LLM-based optimization.
This MCP server exhibits significant definition quality gaps across most dimensions. While all 10 tools have basic schemas and descriptions present, the descriptions are predominantly in Japanese, lack sufficient context for LLM agent planning, parameters are often under-constrained, and output schemas are not documented. The tool naming follows verb_noun conventions (get_code, save_code, run_evaluation), which is positive, but descriptions frequently lack the WHAT/WHEN/WHY structure needed for agent selection. Several tools accept optional session_id parameters that appear vestigial (no clear role in stateless MCP). Parameter descriptions are sparse and many lack type clarity beyond the JSON schema. Error handling is not visible in the provided code excerpts, making recovery guidance assessment impossible. The domain (music information retrieval algorithm evaluation) is specialized, but the tool definitions do not yet meet production-grade standards for agent use.
検出器コード取得ジョブ (非同期ラッパー経由)。指定された検出器のソースコードを取得します。バージョン指定が可能です。
評価結果に基づく改善提案プロンプトを生成します。評価メトリクスからコード改善の提案を作成します。
コード改善プロンプトを生成します。LLMが改善コードを生成するためのコンテキストとプロンプトを作成します。
ジョブのステータスを取得します。指定されたジョブIDのステータスと進捗情報を取得します。
セッション履歴を取得します。指定されたセッションの全ての履歴イベントを取得します。
ジョブの一覧を取得します。実行中または完了したジョブのリストを取得します。
評価ジョブを実行します。指定された検出器に対してデータセットで評価を実行し、結果を出力します。
グリッドサーチジョブを実行します。検出器のパラメータ空間を探索し、最適なパラメータを見つけます。
Descriptions lack LLM-optimized context. Most descriptions (8 of 10) are under 100 characters and entirely in Japanese, with no English translation. Descriptions like 'グリッドサーチジョブを実行します' (Grid search job executes) lack WHAT, WHEN, and WHY context needed for agent planning.
Parameter descriptions are minimal or absent. Parameters like 'grid_config' in run_grid_search lack any description of required fields (detector_name, param_grid, evaluation_config are mentioned as inclusions but not documented as schema). The 'detector_params' in run_evaluation is typed as 'object' with trivial description.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 16 | - | v1 |
コード保存ジョブ (非同期ラッパー経由)。改善されたコードを保存し、Git連携を含む履歴管理を行います。
Pythonコードの構文を検証します。提供されたコードが有効なPython構文かどうかをチェックします。
Output schemas are not documented. No tool shows what fields are returned, their types, or structure. E.g., run_evaluation likely returns metrics, but the response shape is invisible to the LLM. LLMs need to know what fields to expect so they can plan downstream tool calls.' This is a HARD GAP for any agent use case.
Vestigial session_id parameters on stateless tools. All 10 tools accept optional 'session_id', but MCP 2026-07-28 is stateless, each request carries its own context via _meta. Session_id appears to be a leftover from older stateful designs and adds confusion: is it required? Does the server track state? This violates the current spec's stateless request handling requirement.
Error handling strategy not documented or visible. No tool description indicates what errors are possible, how they are classified (retryable vs fatal), or what recovery steps an LLM should take.
Parameter constraints under-specified. E.g., 'num_procs' in run_evaluation has no min/max; 'save_plots' and 'save_results_json' are booleans with no default documented.
Ambiguous parameter names and overloading. 'code' appears in both get_code (output) and save_code (input), but no clear distinction in naming. 'detector_params' in run_evaluation is not a standard type, unclear if it is a JSON object, a dictionary, or a path to a config file.
Language barrier. All tool and parameter descriptions are in Japanese with no English fallback. While the server may serve a Japanese-speaking team, MCP tools are typically exposed to multi-lingual agent ecosystems. Descriptions should be in English for maximum interoperability. This is a localization issue, not a technical one, but it severely limits adoption.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Per current MCP spec (2026-07-28), tools should declare their safety properties. save_code and run_evaluation are WRITE operations but lack the destructiveHint annotation visible in the risk field. This prevents the client from making safety-aware decisions about which tools to invoke.