Comprehensive MCP server for administering Msty Studio Desktop 2.9+ with 55 tools, Studio entity inventory, insights analytics, Nexus bridge, and Bloom evaluation
Server has 15 tools with mixed definition quality. Most tools have descriptions (10-250+ chars), but parameter documentation is inconsistent. Input schemas are present and visible in code, but many lack comprehensive descriptions for all parameters. Several tools have vague or generic descriptions that don't clearly articulate WHEN to use them or WHY they exist. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present. Error handling and recovery guidance are absent from descriptions. Tool composition is reasonable (single responsibilities) but parameter naming could be more explicit about accepted formats.
Comprehensive health report for Studio install, DB, and services.
Check if a model should hand off tasks to Claude based on evaluation history. Returns recommendation with confidence score and reasoning.
Run a Bloom behavioural evaluation on a local Ollama model. Tests for problematic behaviours like sycophancy, hallucination, overconfidence. Requires Anthropic API key for judge model. Default behaviours available: - sycophancy - overconfident-claims - hallucination - scope-creep - task-quality-degradation - certainty-calibration - context-window-degradation - instruction-following Define your own in cv_behaviors.py.
Retrieve past Bloom evaluation results for analysis.
Get quality thresholds and handoff triggers for a task category.
List available behaviours for evaluation, including custom ones.
Generic descriptions lacking WHEN/WHY context. Tools like 'list_configured_tools', 'get_server_status', 'sync_claude_preferences', and 'get_model_providers' have 30-50 char descriptions that don't explain when the LLM should call them or what makes them distinct from similar tools. Example: 'List MCP/Toolbox tools configured in Msty Studio' is what, but lacks context for decision-making.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. All tools are marked READ_ONLY in metadata, but MCP protocol annotations are absent from definitions. This prevents clients from making intelligent retry/safety decisions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 51 | - | v1 |
Validate that a model is suitable for Bloom evaluation.
Detect Msty Studio installation, version, database, and service ports.
Export a Studio tool configuration by name or toolId.
Generate a persona definition JSON suitable for Persona Studio import.
List local service backends and Studio-configured providers (keys redacted).
MCP server status and tool inventory summary.
List MCP/Toolbox tools configured in Msty Studio.
Query the Studio workspace database. Use query='stats' for overview, query='tables' for table list, or a read-only SQL SELECT/WITH/PRAGMA statement.
Report Claude Desktop MCP sync candidates (read-only inventory).
Missing error handling and recovery guidance in descriptions. No tool description indicates what errors are possible, how to recover, or what to try next. Example: 'read_msty_database' doesn't mention 'Query not found' or 'database connection failed' recovery paths.
Parameter descriptions are minimal or absent for many parameters. Example: 'read_msty_database' has a 'query' param with description 'SQL query or special command', but doesn't specify: valid query types, whether subqueries are supported, what 'SELECT/WITH/PRAGMA' means, or maximum query length. Parameter 'limit' lacks guidance on default behavior or maximum allowed value.
Enum parameters lack exhaustive documentation. 'bloom_evaluate_model' accepts 'task_category' with enum [research_analysis, data_processing, advisory_tasks, general_tasks], but descriptions don't explain what each category means or how to choose. Similarly, 'behavior' parameter allows 'custom or Bloom default' but doesn't list defaults or explain how to register custom ones.
Output schemas not documented. Tool descriptions do not specify what fields the LLM should expect in responses. For example, 'bloom_check_handoff' returns 'recommendation with confidence score and reasoning' but doesn't document the exact JSON structure, field names, or data types. This forces LLMs to guess about downstream field access.
Ambiguous parameter naming in 'export_tool_config'. Parameter 'tool_name' description says 'Tool name, toolId, or id to export', this overloads the parameter to accept three different types of identifiers without clear guidance on which to prefer or how to disambiguate collisions (e.g., if a tool is named 'toolId').
Missing pagination guidance for list/query tools. 'read_msty_database' accepts a 'limit' parameter (default 100) but doesn't specify: is there a cursor/offset for pagination? What happens if the result exceeds the limit? How should LLMs handle large datasets?