MCP server that exposes tools for controlling Autodesk Flame
flame-mcp demonstrates solid foundational quality with 10 well-named tools that follow verb_noun conventions. All tools have descriptions (average ~180 chars, within the 10-1024 baseline). Tool annotations are correctly implemented for safety classification. However, significant gaps exist: (1) output schemas are completely undocumented, no tool declares what it returns, forcing LLMs to guess; (2) input schemas are visible in the source but sparse, only tool 1 (execute_python) and tool 3 (learn_pattern) show explicit parameter descriptions beyond the raw schema; (3) error handling is minimal, no recovery guidance or actionable error messages visible; (4) parameter constraints are missing enums where appropriate (e.g., tool 2's 'limit' accepts arbitrary integers with no bounds). The knowledge-base tools (search_flame_docs, learn_pattern, execute_python) are well-conceived for RAG-augmented code synthesis, but parameter validation and output documentation are weak. Token tracking and session stats are interesting but orthogonal to core tool quality.
Execute Python code inside Flame to control projects, edit timelines, run batch operations, and query API data. Code runs in Flame's main thread with full access to the flame module and all Flame Python APIs.
Query project metadata (name, description, resolution, frame rate, bit depth) using Wiretap. Returns structured project data without loading the project into Flame.
Retrieve in-session statistics: cumulative call counts by tool, total tokens saved by dedicated tools vs. RAG, timing histogram (last 20 calls), and server uptime since last reset.
Query timeline metadata (duration, frame rate, resolution, bit depth, number of tracks). Does not require the timeline to be open.
Store a working code pattern in the Flame API knowledge base (FLAME_API.md) so future sessions can discover it via search_flame_docs. Only callable by Opus / Fable models (write-allowed). Prevents hallucination contamination by restricting smaller models to read-only access.
List all clips in a specified reel. Returns clip names and metadata (duration, resolution, frame rate).
Output schemas completely undocumented. No tool declares its return type, structure, or fields. LLMs cannot infer what data is available for downstream composition or chaining.
Missing input parameter constraints. search_flame_docs accepts 'limit' as unconstrained integer (no min/max bounds); no enum for risk categories or tool selection logic. Enables LLMs to pass absurd values (limit=999999).
Error handling lacks actionable recovery guidance. No visible error responses that tell LLMs 'try X instead' or classify errors as retryable/fatal. Code scrubs secrets (scrub_secrets) but does not provide recovery hints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
List all media libraries in the current workspace. Filters out hidden system libraries automatically. Returns library names suitable for list_reels.
List all reels in a specified media library. Returns reel names suitable for list_clips.
Manually reset the in-session statistics (call counts, token usage, timings). Statistics are also auto-reset after an idle gap (default 30 minutes, configurable via config.json).
Search the Flame Python API documentation using RAG (Retrieval-Augmented Generation). Returns relevant code patterns, method signatures, and usage examples. Search is cached in-session; identical queries return results from memory without hitting ChromaDB again.
Discovery tools (list_libraries, list_reels, list_clips) lack pagination. No 'limit' or 'offset' parameters; no indication of result count. If a workspace has hundreds of reels, returning all causes context window exhaustion.
Irreversible execute_python tool lacks dry-run or confirmation step. Code runs immediately in Flame's main thread with full access, no way to preview or confirm before execution.