Model Context Protocol server for Figma design system integration with visual comparison, HTML rendering, and code generation capabilities
This Figma MCP server has significant definition quality issues. While all three tools have names and descriptions, the descriptions are overly technical and lack LLM-optimized guidance. Parameter descriptions are present but many lack proper format constraints and validation guidance. Critically, input schemas are inferred from Python script signatures rather than explicitly registered in TypeScript MCP code, which violates the 'must see actual tool definitions' rule. The server exposes credentials (Figma API token) as parameters, a major security violation. Output schemas are not documented. Error handling and recovery guidance are absent. Tool composition is questionable, the three tools appear to serve a narrow use case (Figma rendering quality evaluation) rather than general Figma integration patterns.
Iteratively refines Figma-to-code fidelity by testing depth increments, comparing visual output, and converging on a target SSIM score
Retrieves a comprehensive bundle of node data including DSL, image fills, variables, and plugin snapshots
Compares two Figma nodes by rendering their images and computing SSIM (Structural Similarity Index) with optional scale/bounds variation
API token exposed as tool parameter. The 'token' parameter in all three tools requires the Figma API token to be passed explicitly. This violates the secret-injection pattern, credentials should be injected server-side via environment variables or vault, never exposed as parameters. Tokens in parameters are logged in traces and visible to users.
Tool definitions not explicitly visible in MCP TypeScript code. The provided source shows plugin code (plugin/code.js) and package.json files, but no explicit MCP server registration, tool definitions, or handler implementations in TypeScript. Tools are inferred to exist from script filenames and descriptions, but actual MCP tool schema registration is not shown.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 46 | - | v1 |
Output schemas not documented. None of the three tools document their return types or output structure. LLMs cannot predict what fields to expect (e.g., does figma_image_similarity return the SSIM score alone, or a full comparison object?). This violates the tool documentation pattern.
Overly technical descriptions lacking LLM guidance. E.g., 'figma_fidelity_loop' description uses domain jargon (SSIM, depth increments, tree depth) without explaining WHEN to use this tool or WHAT problem it solves. Descriptions should answer: What does it do? When should the LLM call it? What does it return? Current descriptions are 45-50 chars on average (below the 194-char baseline for A+ tools).
Many parameters lack validation constraints in descriptions. E.g., 'target_score' in figma_fidelity_loop is described as 'Target SSIM score (0-1)' but lacks guidance on what happens at boundary values (exactly 0, exactly 1, > 1, negative). 'max_depth' has no guidance on what values are safe/performant. Constraints should be explicit in descriptions: '(integer, 1-100, default 10)'.
No error handling or recovery guidance. If figma_fidelity_loop fails because a node is not renderable, or figma_get_node_bundle hits an API rate limit, what does the LLM do? Should it retry? Ask the user? The tool descriptions provide no error categorization or next-step guidance, violating the recovery-guide pattern.
Tool composition unclear for production use. The three tools form a specialized Figma rendering quality evaluation loop, not a general Figma integration. Missing basic operations like 'get_file_info', 'list_files', 'export_node', or 'apply_style'. This server appears incomplete for typical Figma agent workflows.
Complex parameters with undocumented dependencies. E.g., 'geometry' (in figma_fidelity_loop and figma_get_node_bundle) accepts values like 'paths', 'bboxes' but no description explains the difference or when to use each. 'plugin_depth' and 'plugin_timeout_ms' are only valid when 'include_plugin' is true, this dependency is not documented.