Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server has critical definition quality issues across nearly all 55 tools. The vast majority of tools (43 of 55) have empty input schemas with NO parameters defined in the source code, only placeholder '{}' specifications. Tool descriptions are present but generic and lack detail about parameters, return values, and error conditions. Names follow verb_noun patterns reasonably well, but parameter documentation is almost entirely absent. The server appears to be a skeleton implementation where tool stubs exist but core schema details have not been populated. Schema quality is extremely low: only tools like capture_blender_window_screenshot and capture_blender_3dviewport_screenshot show detailed parameter schemas; all others (create_wall, update_wall, get_wall_properties, create_window, update_window, get_window_properties, etc.) expose empty {} schemas. This violates the critical rule that every parameter must have a description and type definition. When reviewed against the baseline that 100% of A+ tools have documented return types and all params have descriptions, this server falls far short.
43 of 55 tools have completely empty input schemas ({}). No parameters are defined, typed, or described. This violates the critical rule: 'Every parameter needs a description explaining what it controls' and 'Parameters without type definitions cannot exceed 30 in schema score.' LLMs cannot infer parameter requirements from empty schemas.
Populate input schemas for all 43 tools with empty {} definitions. Each parameter must have: 'type' (string, integer, boolean, object, array), 'description' (10 - 100 chars explaining what it controls), and optionally 'enum' (for fixed choices) or 'minimum'/'maximum' (for numeric bounds). Example for create_wall: {"type": "object", "properties": {"name": {"type": "string", "description": "Wall name (e.g., 'exterior_wall_1')"}, "height": {"type": "number", "description": "Wall height in meters (e.g., 3.0)", "minimum": 0.1}, "material": {"type": "string", "description": "Material name or 'concrete', 'brick', 'wood'"}}, "required": ["name", "height"]}
Expand tool descriptions to 80 - 200 characters. Include: (1) What the tool does, (2) When to use it vs. similar tools (e.g., 'Use create_wall to add structural elements; use create_opening to cut through existing walls'), (3) What it returns (e.g., 'Returns the wall ID for downstream updates'). Example: 'Create a wall element in the IFC model. Specify name, height, and material. Returns wall_id and coordinates. Use this for structural walls; see create_opening for cutouts.'
Document output schemas for all tools. Include: field names, types, and what each field represents. Example for create_wall: {"type": "object", "properties": {"success": {"type": "boolean"}, "wall_id": {"type": "string", "description": "UUID of created wall for use in update_wall, get_wall_properties, create_opening"}, "coordinates": {"type": "object"}, "message": {"type": "string"}}}. This enables downstream tool chaining.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 0 points across a rubric change (v1 → v2)
42/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
42
<=2025-11-25
v2
2026-03-09
F
42
-
v1
Create a physically-based rendering style.
create_roofwritesource verified35/100
Create a roof element in the IFC model.
create_slabwritesource verified35/100
Create a slab element in the IFC model.
create_surface_stylewritesource verified35/100
Create a surface style for materials.
create_trimesh_ifcwritesource verified35/100
Create an IFC entity using Trimesh-based geometry generation.
Tool descriptions are minimal (15-60 chars for most tools) and generic. They lack critical details: WHEN to use each tool (e.g., 'create_wall' vs 'create_mesh_ifc'), what parameters control, what the output contains, and what errors mean. For example, 'Create a wall element in the IFC model' tells the LLM nothing about whether this tool requires a position, material, height, or whether it returns the wall ID for downstream calls. Baseline: A+ tools average 194 chars and answer 'What does it do? When should I call it instead of a similar tool? What does it return?'
No output schemas documented. LLMs cannot know what fields to expect from tool results, forcing them to guess which fields to use for downstream calls. For example, does 'create_wall' return a 'wall_id', 'object_id', 'uuid', or nothing? This breaks the tool-chain pattern. Baseline: 100% of A+ tools have documented return types.
All 55 tools—none show output schema in available documentation
No error handling guidance. Tools like 'execute_blender_code' and 'execute_ifc_code' are high-risk (DESTRUCTIVE, WRITE) but offer no error recovery hints. LLMs don't know whether errors are retryable, whether they indicate invalid input, or whether they represent unrecoverable failures. This violates 'Error responses must tell the LLM what to do next.'
Destructive operations (delete_roof, remove_style, delete_ifc_objects, remove_opening, remove_filling) have no confirmation or dry-run support. An agent can invoke 'delete_ifc_objects' with no parameters and no chance to preview what will be deleted. Agents make mistakes, irreversible operations should support confirmation.
Code execution tools ('execute_blender_code', 'execute_code', 'execute_ifc_code', 'execute_ifc_code_tool') accept arbitrary Python as a parameter but provide no detail on security boundaries, sandboxing, available modules, or constraints. LLMs cannot reason about what code is safe to generate. Descriptions must explain: What modules/APIs are available? What is forbidden? Are file system, network, subprocess calls allowed?
Multiple tools with overlapping semantics lack clear differentiation. For example, 'execute_blender_code' vs 'execute_code' vs 'execute_ifc_code' vs 'execute_ifc_code_tool', LLMs must reason about which to use for a given task. Names alone do not clarify when to use each. Add descriptions explaining: 'Use execute_blender_code for Blender Python API calls; use execute_ifc_code for IfcOpenShell operations; use execute_code for generic Blender context code.' This prevents wasted reasoning cycles.
create_* and update_* tools have empty schemas and no parameter documentation. Agents cannot know what fields to populate. For example, does 'create_wall' accept 'name', 'height', 'position', 'material', 'layer'? Does 'update_wall' accept 'wall_id' or just operate on the selected object? These ambiguities force LLMs to hallucinate parameters or fail.
Add error recovery guidance to all tools. For example, execute_blender_code: 'If execution fails with SyntaxError, check the code for valid Python syntax. If it fails with NameError, verify the requested Blender module is imported. If it times out, the code may be stuck in a loop, try a simpler operation.' Pattern: error_type → LLM action.
Add a dry-run or confirmation parameter to destructive tools (delete_roof, remove_style, delete_ifc_objects, remove_opening, remove_filling). Example: {"operation": "dry-run" | "confirm", "description": "Pass 'dry-run' to preview deletions; pass 'confirm' to execute. Dry-run shows what will be deleted without making changes."}
Clarify code execution tools in descriptions and examples. For each code execution tool, state: (1) Which Python context it runs in (Blender, IFC OpenShell, etc.), (2) What modules/APIs are available, (3) What operations are blocked (e.g., no network, no file system writes), (4) Example usage. Example: 'Execute arbitrary Python code in the Blender context. Available: bpy (Blender API), math, random. Not available: os.system, subprocess, socket. Use this to manipulate Blender scenes, materials, and objects.'
Merge or differentiate overlapping code execution tools. If 'execute_blender_code' and 'execute_code' are functionally identical, retire one. If they differ, document: 'execute_blender_code for Blender Python API; execute_code for generic operations that don't require Blender context.' Add decision trees to tool descriptions.
Add 'required' fields to all parameter schemas. Example: {"required": ["name", "height"]} for create_wall. This tells LLMs which parameters are mandatory vs. optional, reducing failed attempts.
For tools that operate on existing objects (update_wall, get_wall_properties, get_door_operation_types), clarify how objects are selected: by ID parameter? From current selection? By name search? Document: 'Pass wall_id to specify the target, or leave empty to operate on the currently selected wall in Blender.'
Document pagination and limits for list tools. Baseline: 'Limits apply to prevent context window exhaustion. Results are capped at 20 items per call. Use offset/limit parameters to fetch additional pages.' Example for list_styles: {"type": "object", "properties": {"offset": {"type": "integer", "description": "Start index (0 - N)", "minimum": 0, "default": 0}, "limit": {"type": "integer", "description": "Results per page (1 - 100)", "minimum": 1, "maximum": 100, "default": 20}}, "required": []}. Include total count in output so agents know when they've fetched all results.