A comprehensive AI development environment with state management, web tools, and MCP integration. Provides REST API and MCP endpoint for project management, state graph queries, code execution, and web search/fetch capabilities.
AItelier exposes 3 tools via fastmcp (HTTP transport). All three tools have non-empty descriptions, but definitions exhibit critical deficiencies: (1) Input schemas are present but extremely generic, both state_graph_read and state_graph_write accept an 'action' string and 'arguments' object with no enumeration of valid actions, no field-level constraints, and no type specification for arguments; (2) Descriptions are dense, domain-specific, and poorly optimized for LLM reasoning, they read like internal API documentation rather than agent-facing guidance; (3) No structured output schemas are documented; (4) Parameter descriptions lack clarity on valid values, formats, or constraints; (5) Tool names (state_graph_*) are generic and do not follow verb_noun pattern conventions; (6) Error handling is not addressed in any tool description. The descriptions contain internal jargon ('State DAG', 'writer authorization', 'driver-note', 'frontier') that presumes deep domain knowledge and does not guide LLM behavior. No actionable error recovery guidance is present.
Read the State DAG command schemas and trust boundary. State facts are separate from workflow execution. Use this before state_graph_read/write.
Private State query. Requires writer authorization even though it does not mutate. Actions include list_director_messages, list_issues, get_issue, get_driver_note, driver_note_index, get_driver_note_entry, check_driver_note_index, get_driver_guide_section, driver_note_history, search_driver_note_history, list_projects, get_graph, get_node, facet_lint, frontier, events, get_attempt, list_attempts, evidence, search_design_items, design_impact, wait_for_state_change. Driver-note search returns bounded redacted excerpts in revision order. Design queries return candidates/review hints, not semantic proof. Use cursor-based waits for updates. Exact arguments: state_graph_help.
Manage State facts using typed state_graph_help contracts. Actions include report_issue, link_issue, resolve_issue, send_director_message, acknowledge_director_message, resolve_director_message, update_driver_note, write_driver_note_entry, supersede_driver_note_entry, delist_driver_note_entry, create_project, add_nodes, revise_node, split_node, supersede_node, set_node_facet, start_attempt, recover_attempt, reconcile_attempt, disposition_failed_attempt, start_external_attempt, report_external_attempt, record_evidence, verify_node, import_tasks. Report observations (defects, gaps, hand-offs, questions) with report_issue, not add_nodes; nodes are acceptance-bearing goals. Checkpoints stay ask; completion never implies verification. Evidence must come from an actual verifier, not invented passing results. Failed external attempts may retain scoped evidence only after a terminal quiescent report; they remain failed and unverifiable.
Input schemas lack enumeration of valid action names. Both state_graph_read and state_graph_write accept an 'action' string parameter but do not constrain it to a known set of values (e.g., 'list_director_messages', 'report_issue'). LLMs will hallucinate invalid action names.
The 'arguments' parameter in both read/write tools accepts type 'object' with no schema, no required/optional field list, and no per-field type information. LLMs cannot determine what fields to pass or what types they expect.
Tool naming does not follow verb_noun convention. 'state_graph_help', 'state_graph_read', 'state_graph_write' are generic container names. Specific action names (e.g., 'get_issue', 'report_issue', 'list_director_messages') would be clearer and allow agents to understand intent from the tool name alone.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 40 | 2026-07-28+ | v2 |
Descriptions use dense domain jargon ('State DAG', 'driver-note', 'frontier', 'writer authorization', 'revision order') without explaining what these concepts mean to an LLM. Descriptions should be 50-200 chars, action-focused, and self-contained. Current descriptions range 150-500+ chars and assume deep contextual knowledge.
No output schemas are documented. LLMs cannot predict what fields or structure each action returns, forcing trial-and-error exploration and wasting tokens.
No error handling or recovery guidance. Descriptions do not explain what errors can occur, how to interpret them, or what steps to take if an action fails. For example, state_graph_read mentions 'Requires writer authorization even though it does not mutate' but does not explain what error is returned or how to recover.