AI-Powered AWS Architecture Diagrams with Validation & Persistence via Claude MCP
CloudForge has 6 tools with mixed quality. Naming is clear and action-verb-based (generate_, validate_, list_, get_, delete_). Descriptions exist for all tools and are moderately detailed (ranges 60-450 chars), though the longest description for generate_from_description is verbose with implementation details that should be abstracted. All tools have visible input schemas with proper type declarations (string, boolean). However, critical gaps emerge: (1) Output schemas are NOT documented, tools return dict[str, Any] with no structure specification, forcing LLMs to infer response fields; (2) Error handling is minimal, most tools catch exceptions and return success/failure without recovery guidance or error classification; (3) Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are completely absent despite clear semantic distinctions (validate_diagram and list_diagrams are read-only, delete_diagram is destructive, generate_diagram is write-heavy); (4) Parameters lack validation rules in descriptions, no mention of required vs optional, no constraints on code length, diagram_id format, or tag patterns; (5) Composition issue: generate_from_description requires GOOGLE_API_KEY but the tool description does not state this is a hard dependency or warn that it may fail silently if the key is missing.
Delete a saved diagram.
Generate an AWS architecture diagram.
Generate AWS architecture diagram from natural language description. Uses LangGraph Pipeline with auto-retry for reliability: 1. Blueprint Architect: Analyzes text → structured blueprint 2. Diagram Coder: Blueprint → Python code (with LangChain retry) 3. Validator: AST + security validation 4. Generator: Code → diagram images (PNG, PDF, SVG). Features: Automatic retries on failure (up to 3x), Structured output parsing with Pydantic, Comprehensive error handling, Observability via LangChain/LangSmith. Requires GOOGLE_API_KEY environment variable for Gemini API.
Get a specific diagram.
List all saved diagrams.
Validate diagram code.
Output schemas are completely undocumented. All tools return dict[str, Any] with no specification of response structure. LLMs cannot infer what fields to expect or plan downstream tool calls.
Tool annotations missing. Read-only tools (validate_diagram, list_diagrams, get_diagram) lack readOnlyHint. Destructive tool (delete_diagram) lacks destructiveHint. Annotations help LLMs reason about safety and plan accordingly.
Error handling is minimal and non-recoverable. Exception handlers return success: false / message with no guidance on root cause, next steps, or error classification (retryable vs user-fixable vs fatal). E.g., generate_diagram validation failure lists errors but does not suggest retrying with corrected code.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Irreversible operations lack confirmation. delete_diagram has no dry-run, confirmation, or pre-delete validation. No pattern to prevent accidental data loss when an agent misuses the tool.
Parameter constraints missing. 'code' parameter in generate_diagram and validate_diagram lacks length limits or format hints. 'diagram_id' in get_diagram and delete_diagram has no documented format (UUID vs opaque string). 'tag' in list_diagrams has no enum/pattern. Descriptions do not state validation rules.
generate_from_description has hard dependency on GOOGLE_API_KEY but provides no recovery path. If the key is missing, pipeline_enabled=False is logged as a warning, but the tool itself does not check this or return an actionable error when invoked. LLM will waste a call before discovering failure.
Description for generate_from_description is bloated (450+ chars) with implementation details (LangGraph, Pydantic, LangSmith). This wastes tokens and buries the core function. Should be 50-200 chars and abstract internals.
No pagination support in list_diagrams. Tool accepts optional 'tag' filter but has no limit, offset, or next_cursor parameters. If many diagrams exist, returning all at once risks context explosion and performance degradation.
No tool composition guidance. It is unclear when to use generate_diagram (manual code) vs generate_from_description (LLM-assisted). Description should state: 'Use for hand-written diagrams. For automatic generation from natural language, use generate_from_description.'