Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The server provides 12 tools with explicit schemas and descriptions. Naming follows verb_noun convention (list_, get_, start_, stop_, deploy_, update_, delete_, validate_, diagnose_, cleanup_) which is good for LLM comprehension. Most tools have clear descriptions (100-250 chars), though several are minimal. Parameter schemas are well-defined with regex patterns, enums, and length constraints for app names. However, descriptions lack actionable guidance on when to use which tool, expected outcomes, and error recovery. Output schemas are not documented, LLMs cannot see what fields to expect from responses. Error handling is minimal with no guidance on recoverable vs. fatal failures. Security-sensitive operations (delete_custom_app, deploy_custom_app) lack pre-execution confirmation mechanisms or detailed permission checks in the visible code.
No output schemas documented. LLMs cannot infer what fields each tool returns, forcing them to guess at response structure and complicating downstream tool composition.
Destructive operations (delete_custom_app, cleanup_stale_containers) and WRITE operations (deploy_custom_app, update_custom_app, start_custom_app, stop_custom_app) lack explicit confirmation-request patterns. delete_custom_app has confirm_deletion param but no documented dry-run. Agents can trigger data loss without explicit user confirmation.
Document output schemas for all 12 tools. For example: list_custom_apps should document it returns {apps: [{name, status, created_at, image_version, health_status}], total_count, status_summary}. This enables LLMs to chain tools and extract needed fields.
Add pre-execution confirmation dialogs for destructive operations. deploy_custom_app, update_custom_app, and delete_custom_app should return a confirmation_required result type, or require explicit user approval via MRTR (Multi Round-Trip Request) before proceeding.
Implement comprehensive error handling with recovery guidance. E.g., 'App not found. Available apps: app1, app2. Did you mean one of these?' or 'Deploy failed due to invalid Docker image. Validate the image name with validate_compose first.'
Enhance parameter descriptions with explicit constraints. compose_yaml description should note: 'Must be valid Docker Compose v3+ YAML. Common errors: indentation, undefined services, missing required fields. Call validate_compose to check before deploy.'
Document tool composition and dependencies. Add notes like: 'Typically: validate_compose() → deploy_custom_app() → get_custom_app_status()' and 'Use list_custom_apps to discover available apps before calling get_custom_app_status or start_custom_app.'
Clarify cleanup_stale_containers and diagnose_docker_issues with purpose statements. E.g., diagnose_docker_issues: 'Scans system logs and running containers for common Docker/networking issues. Returns a summary of detected problems and recommended fixes.' cleanup_stale_containers: 'Removes exited containers and dangling images that failed to clean up automatically. Requires confirm_cleanup=true.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 37 points across a rubric change (v1 → v2)
59/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
59
<=2025-11-25
v2
2026-03-09
F
22
-
v1
85/100
List all Custom Apps with status information
start_custom_appwriteauthsource verified82/100
Start a stopped Custom App
stop_custom_appwriteauthsource verified82/100
Stop a running Custom App
test_connectionread onlyauthsource verified80/100
Test TrueNAS API connectivity and authentication
update_custom_appwriteauthsource verified78/100
Update an existing Custom App with new Docker Compose configuration
Error handling is not visible in code samples. No evidence of recovery guides, error classification (retryable/user-fixable/fatal), or actionable error messages. Agents cannot self-correct on failures.
Parameter descriptions lack actionable constraints. E.g., 'compose_yaml' says 'Docker Compose YAML content' but does not explain valid format, common mistakes, or validation rules. 'lines' in get_app_logs accepts 1-1000 but description does not state this range explicitly.
Tool descriptions do not guide composition or explain dependencies. When to call validate_compose before deploy_custom_app? How do list_custom_apps and get_custom_app_status differ in output? Can get_app_logs be called on a stopped app? Undocumented dependencies force agents to explore via trial and error.
cleanup_stale_containers and diagnose_docker_issues have no input parameters but lack clarity on what state they examine. Do they examine the current system, a specific app, or all apps? What makes a container 'stale'?
diagnose_docker_issuescleanup_stale_containers
Add tool annotations (readOnlyHint, destructiveHint, idempotentHint) to signal operation type. mark get_*, list_*, validate_* as readOnlyHint=true; mark delete_*, cleanup_* as destructiveHint=true.
Define pagination for list_custom_apps. Add limit, offset, or cursor parameters and document the response includes total_count and next_cursor for large result sets.
Include IDs in responses that downstream tools need. E.g., deploy_custom_app should return app_id or app_name so subsequent start_custom_app calls have the identifier to use.