Model Context Protocol (MCP) server for Firewalla MSP API - Provides real-time network monitoring, security analysis, and firewall management through 28 specialized tools compatible with any MCP client
The server has 18 tools with inconsistent quality. Naming follows verb_noun conventions well (get_*, pause_*, resume_*, clear_*, etc.), which is a strength. However, there are critical gaps in schema completeness, parameter descriptions, and output documentation. Many tools lack proper type constraints on parameters (e.g., 'query' fields accept free-form strings without enum constraints). Some tools are inferred from test scripts rather than explicit registration code, capping their reliability. Error handling guidance is absent, tools return raw API responses without recovery hints. Parameter descriptions are present but often lack specificity about valid ranges, formats, or dependencies. Output schemas are not documented in tool definitions, forcing LLMs to guess what fields to expect.
Clear the server cache and return count of cleared entries
Clear collected metrics and return count of cleared metrics
Generate a comprehensive system report with environment, memory, cache, and configuration information
Retrieve current security alerts and alarms from Firewalla firewall
Get bandwidth consumption by device
Retrieve server debug information including memory, cache, metrics, and health status
Check online/offline status of devices on Firewalla network
Free-form 'query' parameters lack enum constraints and format validation. Tools like get_active_alarms, get_flow_data, and search_flows accept arbitrary query strings (e.g., 'query:"type:N"') with examples in descriptions but no machine-readable constraints. LLMs cannot determine valid options and may pass hallucinated values.
Output schemas are not documented anywhere in tool definitions. LLMs cannot know what fields get_active_alarms, get_flow_data, or other tools return, forcing them to guess field names for downstream operations and increasing hallucination risk.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 28 | - | v1 |
Query network traffic flows from Firewalla firewall
Get network flow trends over time
Retrieve firewall rules and conditions
Get detailed information for a specific Firewalla alarm
Get network statistics organized by Firewalla box
Temporarily disable an active firewall rule for a specified duration
Resume a previously paused firewall rule, restoring it to active state
Search network flows with advanced filtering
Simulate load on the server to test performance and stability
Test connectivity to Firewalla API and return firewall status
Validate server configuration and return list of configuration issues
Tools like get_bandwidth_usage, get_flow_trends, get_statistics_by_box, and get_debug_info are defined in test scripts (scripts/test-problematic-tools.js, src/debug/tools.ts) rather than visible explicit registration in src/server.ts. This suggests tools may be inferred or dynamically registered, making reliability and schema completeness questionable. Cannot confirm these have proper JSON Schema input definitions.
No error handling guidance. Tools lack recovery hints for failure cases (e.g., 'Firewall connection failed. Ensure API credentials are set in FIREWALLA_* env vars'). LLMs receive raw errors with no actionable next steps, violating the recovery-guide pattern.
No explicit destructive operation hints. Tools like pause_rule, resume_rule, clear_cache, and clear_metrics modify state but lack toolAnnotations (destructiveHint=true or idempotentHint). LLMs cannot distinguish between safe read tools and risky write tools, risking unintended mutations.
Parameter descriptions lack specificity. Many parameter docs are under 50 chars and generic (e.g., 'Maximum results to return' for limit; 'Time period' for period). Do not explain constraints, valid values, or format expectations. Example: 'get_bandwidth_usage' period param should state: 'Valid values: 24h, 7d, 30d (represents hours or days; e.g., 24h = last 24 hours).'
No pagination metadata in output documentation. Tools accepting 'limit' and 'cursor' parameters (get_active_alarms, get_flow_data, search_flows) should document that responses include 'next_cursor' and 'total_count'. Without explicit output schemas, LLMs cannot construct pagination loops reliably.
Confusing tool overlap without clear differentiation. 'get_flow_data' vs 'search_flows' vs 'get_flow_trends' all query network flows but descriptions don't explain when to use each. LLMs will struggle to select the right tool and may make redundant calls.
Debug and admin tools (get_debug_info, test_firewalla_connection, simulate_load, clear_cache, clear_metrics, validate_configuration, generate_system_report) lack permission gates or scope declarations. Agents may inappropriately call these in production contexts without clear warnings or permission checks.