Model Context Protocol server for OPNSense firewall management with inter-VLAN routing diagnostics, ARP table, DNS filtering and HAProxy support via Claude Desktop
Scoring was not performed
Tool descriptions are critically sparse. 'List all firewall rules (cached)', 'Get all dashboard data in one call (optimized)', and 'Get cache performance metrics and recommendations' do not explain WHEN to call each tool, what the return structure is, or how they differ from non-cached alternatives. Descriptions must be 50-200 characters and answer: What does it do? When should the LLM call it? What does it return?
Output schemas are not documented. The code shows input parameter schemas (forceFetch, includeDisabled as booleans) but returns are completely undocumented. LLMs cannot infer what fields dashboard_data contains, what structure cache_metrics has, or how to chain results into downstream tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 33 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 26 | - | v1 |
Parameter descriptions are minimal or missing context. 'forceFetch' is described as 'Force fetch from API' but does not explain: what is the default behavior? Is a cached result preferred? How old can a cached result be? Similarly, 'includeDisabled' lacks explanation of what 'disabled rules' means in OPNsense context.
Tool names do not disambiguate purpose from non-cached alternatives. The server likely offers get_firewall_rules() in addition to list_firewall_rules_cached(). The 'cached' suffix is confusing, it describes implementation (caching), not intent. Better names: list_firewall_rules_fast (intent: speed) or list_firewall_rules_with_cache_control (intent: explicit control).
No error handling guidance. Descriptions do not explain what happens if cache is invalid, if the API is unreachable, or what a warm_cache failure means. Agents need recovery instructions, not silence.
Tools lack output pagination guidance. If get_dashboard_data or list_firewall_rules_cached return large datasets, there is no limit parameter, no next_cursor, and no total_count documented. This risks context window exhaustion.