Inter-session Hebbian memory, 138 MCP tools, and multi-agent factory for AI coding agents. Includes security audit, compliance, orchestration, and delivery tools.
Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
This MCP server has severe definition quality issues across all dimensions. Tool definitions appear to be inferred rather than explicitly registered in production code, the source provided is metadata (pyproject.toml), SDKs, VS Code extension config, and a doc-generation script, but NO actual MCP tool registration code is visible. Additionally, the tools that ARE described in the issue brief show poor naming (verb_noun clarity is mixed), minimal descriptions (most are 50 - 100 chars, below the 194-char baseline), and incomplete schemas. The security_scan tool has a description ('Run a comprehensive security audit on OpenClaw configuration') and basic schema, but most tools lack parameter descriptions. Error handling, recovery guidance, and composition patterns are not evident in any source code provided.
No explicit MCP tool registration code visible. All 8 tools are inferred from issue metadata rather than from source code showing tool.Tool() calls, schemas, and descriptions.
Parameter descriptions are missing or trivial. 'firm_gateway_fleet_status' has no input schema visible (empty {} object). Many tools lack descriptions for their parameters (e.g., 'since_days' in openclaw_hebbian_analyze has a description, but most follow the pattern of bare parameter names with minimal guidance).
Provide the actual MCP server implementation (Python or TypeScript). The rubric requires static code analysis of tool.Tool() registrations, not inferred metadata. Include the file that calls server.add_tool() or equivalent.
For each tool, write a description of 150 - 250 characters that answers: (1) What does it do? (2) When should an LLM call it? (3) What does it return? Example for openclaw_hebbian_status: 'Retrieve current Hebbian memory state: activation weights, pattern counts, and age. Call this to understand memory before deciding whether to analyze trends (openclaw_hebbian_analyze) or update weights (openclaw_hebbian_weight_update).'
Add descriptions to every parameter. Example for 'severity_filter': 'Filter security findings by severity level. Must be one of: LOW, MEDIUM, HIGH, CRITICAL. If omitted, all findings are returned.'
Document output schemas for all tools. Use JSON Schema format in code comments or docstrings. Example for openclaw_security_scan: '{ "findings": [ { "id": string, "type": string, "severity": "LOW"|"MEDIUM"|"HIGH"|"CRITICAL", "description": string, "recommendation": string } ], "total_count": integer, "scan_timestamp": ISO8601_string }'
Add error handling guidance to tool descriptions. Example: 'If the config_path is not found, returns 404 with message "Config file not found at {path}. Verify the path and try again." If the format is invalid, returns 400 with the specific parsing error.'
For destructive tools (firm_export_github_pr, firm_export_slack_digest), document what happens if called twice with the same parameters. Is idempotency guaranteed? Or will it create duplicate PRs/messages? If not idempotent, add a unique_id parameter or clarify the consequence.
Score history
Overall score trend
↑ 0 points across a rubric change (v1 → v2)
40/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
40
2026-07-28+
v2
2026-03-09
F
40
-
v1
45/100
Update Hebbian weights based on recent activations
Tool descriptions are too short and lack context. 'Get Hebbian memory status and current weights' (50 chars) is below the 194-char baseline and does not explain WHEN to call it, WHAT it returns, or HOW it differs from openclaw_hebbian_analyze. When instead of a similar tool? What does it return?
No output schemas documented. The issue metadata provides tool names and input schemas, but no response schemas are visible. Example: What does openclaw_security_scan return? A list of findings? A summary? Field names are unknown.
No error handling or recovery guidance. None of the tool descriptions mention error cases, retryable failures, or what the LLM should do if a call fails. Per pattern:recovery-guide, errors must tell the agent 'what to do next', not just fail silently.
Destructive tools (firm_export_github_pr, firm_export_slack_digest, openclaw_hebbian_weight_update) lack confirmation/dry-run support. Per pattern:confirmation-request, irreversible operations should support dry-run or confirmation. firm_export_github_pr creates a PR (irreversible); openclaw_hebbian_weight_update has a dry_run param but descriptions do not clarify what 'applying' weights actually does or whether it is reversible.
Naming clarity issues. Tool names mix prefixes: 'openclaw_*' vs 'firm_*'. Within the openclaw group, 'hebbian_status' vs 'hebbian_analyze' is unclear, does analyze include current status, or only historical trends? 'openclaw_a2a_discovery' is also vague, does it discover A2A agents, or configure A2A protocol?
Enum constraints missing where appropriate. 'severity_filter' in openclaw_security_scan is a free-form string with an example value 'HIGH, CRITICAL' in the description. Per pattern:constrained-input, this should be an enum with values ['LOW', 'MEDIUM', 'HIGH', 'CRITICAL'] to prevent LLMs from hallucinating invalid severity levels.
Parameter relationships undocumented. 'dry_run' in openclaw_hebbian_weight_update implies that without it, weights are applied, but this mutual exclusivity is not stated.
Tool composition broken. 'firm_export_github_pr' requires GitHub owner, repo, title, and description, but does not document what data it extracts from prior tool calls (e.g., agent decisions from openclaw_hebbian_analyze?). Per pattern:tool-chain, outputs of upstream tools must match the parameters downstream tools expect. It is unclear whether firm_export_github_pr should be called after specific prior tools or if it can stand alone.
firm_export_github_prfirm_export_slack_digest
Clarify tool relationships. Add to openclaw_hebbian_analyze: 'Use after openclaw_hebbian_status to understand trends in memory activations over time. Returns time-series data, not current state.'
For firm_export_github_pr and firm_export_slack_digest, document what source data they consume. Example: 'This tool extracts agent decisions and code modifications from the last successful run (from memory or context). Call after agent execution completes. If no prior execution context, it will fail with "No execution context found."'
Convert 'severity_filter' string parameter to an enum: { "type": "string", "enum": ["LOW", "MEDIUM", "HIGH", "CRITICAL"], "description": "Filter findings by severity..." }
For openclaw_a2a_discovery, clarify: Does it discover new agents, list known agents, or configure agent-to-agent communication? Example redescription: 'Discover available Agent-to-Agent (A2A) protocol endpoints at the given base URL. Returns a list of agent service endpoints, capabilities, and connection requirements. Use this to bootstrap multi-agent workflows.'
Add per-item error reporting for batch-like operations. If firm_export_github_pr fails partway through (e.g., PR creation succeeds but adding labels fails), return { "pr_created": true, "pr_url": "...", "errors": [{ "step": "add_labels", "message": "..." }] } rather than a flat error.
Add explicit timestamp handling. All tools that return temporal data (e.g., openclaw_hebbian_analyze) should return ISO 8601 dates, not Unix timestamps or epoch milliseconds.
Document rate limits and timeouts. If firm_export_github_pr has a rate limit (e.g., 1 PR per minute), state it: 'Rate limited to 1 PR per 60 seconds. If exceeded, returns 429 with retry-after header.'
For tools that accept natural-language identifiers (e.g., channel names in firm_export_slack_digest), document: 'Accepts channel name (e.g., "general") or channel ID (e.g., "C123ABC"). If ambiguous, returns a list of matching channels for user confirmation.'