Harness engineering for career intelligence — invisible infrastructure that gives AI agents the tools, data, and context to operate on your career
Critical gaps in tool definition quality across all 4 tools. While descriptions exist, they are minimal (20-80 chars), parameters lack type information in visible schema, and no output schemas are documented. The tools are organized around a 'nerv' namespace which suggests integration with an external NERV A2A hub system, but the server configuration reveals this is a Python MCP server with HTTP transport that wraps NERV operations. However, the tool definitions in the schema show input parameters with types, but descriptions are sparse. The sample source code shows FastAPI setup and plugin architecture but does NOT reveal the actual tool registration or implementation, meaning tool definitions are partially inferred from the provided schema array rather than directly visible in source.
List all pending tasks assigned to an agent. Falls back to NERV_AGENT_SOURCE env var if agent_id omitted.
Check whether the NERV A2A hub is reachable and healthy.
Return aggregate counts for active memories in NERV memory store.
Get the current state of an A2A hub task by its ID.
Tool implementations not visible in provided source code. File references point to '.opencode/tools/nerv-stats.ts' (TypeScript) but server is listed as Python with custom framework. Tool definitions are inferred from schema array, not directly visible. Per HARD SCORING RULE: 'If you cannot see the actual tool definition in the source (only inferred): cap that tool's overall at 50.' All tools capped accordingly.
Descriptions are below 50 characters for 3 of 4 tools (nerv_memory_stats: 'Return aggregate counts...', nerv_hub_health: 'Check whether the NERV...', nerv_task_status: 'Get the current state...'). Per HARD SCORING RULE: descriptions under 20 chars must score 0-20; these are 20-50 chars, scoring 25-35. Baseline for A+ tools: 194 chars average. Current descriptions lack context on WHEN to use tools, prerequisite steps, or recovery guidance.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 45 | <=2025-11-25 | v2 |
No output schemas documented. Tools return data but LLM has no way to know what fields to expect, breaking the tool-chaining pattern. E.g., nerv_task_status accepts task_id but we do not know if it returns {task_id, status, created_at, ...} or {id, state, timestamp, ...}. Without output schema, agent cannot plan downstream calls or extract values for subsequent tools.
No error handling guidance. Tools expose READ_ONLY risk classification but do not document failure modes: What if task_id does not exist? What if NERV hub is unreachable? What should the agent do next? Per pattern:recovery-guide, errors must tell the LLM what to do next.
Parameter descriptions are minimal or missing context. nerv_check_pending_tasks has agent_id with description 'The agent ID to check. If omitted, uses NERV_AGENT_SOURCE from environment.', this is good, but other tools' parameters are under-documented. No constraints, format hints, or examples of valid values. Baseline: 100% of A+ tools have parameter descriptions.
All tools use 'nerv_' prefix which creates namespace clarity but also suggests they are wrappers around external NERV system. Descriptions do not explain this dependency or prerequisites: Must NERV hub be running? What if it is not reachable? This violates the rule 'Include dependency hints: If you only have a name, call search_users() first.' Here: If hub not responding, what is the recovery path?