A type-checker and contract layer for AI agent tool calls, deny-by-default, in-process, Pydantic-only. Strict argument validation, ghost-argument stripping, and self-healing retries for MCP servers and agent frameworks.
The server defines 3 tools with basic schemas and descriptions. All tools have names that start with action verbs (fetch_, search_, send_), which is good. Descriptions are present but generic and lack crucial context about error conditions, dependencies, and use cases. Input schemas are present with type definitions, but parameter descriptions are minimal. Output schemas are undocumented. The tools show asymmetric parameter quality: fetch_user has minimal documentation, search_products includes a default value and rate-limit caveat in description (good), but send_notification lacks clarity on PII masking behavior and what constitutes 'strict mode'. No documented error handling or recovery guidance. This is typical C/C+ tier work, present but incomplete.
Fetch user data asynchronously. Simulates an async database or API call.
Search products asynchronously. Rate-limited to 10 calls per minute.
Send a notification asynchronously. Strict mode rejects unknown arguments. PII masking enabled for output.
Output schemas not documented. LLM cannot infer what fields fetch_user, search_products, and send_notification return. This forces the LLM to guess at downstream tool compatibility and risks broken tool chains.
fetch_user description (41 chars: 'Fetch user data asynchronously. Simulates an async database or API call.') is too terse. Does not state: when to call it vs similar tools, what fields are returned, what happens on user-not-found, or whether the call is idempotent. LLM cannot reliably select this tool.
Parameter descriptions are generic or missing context. 'user_id' description ('The user ID to fetch') restates the name rather than explaining format, range, or expected values. Rubric baseline: average param description is 72 chars; here most are under 30 chars.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 54 | - | v1 |
send_notification description mentions 'PII masking enabled for output' and 'Strict mode rejects unknown arguments' but does not explain: what PII is masked, when strict mode applies, what errors result from unknown args, or how to recover. LLM cannot reason about these behaviors.
No error handling or recovery guidance. Tools do not document: which errors are retryable, which are user-fixable, what to do if a required field is missing, or how to interpret failure responses. Rubric requires 'Error responses must tell the LLM what to do next.'
search_products declares 'limit' with default=10 but no maximum. Rubric requires bounds for numeric parameters (p10 - p90). LLM could pass arbitrary large limits, potentially overwhelming the API or timeout.
send_notification 'priority' parameter has no enum constraint. Description says 'default="normal"' but does not list valid values. LLM will hallucinate values like 'critical', 'asap', 'urgent' when only 'low|medium|high|urgent' (or similar) are valid.
Tool descriptions do not state whether calls are idempotent. Rubric: 'Make tools produce the same result on repeated calls with the same input. Agents retry on ambiguous failures, non-idempotent tools risk duplicate side effects.' send_notification is a WRITE tool, critical to document retry safety.