A terminal LLM chat app with support for langchain tools and mcp servers
The aterm server defines 2 tools (add, multiply) with minimal but present schema and description information. Both tools have basic descriptions (11-15 chars), matching parameters with integer types, and clear argument documentation. However, descriptions fall below the 20-character guideline and lack context about when to use these tools. Parameter descriptions are adequate but generic. No output schema documentation. No error handling guidance. The tools are mathematically trivial test cases rather than production-grade patterns. Schemas are properly typed (integer) but bare. This is solidly in the 'fair' range for definition quality, adequate for testing but insufficient for production use.
Add a and b.
Multiply a and b.
Tool descriptions are under 20 characters, violating critical baseline. 'Add a and b.' and 'Multiply a and b.' lack context about when to use these tools, expected inputs/outputs, and any prerequisites.
No output schema documented. LLMs cannot plan downstream tool calls or know what fields to extract from results. Both tools implicitly return an integer, but this is never stated in machine-readable form.
No error handling or recovery guidance. Tools may fail due to type mismatches, overflow, or invalid input, but provide no actionable error messages to guide the LLM.
Generic parameter descriptions ('first int', 'second int') do not explain constraints, ranges, or expected values. No mention of min/max bounds, overflow behavior, or negative number handling.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | - | v1 |
Tool names ('add', 'multiply') are single words without verb_noun structure. While semantically clear, they do not follow the verb_noun convention (e.g., 'compute_sum', 'compute_product') that helps LLMs parse intent reliably across large tool sets.