Collection of agentic systems including accounting firm, call center, deep research, hospital management, professor/tutor, and quantitative hedge fund applications
Scoring was not performed
Missing output schemas for all tools. LLM cannot infer what fields to expect in responses, breaking downstream tool chaining and context planning.
Terse tool descriptions (< 100 chars). Descriptions lack WHEN to use guidance, prerequisites, or state-modification warnings. imageProcessing: 'Process and extract text from images using OCR' (61 chars). wikipedia: 'Search Wikipedia for articles and information' (44 chars). Production baseline: 194 chars average.
Parameter 'imageData' in imageProcessing lacks guidance on valid base64 encoding. No example format or size constraints documented. Invites malformed input.
Parameter 'language' in imageProcessing has default 'eng' but no enum constraint. Description says 'e.g. eng, fra, deu' but LLM may infer these are examples, not exhaustive. Should be an enum.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 24 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 35 | - | v1 |
Parameter 'limit' in wikipedia accepts 1 - 10 but no guidance on what happens when results are fewer than requested. Does it return padding, empty array, or partial list? Ambiguous.
No error handling guidance. What happens if OCR fails? If Wikipedia query returns nothing? If base64 is invalid? No recovery path or actionable error messages documented.
imageProcessing lacks idempotency guarantee. Calling with same image twice, does it return same result or different OCR text due to image enhancement randomness? Not specified.