Server has 5 tools with reasonable naming (verb_noun pattern) and documented schemas, but falls short of production quality. Descriptions are present but generic (averaging ~85 chars, below the 194-char baseline). Parameter descriptions exist but lack depth, e.g., 'message' in display_message has no guidance on length limits or encoding. Schemas are properly typed JSON Schema but lack output documentation. No error handling guidance visible. No tool annotations (readOnlyHint, destructiveHint). Risk profile is clear (READ_ONLY vs WRITE) but not surfaced to LLM via annotations. This is a Fair (C-grade) server, common gaps that prevent B-grade classification.
Display an image on the micro:bit LED matrix.
Display a text message on micro:bit LED matrix
Get the current temperature reading from the micro:bit sensor
Play music on the micro:bit using an array of notes
Wait for a button press on the micro:bit
Missing output schemas for all tools. LLMs cannot reason about results or plan downstream calls without knowing return types. E.g., get_temperature should document 'returns {temperature: number, unit: string}', wait_for_button_press should document 'returns {button: string, elapsed_seconds: number}'. This violates pattern:tool and mxe:include-chaining-ids.
Tool annotations missing. No readOnlyHint or destructiveHint in tool definitions, even though schema metadata flags READ_ONLY vs WRITE risk. Tools like display_message and display_image should be annotated with destructiveHint:true (they modify device state), while get_temperature and wait_for_button_press should be annotated with readOnlyHint:true. This prevents LLMs from reasoning about idempotence and retry safety.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter descriptions lack actionable detail. 'message' in display_message has no guidance on length limits, character encoding, or overflow behavior. 'notes' in play_music includes inline examples (C4:4, D4:2) instead of a formal regex pattern or enum. This forces LLMs to guess valid formats and increases hallucination risk.
No error handling guidance. Tool descriptions do not specify what errors can occur (timeout, invalid format, device disconnected) or what the LLM should do (retry, try different input, ask user). This violates pattern:recovery-guide and prevents agent self-correction.
Tool descriptions are generic and short (average 85 chars, baseline 194). get_temperature (62 chars) omits units, precision, and when to call it. wait_for_button_press (126 chars) lacks details on blocking behavior and return format. Descriptions should follow the pattern: WHAT does it do? WHEN should LLM call it? WHAT does it return? See pattern:tool-description.
Inline examples in descriptions (play_music shows 'C4:4', 'D4:4'; display_image shows '00300:03630:36963:03630:00300'). LLMs tend to reuse examples literally rather than adapt them, causing failures.
display_message and display_image lack details on message length limits and image matrix overflow behavior. What happens if message is too long? Are multi-line messages supported? How many characters fit? Constraints should be explicit in parameter description.