Guided electronics diagnosis with WebMCP tools for repair case management and diagnostic testing.
ProbeLoop demonstrates solid tool design with clear naming, comprehensive schemas, and thoughtful parameter constraints. All 9 tools have explicit descriptions (avg 156 chars) and well-typed input schemas with enums where appropriate. Tool annotations (readOnlyHint/untrustedContentHint) are correctly applied. Key strengths: version-based optimistic locking, const constraints for safety (observed_by='human', power_disconnected=true), and domain-specific enum values. Weaknesses: output schemas are not formally documented (only inferred from code); error handling is minimal (no recovery guidance); some descriptions lack actionable context for LLM selection (e.g., when to call list_safe_tests vs recommend_next_test).
Highlight one component on the board so the person and the agent are looking at the same spot. Changes only the view and bumps the case version.
Export a complete case report including device info, symptom, diagnosis, repair details, and verification results.
Read the current repair case: version, phase, ranked hypotheses, selected test, recorded measurements, proposed repair, and verification status. Call this first and after any change.
List the diagnostic tests that have not been run yet, ranked by expected information gain. Includes instrument, target component, time, risk, required power state, and the valid outcomes.
Return the single most informative remaining test, with the reason and the instruction the person at the bench should follow.
Record a reading the person reported for the selected test and update the diagnosis. The reading must come from the person, with USB power confirmed disconnected; the tool cannot measure anything itself.
Output schemas not formally documented. Code infers response structure (conciseState, jsonResult helpers) but no explicit schema definitions in tool registration. LLMs cannot plan downstream calls without knowing return field types.
Error handling lacks recovery guidance. No tool returns actionable error messages (e.g., 'Version mismatch: expected 5, got 4. Call get_case_state() to refresh.'). Agents cannot self-correct on failure.
Descriptions lack dependency hints and selection context. 'list_safe_tests' and 'recommend_next_test' both return test recommendations but descriptions don't explain when to call each. LLMs may pick the wrong one.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 79 | 2026-07-28+ | v2 |
Record the results of the post-repair verification test to confirm the repair was successful.
Set up one diagnostic test: shows the probe points on the board and the meter instruction to the person. Does not take or invent a reading.
Propose a repair plan based on the current diagnosis. The person approves or rejects it on the page; there is no tool for that step.
No confirmation step for irreversible operations. 'stage_repair_plan' proposes a repair but description says 'person approves on page', no dry-run or confirmation tool exists. Agents cannot verify intent before committing.
Parameter descriptions lack format/constraint details. 'note' fields accept up to 240 chars but description doesn't mention the limit. 'component_id' enum is hardcoded but no description explains what each component is (J1=fuse?, U2=IC?).