BlindOracle MCP server — verifiable agent commerce. ERC-8004 passports, x402 + Fedimint payments, ProofDB delegation chains, MASSAT security audits. Composes Chainlink + Kalshi + Polymarket oracles for prediction-market settlement.
BlindOracle MCP exhibits significant definition quality gaps. Tool names lack clear action verbs (analyze_all_markets, list_target are vague). Descriptions are present but generic, analyze_all_markets lacks WHEN/WHY context for LLM selection. Parameter descriptions exist but are minimal (e.g., 'Agent identifier' repeated across tools without explaining what agent_id resolves to or how it's obtained). No output schemas documented. The run_tests, git_status, and list_target tools reference 'RQ-201 mirror' without explaining what that means to an LLM. No error handling guidance, no recovery paths, no validation rules stated in descriptions. Tool allowlist in core/tool_allowlist.py shows internal tool definitions but these do NOT match the 4 exposed tools, a sign of incomplete or inferred tool registration.
Analyze all prediction markets with oracle verification
RQ-201 mirror: git status filtered to marketplace_sandbox/<agent_id>/.
RQ-201 mirror: names+sizes under marketplace_sandbox/<agent_id>/target/[subpath].
RQ-201 mirror: pytest tests/marketplace/<agent_id>/ — 60s timeout, 2KB/1KB output caps.
Tool names lack clear action verbs. 'analyze_all_markets' is vague (analyze what aspect?). 'list_target' is unclear (list what from target?). LLMs cannot infer intent from these names alone.
Descriptions are generic and lack WHEN/WHY context. 'Analyze all prediction markets with oracle verification' does not explain when to call this vs other tools, what prerequisites exist, or what the output structure is. No guidance for LLM selection.
No output schemas documented. LLMs cannot plan downstream tool calls or extract required fields. Baseline: 100% of A+ tools have documented return types.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 40 | <=2025-11-25 | v2 |
Parameter descriptions are minimal and repetitive. 'Agent identifier for test execution' appears across tools without explaining how agent_id is obtained, what valid values are, or what happens if it's invalid. No constraints (enum, regex, length) stated.
No error handling guidance. Tools reference 'RQ-201 mirror' and '60s timeout, 2KB/1KB output caps' but do not explain what happens on timeout, what errors are retryable, or how to recover. LLMs have no recovery path.