Web crawling and conversion system with FastAPI and MCP server support
Single tool 'crawl' has a clear verb-noun name and proper JSON Schema with typed parameters. However, the tool description is minimal (67 chars, below the 194-char baseline), lacks WHEN/WHY guidance, and does not explain the whitelist validation behavior or error recovery paths. Parameter descriptions are present but generic ('URL to crawl' lacks format/constraint details). No output schema is documented, the response is inferred as markdown text but not formally specified. Error handling is basic (ValueError wrapping) with no recovery guidance. The tool is READ_ONLY (safe) but lacks idempotent/destructive annotations. Overall: functional but below production baseline.
Crawl a URL and return filtered content as markdown
Tool description is too brief (67 chars vs 194-char baseline). Lacks WHEN to use, prerequisites (whitelist requirement), and what the output contains. LLMs cannot determine selection criteria.
Output schema is not documented. Response is inferred as markdown text but no formal schema (type, fields, length limits) is provided. LLMs cannot plan downstream operations or validate results.
Parameter 'url' description lacks format/constraint details. Should specify: valid URL format, whitelist requirement, and example domains. Current description 'URL to crawl (must be in whitelist)' is vague about what domains are allowed.
Error handling returns generic ValueError messages without recovery guidance. 'Failed to crawl URL: ...' tells LLM nothing about retry eligibility, user-fixable issues, or next steps.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | <=2025-11-25 | v2 |
Tool lacks annotations (readOnlyHint, idempotentHint). Although 'crawl' is read-only and idempotent, these are not declared in the schema, forcing LLMs to infer safety properties.