The Uingest MCP server provides 4 tools with HTTP transport and solid descriptions. However, there are significant gaps in schema completeness, parameter descriptions, output documentation, and error handling. Tool names follow verb-first conventions, which is good. Descriptions are adequate (100-250 chars range) and explain what each tool does, but lack depth on WHEN to use them and comprehensive parameter guidance. Input schemas are partially visible but incomplete, several parameters lack explicit type declarations or constraints in the schema definition. Output schemas are entirely undocumented. Error handling is absent from tool definitions. No tool annotations (readOnlyHint/destructiveHint) despite clear read/write risk designations in the provided metadata. The codebase shows good engineering (async context managers, database integration) but the tool interface itself lacks production-grade rigor.
Crawl a single web page and store its content in PostgreSQL. This tool is ideal for quickly retrieving content from a specific URL without following links. The content is stored in PostgreSQL for later retrieval and querying.
Retrieve all unique sources (domains) from crawled documents in PostgreSQL. This tool returns a list of all domains from which documents have been crawled and stored, useful for understanding what content is available and for filtering searches by source.
Search for documents stored in PostgreSQL based on semantic similarity. This tool uses vector embeddings to find the most relevant document chunks matching your query. It supports filtering by source (domain) and returns the most similar matches ranked by relevance score.
Intelligently crawl a URL based on its type and store content in PostgreSQL. This tool automatically detects the URL type and applies the appropriate crawling method: - For sitemaps: Extracts and crawls all URLs in parallel - For text files (llms.txt): Directly retrieves the content - For regular webpages: Recursively crawls internal links up to the specified depth All crawled content is chunked and stored in PostgreSQL for later retrieval and querying.
Output schemas completely undocumented. No tool defines what it returns, field names, types, structure. LLMs cannot plan chained operations without knowing response shape.
Parameter descriptions are missing or minimal. 'url' param in crawl_single_page has description, but no format guidance (http/https only?), max length, or handling of invalid URLs. 'max_depth', 'max_concurrent', 'chunk_size' in smart_crawl_url lack constraints, what are valid ranges? Defaults stated but no bounds.
No tool annotations despite clear risk levels. crawl_single_page and smart_crawl_url are marked WRITE (destructive), but no destructiveHint annotation. search_documents_by_query and get_available_sources are READ_ONLY but lack readOnlyHint. These annotations guide agent retry logic and safety checks.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No error handling guidance. Tools lack documented error responses, recovery suggestions, or categorization (retryable vs fatal). If a crawl times out or database insert fails, the agent has no recovery path.
Missing parameter constraints and validation rules. No enums for string params, no min/max for integers. smart_crawl_url's 'source' filter param is not optional/required clarity. Numeric defaults (max_depth=3, max_concurrent=10, chunk_size=5000) stated but no bounds to prevent LLM from passing absurd values (max_depth=1000, chunk_size=1000000).
search_documents_by_query 'source' parameter is described as optional but no schema marking (required: false not visible). No guidance on what 'source domain' means, is it a hostname, full domain with TLD, or normalized value? Will 'github.com' match 'api.github.com'?
No pagination support. search_documents_by_query returns top N results (match_count default 5) but no cursor, offset, or total_count. If results exist beyond the limit, agent cannot discover them. get_available_sources returns all domains, unbounded; no total count or pagination for large datasets.
No dependency hints or chaining guidance. smart_crawl_url and crawl_single_page store content in PostgreSQL, then search_documents_by_query retrieves it, but this relationship is not documented in tool descriptions. Agent may not know to call crawl first before search.