Production-ready MCP server combining Qdrant vector search, Neo4j knowledge graphs, and Crawl4AI web intelligence
This server presents mixed definition quality. Nine tools are registered with generally complete input schemas and descriptive text. However, critical gaps emerge in parameter descriptions (many lack constraint details), output schema documentation (completely absent), error handling guidance (minimal), and composition patterns. The tool set shows reasonable naming (verb-first: create_, search_, extract_, store_) and covers 3 distinct domains (graph, vector, web). However, lack of output documentation, missing enums for constrained fields, and absence of error recovery guidance limit production readiness. Per-tool analysis: 6 tools score 60-70 (graph/vector tools with decent inputs but no output docs); 2 tools score 50-60 (web tools with nested object parameters missing schema detail); 1 tool (extract_web_content) scores 45 due to minimal parameter descriptions in nested request object.
Crawl and extract content from a web page or website. This tool provides comprehensive web crawling capabilities with support for single page or multi-page crawling with depth control, multiple output formats (Markdown, HTML, JSON), different crawling strategies (BFS, DFS, Best-First), content extraction and link discovery, and screenshot capture and page interaction.
Create a new node in the knowledge graph.
Create a relationship between two nodes in the knowledge graph.
Create a new vector collection for storing embeddings. This tool creates a new collection with specified configuration for storing and searching vector embeddings.
Extract knowledge from text using AI-powered analysis.
Extract specific content from a web page using various strategies. This tool provides targeted content extraction with support for LLM-based intelligent extraction with custom instructions, CSS selector-based precise element extraction, regex pattern-based text extraction, structured data extraction with JSON schema validation, and multiple output formats and content types.
Output schemas completely undocumented. No tool documents what it returns, field types, or structure. LLMs cannot plan downstream tool calls or extract specific fields without knowing output format.
Constrained parameters lack enums. Fields like 'node_type' (Entity, Concept, Person, etc.), 'relationship_type', 'distance_metric' (Cosine/Dot/Euclidean), 'output_format', 'strategy', and 'extraction strategy' are described as free-form strings. LLMs will hallucinate invalid values; use JSON Schema enums instead of text descriptions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 64 | - | v1 |
Search the knowledge graph for nodes and relationships.
Perform semantic similarity search across vector embeddings. This tool finds documents similar to the query text using vector similarity search with configurable filters and thresholds.
Store a document with vector embedding in Qdrant. This tool converts text content into vector embeddings and stores them in the specified collection for later semantic search retrieval.
Nested object parameters lack internal schema documentation. 'crawl_web_page' and 'extract_web_content' accept 'request' objects with nested properties but provide no formal JSON Schema for the nested structure. The description mentions properties but LLMs cannot validate structure.
Missing parameter constraints and ranges. Numeric parameters (abstraction_level, confidence_score, weight, max_depth, limit, wait_time, vector_size) lack explicit min/max bounds in descriptions. LLMs will pass unbounded values that may break backend logic.
No error handling guidance. None of the 9 tools document how they fail, what errors are retryable, or what the LLM should do on failure. A tool that calls external services (crawl_web_page, extract_web_content, embedding models) with no error guidance leaves agents blind.
No idempotency or duplicate prevention documentation. create_graph_node and create_graph_relationship lack guidance on whether duplicate calls create duplicate records or merge. If non-idempotent, agents retrying on timeout risk duplicate nodes/edges.
Parameter descriptions lack actionable detail. Examples: 'Additional node properties' (properties param) is vague; 'Relationship properties' is vague; 'Optional metadata dictionary' lacks guidance on expected keys.
Composition gaps: no mention of chaining IDs. If search_graph returns results, what fields does it include so downstream tools (create_graph_relationship) can reference them? If crawl_web_page returns pages, do they include URLs for extract_web_content to use?
extract_web_content has minimal parameter descriptions. The 'request' object contains properties like 'instruction', 'schema', 'output_format' with no guidance on valid values or expected formats. LLMs will guess.