Wiki-first RAG system with dual semantic + graph memory providing MCP tools for knowledge graph traversal, memory retrieval, QA history search, and expert ranking
This server implements 6 tools with moderate quality. Naming is generally clear (verb_noun pattern observed in 5/6 tools, except 'suggest_follow_ups' which uses a less conventional verb). Descriptions are present but inconsistent in depth and LLM-optimization: some are comprehensive with usage guidance (search_relationships: 240 chars with 'When to use' / 'When NOT to use' sections), others are sparse (suggest_follow_ups: 188 chars but mixes implementation concern 'No bullets, no markdown' into user-facing description). Input schemas are visible and properly typed across all tools, but lack critical output documentation. Parameters have descriptions but miss ranges, enums where applicable, and inter-parameter dependency documentation. Error handling is absent, no recovery guidance, categorization, or actionable error messages are evident in tool definitions. Security posture is reasonable (read-only tools, no secrets in params), but audit/permission declarations are missing.
BM25 keyword search over atomic facts (Weaviate Tier 2 / tier=atomic). Cost: ~$0.001. Target latency: <200ms. Results are MMR re-ranked (λ≈0.6) to improve diversity when multiple paraphrased queries hit the same top facts.
Search external web knowledge via Tavily or Olostep. Cost: ~$0.01. Target latency: ~1s. Requires TAVILY_API_KEY or OLOSTEP_API_KEY environment variable.
Search for images, PDFs, and links shared in the channel. Cost: ~$0.001. Target latency: <200ms.
Search past Q&A pairs semantically for similar questions in this channel. Cost: $0. Target latency: <100ms.
Traverse the knowledge graph for relationships between named entities. **Purpose.** Resolve each entity name to a canonical node (fuzzy match), then merge the ``hops``-radius neighbourhoods into a single deduplicated subgraph of nodes and edges. Useful for answering *"how is X connected to Y?"* style questions. **When to use.** - The question names 1-3 concrete entities (people, projects, tech) and asks how they relate. - You need structured edges with relationship types and confidence, not free-text facts. **When NOT to use.** - The question is broad / exploratory ("tell me about X"): prefer ``search_channel_facts`` or ``get_topic_overview``. - You need temporal ordering of decisions: use ``trace_decision_history``. - You need ranked people by expertise: use ``find_experts``. Cost: ~$0.005. Target latency: ~500ms.
Output schemas are completely undocumented across all 6 tools. LLMs cannot predict result structure, field names, or types, forcing them to reason about data extraction and increasing hallucination risk.
Enum-valued parameters (mode in search_external_knowledge, time_scope in search_channel_facts, media_type in search_media_references) document valid values in descriptions but do not formalize them as JSON Schema enum constraints. LLMs may hallucinate invalid values.
Numeric parameters (hops, limit) lack min/max bounds in schema. search_relationships.hops, search_qa_history.limit, search_channel_facts.limit, search_media_references.limit are unconstrained, allowing LLMs to pass extreme values that may cause performance degradation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 64 | 2026-07-28+ | v2 |
Call this tool to suggest 3 short follow-up questions as plain strings. No bullets, no markdown, no numbering prefixes. Each string is a complete question under 100 characters. Call this ONCE at the end of your answer. Do NOT write a FOLLOW_UPS JSON block in your prose — this tool replaces that mechanism.
No error handling or recovery guidance in tool definitions. Tools may fail (search_external_knowledge requires API keys; graph traversal may timeout; semantic search may return no results) but define no error responses, retry policies, or corrective actions for the LLM.
suggest_follow_ups description mixes implementation detail ('No bullets, no markdown, no numbering prefixes') with user-facing guidance. This teaches LLM about internal constraints rather than when/why to call the tool.
Tool descriptions lack inter-tool disambiguation. search_qa_history, search_channel_facts, and search_relationships all search for knowledge but differ in scope (Q&A vs facts vs graph). Descriptions do not guide LLM on which to prefer for a given query type.
Environment variable dependencies (TAVILY_API_KEY or OLOSTEP_API_KEY for search_external_knowledge) are documented in description but not formalized. No schema indicates what happens if keys are missing, and no error guidance teaches LLM recovery.