An API to search vectorized Cursor chat history stored in LanceDB using embeddings generated by Ollama.
This MCP server exposes 2 tools via FastAPI HTTP endpoints. Both tools have basic descriptions and input schemas, but critical issues significantly impact overall quality. The server implements HTTP transport correctly, but the tool definitions lack the rigor required for production-grade agent use. Tool naming follows verb_noun convention (search_chat_history, health_check), but descriptions are minimal and parameter documentation is incomplete. The schema definitions are present but lack proper typing constraints. Error handling exists but responses do not provide recovery guidance for agents. Overall, this is a D-grade server with foundational issues that would require significant revision before production deployment.
Basic health check endpoint that returns the status of Ollama and LanceDB connections.
Search vectorized Cursor chat history stored in LanceDB using semantic similarity. Generates embeddings for the query text and returns matching chat history items.
Tool descriptions lack action/context clarity. 'search_chat_history' describes WHAT (generates embeddings, returns items) but not WHEN to use it vs alternatives, or what prerequisites are required (LanceDB/Ollama must be running). 'health_check' description is trivial (15 chars) with no context on when/why an agent should call it.
Parameter 'top_k' lacks description of valid range constraints. Schema specifies type=integer with default=15, but no documentation of minimum/maximum bounds (1 - 100?). LLM may pass absurd values like 10000 that break performance or memory.
health_check has no input parameters but also provides minimal guidance on interpretation. Response format is documented implicitly (status, ollama_connection, lancedb_c...) but cut off in source. Parameter descriptions are MISSING for search_chat_history's 'query_text' (while it has a description in schema, the description is generic and does not explain what kinds of queries work best, or dependencies on Ollama embedding model).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Error handling returns HTTPException with raw messages ('Ollama client not available', 'An unexpected error occurred: ...') but does NOT guide the agent on recovery. No guidance on retryability, user-fixable vs fatal, or next steps. Missing pattern: error responses must tell the LLM what to do next.
Output schema for search_chat_history is documented implicitly (SearchResponse model with results list and message field), but SearchResultItem only returns text and source_db. No score/relevance field exposed, limiting agent's ability to filter/rank results by confidence. Response truncation in source makes full validation impossible.
No pagination or result limiting documented for search_chat_history beyond the top_k parameter. If an embedding match returns many results, top_k is applied in LanceDB, but no guidance on typical result set size, token overhead, or context window impact. Missing baseline mention that results should be capped at 20-50 items.
Composition: search_chat_history and health_check are minimally integrated. No chaining IDs or cross-references. If an agent wants to search history but health_check fails first, it lacks guidance on whether to retry or escalate. No tool-to-tool flow documented.