Semantic Search for Zotero using local AI embeddings with MCP and REST endpoints
ZotSeek provides 3 well-scoped read-only tools with clear names, detailed descriptions, and comprehensive input schemas. All tools start with action verbs (search, find, index) and have descriptions in the 150-350 character range. Input parameters are fully typed with constraints (enums, min/max bounds). Output schemas are partially documented through TypeScript interfaces visible in http-tools.ts. The main gap is the absence of explicit tool annotations (readOnlyHint, idempotentHint) in the registration code, though the tools themselves are demonstrably read-only. Error handling is implicit (relies on underlying search engines) but not explicitly documented per-tool. Overall, this is a solid B+ implementation with production-ready tool definitions.
Find papers semantically similar to a given Zotero item by item key, with optional library scoping and result limit
Get the current indexing status including number of indexed papers, active model ID, readiness state, and model loading status
Semantic search across Zotero library using hybrid, semantic, or keyword search modes with configurable parameters for similarity threshold, result count, and granularity
Missing tool annotations (readOnlyHint, idempotentHint, destructiveHint) in MCP tool registration. While the code comments and interfaces clearly indicate all tools are read-only, the annotations are not visible in the source. This forces clients to infer safety from the tool name alone.
Error handling guidance is not explicitly documented. The ToolResultItem and SearchToolArgs interfaces show input validation (clamping, enum checks in code), but error messages and recovery paths are not documented in tool descriptions. An LLM cannot predict what happens if query is empty, library_key is invalid, or min_similarity exceeds the valid range.
Output schema for 'search' and 'find_similar' is inferred from ToolResultItem interface but not explicitly documented in tool descriptions. The description does not specify the structure of the returned results, pagination behavior, or whether results are sorted by score. LLMs benefit from explicit output schema documentation.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 83 | <=2025-11-25 | v2 |
The 'granularity' parameter in 'search' has an enum (papers|passages) but the description does not explain the practical difference or when an LLM should choose one over the other. 'papers for item-level or passages for chunk-level results' is clear but incomplete, what's the performance/relevance tradeoff?