Model Context Protocol server for Bear Notes with RAG capabilities, providing semantic search, note retrieval, and tag management
Bear Notes MCP server has 4 read-only tools with adequate naming and schema structure, but lacks sufficient descriptions for reliable LLM tool selection. Tool descriptions are present but generic (34-74 characters vs. baseline 194 chars). Parameter descriptions are brief but present. No output schemas documented. No error recovery guidance. The server is functional but below production baseline for agent tooling.
Retrieve a specific note by its ID
Get all tags used in Bear Notes
Retrieve notes that are semantically similar to a query for RAG
Search for notes in Bear that match a query
Tool descriptions are significantly below baseline length (34 - 74 chars vs. 194 avg). Descriptions lack context on WHEN to use the tool, WHAT it returns, and error recovery paths. Generic descriptions like 'Search for notes in Bear that match a query' do not differentiate from keyword-only search.
Output schemas are not documented anywhere. The tool handler returns wrapped results like { toolResult: { notes, searchMethod } } but the LLM has no declared schema for the 'notes' array structure, field names, or data types. This violates the 'documented return types' baseline (100% of A+ tools declare returns).
Error handling responses are returned as part of toolResult with an 'error' field, but no error classification (retryable vs. user-fixable vs. fatal) or recovery guidance is provided. E.g., if 'get_note' fails with 'Note not found', the description should suggest 'Try search_notes() first' or list available tags.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Parameter 'semantic' in search_notes defaults to true but offers no guidance on fallback behavior if semantic search is unavailable. The handler gracefully degrades to keyword search, but the tool description does not state this. LLM cannot know when to expect semantic vs. keyword results or when to retry with semantic=false.
Tool 'retrieve_for_rag' is conditionally registered (only if hasSemanticSearch=true) but no documentation warns the LLM about this availability constraint. If semantic search fails to initialize, the tool disappears from ListTools, but the LLM has no guidance on why or what to do instead.
Parameters lack format/constraint documentation. E.g., 'limit' in search_notes and retrieve_for_rag have no min/max bounds. Can limit be 0? 1000000? The handler accepts any number without validation. Unbounded numeric parameters let LLMs pass absurd values.