A librarian that mediates between agents and a skill collection. The collection stays raw text on disk. Agents ask the librarian; the librarian retrieves, diversifies, and learns from outcomes.
skill_librarian_mcp has 5 tools with clear naming (verb_noun pattern) but significant issues with schema completeness, parameter documentation, and output schema visibility. Tool descriptions are present but vary in quality. Schemas are partially defined in the source code but lack complete JSON Schema formalization for all parameters. Error handling is minimal and does not guide recovery. The codebase shows good internal logic (MMR, embeddings, logging) but the MCP tool interface itself falls short of production standards.
Wider, looser sweep with wildcards for ideation
Recommend skills for an intent (MMR-diverse, recency-decayed)
Rescan + re-embed changed SKILL.md files
Agents report back whether a skill worked (append-only)
Usage digest: hot skills, dead skills, gap queries
Output schemas are not documented for any tool. The source shows tool logic but does not define what fields are returned, their types, or structure. LLMs cannot predict downstream data structure or plan follow-up calls without documented return types.
Parameter descriptions are present but lack specificity. 'Maximum number of skills to return' is generic; missing are: actual default values, valid range (min/max), and constraint justification. Descriptions should state 'default 5, min 1, max 20' not just the purpose.
Error handling is absent from the visible tool definitions. No tool documents what happens on failure (e.g., no embeddings found, no skills in library, Ollama offline). Error responses do not guide recovery. See embed() function which raises RuntimeError with messages, but these are implementation details not reflected in tool specs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 62 | 2026-07-28+ | v2 |
librarian_report tool description ('Agents report back whether a skill worked') does not clearly state it is a write operation with append-only semantics. The Risk tag says WRITE but description does not. LLMs need explicit 'this modifies state' language.
librarian_reindex has no parameters but the description does not explain what triggers a re-index (all changed files? those modified since last run?). The tool is a write operation but the description omits 'This modifies the embeddings database' language.
No pagination or result limiting is documented. librarian_find and librarian_brainstorm accept 'limit' parameters but do not document max acceptable values, what happens if limit is exceeded, or whether results are paginated. Large result sets could blow context window.