Cognitive Pattern Caching for LLM Systems - semantic caching of reasoning patterns with FAISS indexing. Exposes the Reflection Bank system as MCP tools for Claude Code integration.
Reflection Bank MCP exposes 5 tools with generally complete schemas and descriptions. All tools have clear names starting with verbs (suggest_, record_, add_, search_, get_), proper input schemas with typed parameters, and parameter descriptions. However, output schemas are not documented, the code shows tool registration but Tool objects lack explicit return type documentation. Descriptions are adequate (100-200 chars range) but lack explicit guidance on WHEN to use each tool vs alternatives. No error handling specifications visible in the code. The tool composition is reasonable (separate read/write/admin concerns), but the server is STDIO-only, which is a hard transport limitation.
Add a new learned behavior to the Reflection Bank. Use this to save proven reasoning patterns for future reuse. The behavior will be indexed for semantic search.
Get Reflection Bank usage statistics and health metrics. Returns behavior count, usage stats, success rates, and index health.
Record that a behavior was used and whether it was successful. This feedback loop improves future suggestions by tracking effectiveness. Call this after applying a suggested behavior to a task.
Direct semantic search over behaviors. Returns behaviors matching the query with distance scores. Use this for exploratory search when you want raw results.
Get relevant behaviors for current task context. Returns behavior patterns that match the task/context with relevance scoring. Use this to discover proven reasoning patterns before starting complex tasks.
Output schemas are not documented. The Tool definitions lack explicit return type specifications (no structured outputSchema field visible). LLMs cannot plan multi-step workflows without knowing what fields to expect from each tool's response.
Missing error handling specifications. No visible documentation of how tools respond to edge cases (e.g., behavior_id not found in record_usage, invalid context/task in suggest_behaviors, or index corruption in search_behaviors). LLMs need guidance on recovery actions.
Tool selection ambiguity: suggest_behaviors and search_behaviors both retrieve behaviors but with different semantics (context+task vs free-form query). The descriptions do not explicitly state when to prefer one over the other.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 73 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
get_stats lacks a meaningful description beyond listing outputs. It does not explain WHY an LLM would call it or WHEN (e.g., before adding a behavior? to validate index health before search?). Description is at the lower end of adequacy (75 chars).
No pagination/limits specified for search_behaviors or suggest_behaviors. Top_k defaults are reasonable (3 and 5), but the descriptions do not warn the LLM that large result sets may waste tokens or exceed context windows. Best practice would include guidance: 'Results capped at 20 per call; use top_k to reduce further.'
add_behavior accepts an array of steps but does not specify constraints (min/max length, max string length per step). LLMs may submit absurdly long step lists or empty arrays without validation feedback.