MCP Server for Yitam Context Engine that provides tools for LLM to manage its own memory and context, along with integration with MCP tools from external servers
The server defines 7 tools with basic JSON schemas and descriptions, but multiple critical gaps limit production readiness. All tools have names starting with action verbs (store_, retrieve_, mark_, summarize_, search_, get_, forget_) which is positive. However, descriptions are minimal (14-64 chars, well below the 194-char baseline for A+ tools), parameter descriptions lack detail and guidance, and output schemas are completely undocumented. No evidence of error handling guidance, confirmation patterns for destructive operations, or pagination support. Schema quality is inconsistent, some tools have enums and constraints, but required parameters lack min/max bounds or format guidance. The tool naming is clear and task-focused, but composition is weak: multiple tools operate on the same 'chatId' resource with overlapping concerns (memory, context, conversation stats), and there is no evidence of idempotent design or chaining support (output schemas missing, so no way to know if follow-up tools get required IDs). Security is a concern: forget_context is destructive (DELETE pattern) but has no confirmation mechanism, dry-run option, or permission gate. Error responses are not visible in the source, so no recovery guidance can be assessed.
Remove or reduce importance of old context that is no longer relevant
Get statistics about the conversation and context usage
Mark a specific message as important for future reference
Retrieve relevant context for the current conversation
Search for relevant information in conversation memory
Store important information or facts from the conversation
Create a summary of a conversation segment
Output schemas completely undocumented. No visibility into what fields retrieve_context, summarize_conversation, search_memory, or get_conversation_stats return. LLMs cannot plan multi-step workflows or extract required IDs for tool chaining.
Descriptions are extremely brief (14-64 characters vs. 194-char baseline). For example, mark_important is 'Mark a specific message as important for future reference', lacks context on WHEN to use it, HOW it integrates with other tools, or WHAT downstream effects it has. Insufficient for LLM decision-making.
Destructive operation (forget_context) has no confirmation mechanism, dry-run support, or permission gate. An agent could permanently delete conversation context without safeguard. Violates confirmation-request pattern.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 57 | - | v1 |
No pagination support on list/search tools (search_memory accepts 'limit' but no offset/cursor/next_cursor). Large result sets will exhaust context window. Missing total_count or next_cursor in response documentation.
Numeric parameters lack bounds. retrieve_context 'maxTokens' has no minimum/maximum (unbounded could cause DoS). search_memory 'threshold' is 0-1 (good) but 'limit' is unbounded. forget_context 'olderThanDays' and 'importanceThreshold' have no constraints.
Parameter descriptions lack guidance on expected formats, ranges, or dependencies. For example, 'sourceMessageId' in store_memory says '(optional)' but does not explain what happens if it is invalid or how it relates to the conversation context. 'factType' enum is present but description does not clarify when to use each type.
No error handling guidance visible. No indication of what errors these tools can raise, what they mean, or how an LLM should recover. Example: what if 'chatId' does not exist? Can retrieve_context be retried, or is it a user-fixable issue?
Tools do not appear idempotent. Calling store_memory twice with identical arguments likely creates duplicate facts. No idempotent design pattern evident, increasing risk of mid-chain retry damage.
Tool composition weak: retrieve_context and search_memory solve similar problems (finding relevant information) but with different approaches and parameter names. LLMs may conflate or misselect between them. Missing clear composition guide.