MemoryGate is an MCP server that provides memory management and retrieval capabilities, including hot/cold storage tiering, archival, concept management, and semantic search with embeddings.
Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
MemoryGate provides 24 well-structured tools with explicit schemas and descriptions visible in core/mcp/server.py. All tools have non-empty descriptions (10-400+ chars range). Parameters are typed and mostly described. However, there are consistent gaps: (1) output schemas are not documented for any tool, making it unclear what fields/structure agents should expect when calling them; (2) several parameter descriptions are generic or lack constraints (e.g., 'threshold', 'metadata', 'actor' lack format/validation guidance); (3) error handling guidance is absent from descriptions; (4) some parameter relationships are undocumented (e.g., memory_ids/summary_ids/cluster_ids are mutually exclusive but not stated); (5) no dry-run confirmation patterns visible for destructive operations despite dry_run flags. Tool naming is verb-first and clear (memory_search, archive_memory, etc.), following conventions well. The server demonstrates solid baseline quality (near production-ready for a memory management system) but lacks the polish and LLM-optimization that would elevate it to A-grade.
Output schemas not documented. For all 24 tools, there is no documentation of what fields the response contains, what structure nested objects have, or what data types are returned. LLMs cannot plan downstream tool calls or extract chaining IDs without knowing the response structure.
Document output schemas for all 24 tools. For each tool, add a comment block or markdown section specifying: return type (object, array, string, etc.), required fields, field types, nested object structures, and example responses. E.g., memory_search should document: {results: [{id, text, confidence, domain, edges?, chains?}], total_count}.
Add confirmation-step guidance to destructive tools. Update descriptions of archive_memory, memory_delete_concept, and memory_delete_pattern to explain: 'Call with dry_run=true first to preview changes, then call with dry_run=false to confirm.' Include examples of what dry_run output looks like.
Convert vague enum parameters to explicit enums in the schema. For 'mode' in archive_memory, replace description with JSON schema enum: ['archive_and_tombstone']. For 'edge_direction' in memory_search, define enum: ['both', 'incoming', 'outgoing']. For 'concept_type' and 'pattern_category', enumerate allowed values.
Clarify mutually exclusive parameters. In archive_memory and rehydrate_memory descriptions, add: 'Provide exactly one of: memory_ids, summary_ids, cluster_ids, or threshold. Passing multiple will result in an error.'
Document error scenarios and recovery steps. For each tool, add an 'Errors' section in the description: 'If no results found, try broadening the query or reducing min_confidence. If threshold invalid, ensure it is a dict with keys: {field: string, operator: string, value: number}.' Provide at least 2-3 error cases per tool.
Expand short parameter descriptions. Rewrite 'query' → 'The natural language search query. Supports text search across observation summaries and metadata. Max 500 chars.' Rewrite 'limit' → 'Maximum number of results to return (1-100, default 5). Specifies hard limit on response size.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 41 points across a rubric change (v1 → v2)
65/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
C
65
<=2025-11-25
v2
2026-03-09
F
24
-
v1
auth
source verified
75/100
Initialize a new conversation session for memory tracking
memory_list_conceptsread onlyauth50/100
List concepts with optional filtering
memory_list_patternsread onlyauth50/100
List patterns with optional filtering
memory_recallread onlyauthsource verified77/100
Recall memories for a domain with confidence filtering
memory_register_aiwriteauth50/100
Register a new AI instance
memory_relate_conceptswriteauth50/100
Create or update a relationship between two concepts
memory_resolve_conceptread onlyauth50/100
Resolve a concept by name or alias
memory_searchread onlyauthsource verified78/100
Search memory with vector similarity and filtering options
memory_statsread onlyauthsource verified57/100
Get statistics about stored memories
memory_storewriteauthsource verified75/100
Store a new observation/memory with confidence and evidence
Destructive operations lack confirmation or dry-run guidance in descriptions. archive_memory and memory_delete_concept/memory_delete_pattern have dry_run flags but descriptions do not explain the confirmation flow, making it unclear whether an agent should call with dry_run=true first to preview changes.
Generic parameter descriptions lack validation constraints. Parameters like 'threshold' (archive_memory), 'metadata' (multiple tools), 'actor', 'reason', and 'mode' have vague descriptions without specifying allowed values, formats, ranges, or examples. 'threshold' is particularly opaque, it is described as 'Threshold criteria for archival' (object type) with no schema or format guidance.
Mutually exclusive parameters not documented. archive_memory accepts memory_ids, summary_ids, and cluster_ids, but descriptions do not state that only one should be provided, or what happens if multiple are supplied. This forces LLMs to guess or causes silent incorrect behavior.
No error handling or recovery guidance in descriptions. Descriptions do not explain what errors are possible, what they mean, or how an agent should respond. For example, search_cold_memory does not document what happens if min_score > max_score, or how to interpret an empty result vs a service error.
Parameter description length inconsistency. Some parameters (e.g., 'limit', 'query') have very short descriptions (5-10 chars) that lack context. Baseline expectation is 50 - 100 chars minimum for descriptive parameters. E.g., 'query' in memory_search is described as 'Search query string' (19 chars), no guidance on format, length, or special syntax.
Pagination and result limiting underdocumented. memory_search has a 'limit' parameter but does not state whether it respects a server-side cap, what happens if limit > cap, or whether a total_count or next_cursor is returned. memory_list_concepts and memory_list_patterns default to 100 items with no mention of pagination or offsets.
memory_stats and memory_health have minimal descriptions ('Get statistics about stored memories' and 'Health check endpoint for MemoryGate service', 31 and 38 chars respectively). Descriptions should explain what metrics are returned, when to call them, and what constitutes a healthy system.
Tool annotations present but incompletely applied. readOnlyHint and destructiveHint are used correctly, but there are no idempotentHint annotations. Tools like memory_store, memory_store_concept, and memory_register_ai should declare idempotency if they support upsert semantics or if repeated calls with the same input produce the same result.
Add pagination support documentation. For memory_search, memory_recall, and list_* tools, describe: 'Results are limited to [N] items. For more results, use offset/limit parameters (if supported) or refine filters.'
Enhance memory_stats and memory_health descriptions. Rewrite memory_stats: 'Returns aggregate statistics: total_memories, avg_confidence, total_observations, storage_breakdown by domain/type, and timestamp. Use to monitor system health and storage usage.' Rewrite memory_health: 'Health check endpoint. Returns {status: healthy|degraded|offline, uptime_seconds, last_check_timestamp, database_connected: boolean}. HTTP 200 = healthy, 503 = offline.'
Add idempotentHint annotations to write tools that support idempotent semantics. If memory_store deduplicates identical observations, add {'idempotentHint': True} to the annotations.
Document chaining and required downstream fields. If memory_search returns concept_ids, and memory_relate_concepts requires concept_ids, the search description should say: 'Returns concept_ids usable with memory_relate_concepts() and memory_alias_concept().' This prevents agents from making unnecessary lookup calls.
Add examples to descriptions of complex parameters. For 'threshold' in archive_memory, add: 'Example: {"field": "confidence", "operator": "<", "value": 0.5} archives all memories below 50% confidence.' This is more actionable than the bare description.
Clarify permission/scope requirements. Add a security section to each tool description noting what permissions are required (e.g., 'memory_search requires read:memory scope, archive_memory requires write:memory scope'). This enables least-privilege agent configuration.
Validate and document date/time parameter formats. For date_from/date_to in search_cold_memory, specify: 'ISO 8601 format (YYYY-MM-DD) or Unix timestamp in seconds.' Do not accept ambiguous formats.