Trusted engineering context for agentic software engineering via temporal knowledge graph (Neo4j) and Gemini API. Serves MCP tools for codebase analysis, decision tracking, and impact prediction.
memex demonstrates good intent with 11 well-scoped tools covering read, write, and analysis operations on a knowledge graph. Tool naming follows verb_noun conventions consistently (get_*, search_*, record_*, link_*, explain_*, predict_*). Descriptions are substantive (most 150-300 chars, well above the 20-char floor) and explain WHAT the tool does and WHEN to use it. Input schemas are present and typed for all tools with parameter descriptions. However, there are structural gaps: (1) Output schemas are not documented, the source shows only input schemas; (2) Error handling guidance is absent, tools mention they return Markdown or structured output, but no description of failure modes or recovery paths; (3) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite the split between 8 read-only and 3 write tools; (4) Parameter relationships (e.g., repo vs project mutual exclusivity in several tools) are noted in descriptions but not consistently; (5) No documented pagination or result limits for search_context and get_engineering_context beyond 'top_k' capping at 8. The write tools (record_problem, record_decision, link_entity) mention dedup and write policy layers but do not document what those policies reject or how users should respond. Overall, the tools are well-named and described, but fall short of A-tier due to missing output schemas, no structured error recovery guidance, and lack of tool annotations.
Cross-references a commit's diff with Decision/Problem nodes linked to the affected files and returns a Gemini-Pro-synthesised explanation. Returns a Markdown string under ~2000 tokens grounding the change in graph context.
Return the same bounded packet projection used by the provider path. Returns bounded engineering context based on query, repository/project scope, and optional task/session IDs. Supports historical context retrieval.
Returns open Problem nodes (not yet resolved). Optionally filters by module and includes structured metadata about each problem.
Returns a structured markdown briefing of the project including node counts, active modules, recent decisions, open problems, and stale edges. Optionally includes cluster-level context when available.
Returns Decision nodes created within the last N days, newest first. Runs conflict detection to flag contradictory Decisions with overlapping validity windows. Optionally filters by module and returns only corroborated decisions.
Output schemas not documented. Tool descriptions state return types (Markdown strings, lists, structured metadata) but no JSON Schema for response fields is provided. LLMs cannot plan downstream chaining (which fields to extract, which tools to call next) without documented output structure.
Error handling not documented. Write tools (record_problem, record_decision, link_entity) mention 'intent-confirmation dedup layer' and 'write policy' but do not document: (a) what the dedup layer returns on near-duplicate detection (options for corroborate/supersede/force-proceed), (b) what write policy violations return, (c) which errors are retryable, (d) recovery guidance for the LLM. Agents cannot self-correct without explicit error categorization.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 64 | 2026-07-28+ | v2 |
Returns everything the graph knows about a specific symbol including callers, callees, linked decisions and problems. Falls back to search if symbol not found.
Creates or updates an edge between two entities in the graph. Supports relationship types: CALLS, IMPORTS, EXPORTS, DEPENDS_ON, RELATES_TO, MOTIVATES, MENTIONS, CAUSED_BY. Enforces write policy based on relationship type and principal role.
Returns modules likely affected by changes to a file_path based on historical coupling (calls + exports + imports + decision links). Pure graph traversal — no LLM call. Returns a ranked Markdown list under ~2000 tokens.
Creates or updates a Decision node in the graph. Includes intent-confirmation dedup layer (Layer B) — if a near-duplicate Decision is found, returns options for corroborate/supersede/force-proceed before writing. Enforces write policy (Layer A) based on node type and principal role.
Creates or updates a Problem node in the graph. Includes intent-confirmation dedup layer (Layer B) — if a near-duplicate Problem is found, returns options for corroborate/supersede/force-proceed before writing.
Full-text semantic search over the knowledge graph. Returns matching entities with scores ranked by relevance. Integrates with both keyword and vector search capabilities.
Tool annotations missing. The server clearly delineates read-only tools (8) and write tools (3), but neither the tool definitions nor descriptions include readOnlyHint, destructiveHint, or idempotentHint. This forces LLMs to infer action type from description parsing rather than structured metadata.
Pagination/result limits not enforced or documented at tool level. get_engineering_context and search_context both perform full-text search but only mention top_k capping at 8; no documented pagination mechanism (cursor, offset) for results beyond that limit. Large knowledge graphs could return unbounded results.
Parameter mutual exclusivity documented in text but not enforced. Several tools (get_engineering_context, get_project_context, get_symbol_context, get_recent_decisions, get_open_problems, search_context) note 'repo' and 'project' are 'mutually exclusive' in descriptions, but no validation schema (JSON Schema oneOf/enum constraints) prevents LLMs from passing both.
link_entity relationship_type parameter could benefit from stricter enum validation. Description lists 8 valid types (CALLS, IMPORTS, EXPORTS, DEPENDS_ON, RELATES_TO, MOTIVATES, MENTIONS, CAUSED_BY) but source does not confirm these are enforced as an enum in the JSON Schema, risking typos or hallucinated relationship types.