A VSCode/Cursor extension that acts as a "Long-Term Pair Programmer". As you code, it quietly observes and extracts EventLog Memory (architectural decisions, bug fixes) and Profile Memory (your preferred coding style). When you ask it a question, it uses EverMemOS to retrieve project-specific context that standard LLMs forget.
CodeMem provides 6 tools with reasonable naming and descriptions, but exhibits significant gaps in schema completeness, parameter documentation, and error handling. Tool names follow verb_noun convention (add_episodic_memory, search_project_memory, etc.), which is correct. However, several parameters lack descriptions, some enums are inconsistent across tools (e.g., memory_type uses different values in different tools), and output schemas are not documented. The server demonstrates moderate quality but falls short of production readiness.
Extract and store an episodic memory (event, decision, or bug fix) to EverMemOS. Automatically tags with git context (branch, commit), affected files, and component layer.
Delete memories by ID or in bulk by type/group. Supports selective deletion or group-wide purge.
Retrieve profile memories that capture the user's coding style, preferences, and patterns. Returns accumulated profile data across sessions.
List recent memories with optional filtering by type and scope. Supports pagination.
Search for relevant memories across the project using hybrid retrieval (keyword + vector). Filters by memory type, scope (session/repo/all), and optional date range. Deduplicates and filters out superseded memories.
Update profile memory with user coding style observations and preferences.
Inconsistent enum values for memory_type across tools. add_episodic_memory uses [decision, bug_fix, feature, refactor, incident], while list_recent_memories uses [episodic_memory, event_log, profile, foresight]. LLMs will be confused about which values are valid for which tool.
update_profile_memory has no parameter descriptions at all. The 'style_observations' and 'preferences' parameters are both documented as objects with no explanation of what keys/fields they should contain. LLMs cannot infer the expected structure.
No output schemas documented for any tool. The rubric requires LLMs to know what fields to expect. Without documented return types, agents cannot reliably chain calls or extract the right data for follow-up operations.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 44 | - | v1 |
delete_memories accepts '__all__' as a memory_id to trigger bulk delete, but this magic string is mentioned only in the description. No clear schema constraint prevents invalid IDs. An LLM could pass any string and expect bulk deletion. Needs explicit documentation or a separate bulk_delete parameter.
Destructive operation (delete_memories) lacks a confirmation step or dry-run mode. An agent could accidentally delete all memories without recovery. Needs confirmation_required parameter or a separate confirm_delete tool.
get_profile_memory and update_profile_memory use 'scope' parameter but list_recent_memories uses 'scope' as an enum with 'session|repo|all'. Scope behavior and semantics should be consistent across all tools that use it. Documentation does not clarify what each scope means.
search_project_memory accepts retrieve_method with enum [hybrid, agentic, keyword, vector], but the description says 'auto-selected based on query complexity if not specified'. No guidance on when to use which method. LLMs will not know how to choose.
No error handling documented. Tools do not describe what errors they may return, whether errors are retryable, or what the LLM should do on failure.
list_recent_memories and search_project_memory do not document pagination limits (max page size, max result count). Without explicit limits, LLMs could request 10,000 results, exhausting context windows or overwhelming the backend.