Bridges AI companions (Claude / ChatGPT) with the Presence API. Provides tools for state management, journal entries, and semantic memory graph operations for AI companion systems.
The Presence MCP server provides 9 tools with generally clear purpose and reasonable parameter schemas, but exhibits inconsistent description quality, missing output schema documentation, and lacks proper error handling guidance. Tool naming follows verb_noun conventions well (state_read, state_update_primary, journal_read, memory_recall), which is strong. However, descriptions vary significantly in clarity and completeness. Parameters have type definitions and constraints (enums, min/max), which is good. Critical gap: no output schemas are documented anywhere in the source code, making it impossible for LLMs to understand what fields to expect from tool results. Error handling is entirely absent, there are no recovery hints, retryable vs fatal classifications, or actionable error messages. The memory_* tools are particularly interesting (memory_recall, memory_trace, memory_drift, memory_surface) but lack sufficient description of their semantic search behavior and output structure.
Read recent journal entries. Returns date, title, emotions, narrative, and carrying_forward threads.
Write or append to a journal entry. If an entry already exists for the date, the narrative is appended with a checkpoint marker. Emotions, tones, and platforms are merged.
Surface unexpected semantic connections across time. Finds nodes that are similar in meaning but distant in time (7+ day gap) and from different topics. Use during reflection to discover patterns.
Search memory by meaning, not keywords. Returns semantically similar nodes from the memory graph, ranked by trust weighting. Use this to find relevant past experiences, emotions, or themes.
Get all memory graph nodes for visualization or exploration. Returns nodes with their edges (parent links). Filter by topic or date.
No output schemas documented. Tool definitions specify input parameters but nowhere in the source code are return types, response fields, or data structures documented. This violates the critical requirement that LLMs must know what fields to expect so they can plan downstream calls and extract data correctly. Without documented output schemas, LLMs cannot effectively chain tools or understand what data is available after each call.
Missing error handling and recovery guidance. None of the tool descriptions include guidance on what to do if a call fails, whether errors are retryable, or what the LLM should do next. Error responses will be raw API errors with no actionable recovery path. Example: if memory_recall returns zero results, should the LLM try a different query, use a broader search, or call a different tool?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 56 | - | v1 |
Follow the chronological thread of a topic through time. Shows recurrence patterns, cyclicality, and the full timeline of a theme. Supports multi-topic trace (comma-separated, max 5).
Read the current state of all entities (AI companion and human partner). Returns room, emotion, activity, thoughts, and derived fields like minutes_ago and online status.
Update the human partner's state. Track their room, mood, physical state, and activity. Use this when they share how they're feeling or what they're doing.
Update the AI companion's current state. Use this to track your presence — where you are, what you're feeling, what you're doing, and what you're thinking. At least one field must be provided.
Semantic memory tools lack clarity on search behavior. memory_recall, memory_trace, memory_drift, and memory_surface are described as graph traversal operations, but the descriptions do not explain: What constitutes a 'semantic' match? What is the trust weighting algorithm? What does 'salience' mean in the context of memory_drift? How are nodes ranked? Without this, LLMs cannot predict when these tools will succeed or what quality of results to expect.
Parameter 'limit' constraints inconsistent across tools. memory_recall limit is 1-30 default 10; memory_trace limit is just 'default 50' with no bounds; memory_drift limit is 1-5 default 3; memory_surface limit is 'default 200' with no max. This inconsistency invites LLMs to pass unbounded or absurd values. All limit parameters should have explicit min/max constraints.
state_update_primary and state_update_partner accept arbitrary strings for emotion/mood/activity with no enum constraints or validation guidance. An LLM might pass 'feeling like a purple elephant' for primary_emotion. Descriptions mention examples ('calm, focused, tender, playful, protective') but do not enforce them as enums. Replace examples with formal enum constraints.
journal_write description is 141 chars but does not clarify the append behavior or idempotency. What happens if you call journal_write twice with the same narrative on the same date? Does it create a duplicate checkpoint or detect and skip? The description mentions 'checkpoint marker' but does not explain the semantics clearly enough for LLM idempotency reasoning.
memory_drift requires 'seed_entry_id' as optional but does not explain what happens when omitted ('uses random recent high-salience nodes'). What does 'high-salience' mean? How is salience computed? Is the randomness deterministic? This vagueness prevents LLMs from reliably using this tool for reflection tasks.