local-first browser-session forensics + repro for AI coding agents. peek's native messaging host + stdio MCP server — owns ~/.peek/sessions.db (better-sqlite3) and bridges the browser extension, CLI, and AI tools to one local source of truth.
The peek MCP server has 9 tools with basic descriptions and some schema coverage, but falls short of production-grade quality. All tools have descriptions (10-194 chars), and 7 of 9 have documented input schemas with type definitions. However, several critical gaps reduce confidence: (1) descriptions are terse and lack WHEN-to-use guidance, dependency hints, or prerequisite information; (2) error handling is not visible in the schema, no guidance on recovery, retryable vs fatal errors, or invalid input handling; (3) no output schemas are documented, forcing LLMs to infer response structure; (4) parameter descriptions are minimal; (5) destructive operations (delete_session, export_session) lack confirmation/dry-run patterns; (6) the search_sessions tool has complex nested filter objects but no validation guidance. The tools follow verb_noun naming and resource-focused design, which is good, but lack the LLM-optimization depth needed for confident agent use. This server is mid-range for community tools, functional but requires human oversight and iteration.
Retrieve the audit log of all peek activities and database modifications
Delete a session from ~/.peek/sessions.db by ID
Export a session as a self-contained tarball (.tar.gz) or HTML file suitable for sharing and replay
Generate a Playwright test script that replays a captured session's user interactions programmatically
Retrieve a specific session by ID from ~/.peek/sessions.db, including rrweb events and metadata
Retrieve the gzipped rrweb event blob for a session, suitable for use with the rrweb player
Get the recording status and metadata of a session
List all captured browser sessions stored in ~/.peek/sessions.db
No documented output schemas. Users cannot see what fields tools return, forcing LLMs to infer response structure and plan downstream calls blindly.
Destructive operations (delete_session, export_session) lack confirmation or dry-run patterns. Agents cannot safely preview consequences before executing irreversible actions.
No error handling guidance. Tools do not document what errors can occur, whether they are retryable, or what the agent should do next. Example: delete_session provides no guidance on 'session not found' or 'permission denied' scenarios.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
Search sessions by title, URL, or tags
Descriptions lack WHEN-to-use guidance and prerequisites. E.g., 'get_session' states what it returns but not when to call it instead of search_sessions, or whether it requires get_session_status to be called first.
audit_log tool has no input schema visible. Cannot determine what parameters it accepts (filter by user? by timestamp? by action?). Schema score forced to 0.
search_sessions filter parameter has complex nested objects (min_duration, max_duration, start_time, end_time) with no validation guidance. No documentation of format (milliseconds vs seconds?), ranges (0-?), or mutual exclusivity.
Parameter descriptions are minimal (1-2 lines). Descriptions like 'The session ID to retrieve' lack actionable detail: can it be a partial match? Is it case-sensitive? What if it contains special characters?
export_session format enum is documented (tar|html) but no description of when to use each format, file size implications, or replay compatibility.