Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
ClaudeR MCP exhibits mixed quality across 16 tools. The server provides clear descriptions for most tools (94% have descriptions) and reasonable parameter schemas, but falls short in consistency, output schema documentation, and error handling guidance. Tool naming follows conventions (verb-based), but several tools lack depth in parameter descriptions and validation guidance. No per-tool output schemas are documented in the visible code, forcing LLMs to infer response structures. Parameter descriptions are functional but generic, many lack format hints, constraints, or examples of valid values. The annotation workflow (load_annotation_data → annotate → export_annotations) demonstrates composition understanding, but checkpoint/restore operations lack sufficient guidance on expected outcomes and failure modes. Security considerations exist (token-based session binding, authenticated URLs) but are not prominently documented as part of the tool descriptions.
Tools (16)
annotatewritesource verified68/100
Process the next row of loaded annotation data through the agent and store the result
checkpoint_sessionwriteauthsource verified68/100
Save a snapshot of the R global environment to a timestamped .RData file for later restoration
connect_sessionwritesource verified70/100
Bind this MCP connection to a specific R session by session name
execute_rwriteauthsource verified72/100
Execute R code through the RStudio addin HTTP server
execute_r_with_plotwriteauthsource verified77/100
Execute R code and capture the resulting plot as a base64-encoded PNG image
export_annotationswrite50/100
Export completed annotations to a CSV file
generate_codebookwritesource verified70/100
Generate a codebook and reproducibility README for a research project by scanning scripts for packages, data inputs, and outputs
Output schemas not documented in source code. No visible tool response schemas, return types, or field descriptions. LLMs must infer the structure of execute_r, list_sessions, search_citations, and checkpoint responses from context, increasing hallucination risk.
Parameter descriptions lack validation constraints and format hints. 'code' parameter in execute_r says 'R code to execute' but does not specify: character encoding, maximum size, whether multiline is allowed, or error behavior for invalid syntax. 'label' in checkpoint_session lacks length/format constraints. 'query' in search_citations omits min/max length.
Document the output schema (fields, types, structure) for each tool. Create a JSON schema or brief markdown table showing what execute_r, list_sessions, search_citations, and checkpoint_session return. Example: 'list_sessions returns [{session_name: string, port: int, pid: int, started_at: ISO8601, token?: string}, ...]'.
Add validation constraints to all parameters. For execute_r: 'code (string, 1 - 100,000 chars, multiline allowed, UTF-8 encoded)'. For search_citations: 'query (string, 1 - 500 chars, no wildcard required)'. For checkpoint_session: 'label (string, 0 - 100 alphanumeric + dash + underscore)'.
Add error handling guidance to each write operation. Example for execute_r: 'If code raises an error, returns {success: false, error_message: "..."}. LLM should interpret this as a user-fixable issue and offer to modify and retry. If timeout (60s), tool is unresponsive, suggest reconnecting via connect_session.'
Document session binding semantics in connect_session and execute_r descriptions: 'Once a session is bound via connect_session, all execute_r and execute_r_with_plot calls route to that session. If the session crashes, subsequent calls raise SessionBindingError (session lost). Retry connect_session with list_sessions to find a live target.'
Clarify annotation workflow state machine. Add to load_annotation_data: 'Loads CSV into memory. Returns {total: int, current_row_index: 0}. Call annotate() once per row in order.' Add to annotate: 'Processes current row. Increments index. If all rows complete, returns {complete: true}.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 54 points across a rubric change (v1 → v2)
54/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
54
<=2025-11-25
v2
2026-03-09
F
0
-
v1
get_bibtexread onlysource verified73/100
Fetch the canonical BibTeX entry for a DOI via content negotiation with doi.org
list_session_checkpointsread onlyauth50/100
List available session checkpoints with file, time, and size metadata
list_sessionsread onlysource verified67/100
List all available live R sessions with their metadata (name, port, PID, start time)
Error handling and recovery guidance missing. None of the tool descriptions indicate what happens on failure (e.g., when R code fails to execute, when a session is lost, when checkpoint size exceeds max_gb, when DOI lookup fails). No guidance for LLM on retry vs. user correction vs. fatal errors.
Session binding and state management not clearly documented. connect_session and get_r_addin_url behavior (fail-closed semantics, SessionBindingError handling, _target_session persistence) are visible in code but NOT described in tool descriptions. LLMs cannot understand when a session is lost or how to recover.
Composition workflow for annotation (load_annotation_data → annotate loop → export_annotations) lacks guidance. No description clarifies: how many rows can be loaded, what happens if annotate is called out of order, whether partial annotations are saved, or what export_annotations returns. Forces LLM to guess state machine.
No idempotency documentation. execute_r, checkpoint_session, and annotate are write operations, but descriptions do not state whether they are idempotent or have side effects on retry. Agents that retry on ambiguous failures risk duplicate R commands, duplicate checkpoints, or duplicate annotation entries.
Deprecated MCP 1.x decorator API. pyproject.toml specifies 'mcp>=1.0.0,<2' with explicit comment that MCP 2.0 removed Server.list_tools/call_tool decorators. Server is built on deprecated 1.x patterns. Spec 2026-07-28 has moved to stateless request handling. Migration to 2.x is blocking.
all
Add idempotency statements. For execute_r: 'Each call executes code independently. Identical calls produce identical output (idempotent if code is pure). Non-pure code (e.g., incrementing a counter) may produce different results on retry.' For checkpoint_session: 'Each call creates a new timestamped checkpoint. Retry with identical parameters creates a duplicate checkpoint (not idempotent).'
Upgrade from mcp 1.x to 2.x. Replace Server.list_tools/@server.call_tool decorators with the current stateless request model. Ensure each request is self-contained (no _target_session state leakage). Target MCP spec 2026-07-28.
Clarify session discovery fallback behavior. Add to list_sessions: 'Returns empty list if no sessions found. If multiple sessions exist and none are explicitly bound via connect_session, subsequent execute_r calls raise SessionBindingError (refusing implicit routing). User must call connect_session first.'
Add parameter relationship documentation. For restore_session: 'If checkpoint is omitted, uses most recent. If clear=true (default), environment matches checkpoint exactly (backup-and-restore). If clear=false, merges checkpoint over current (may leave newer objects).' This prevents silent misuse.
Document resource limits for data operations. For generate_codebook: 'Scans up to max_files (default 20) data files. If project exceeds this, prioritize input files via data_files param. Output file size capped at 1MB; reports omitted if exceeded.' For load_annotation_data: 'CSV loaded into memory; max 100,000 rows or 500MB (whichever first). Returns error if exceeded.'
Add structured response examples in documentation (outside tool descriptions). Publish a schema guide showing: list_sessions response structure, search_citations result format, checkpoint metadata fields. This helps agents learn the interface faster.
Implement tool annotations (new in MCP 2.x). Mark execute_r and execute_r_with_plot as destructiveHint: true. Mark list_sessions, search_citations, get_bibtex as readOnlyHint: true. Mark restore_session (with backup=true default) as idempotentHint: true. This helps hosts decide when to require confirmation.