MCP server that bridges Claude Code to Telegram for remote user interaction
Two tools with basic descriptions and schemas present, but multiple critical gaps prevent higher scoring. Tool naming is clear and action-oriented (ask_user, notify_user), but parameter descriptions are minimal and lack constraints. Input schemas use Zod validation, which is good for runtime safety, but descriptions are sparse. No documented output schemas, no error recovery guidance, and critical issues around timeout handling and idempotency. Average tool description length is ~120 chars (within baseline), but parameter annotations are thin (5-25 chars). Missing per-tool security implications, no dependency hints, and no guidance on what to do when ask_user times out. This is solidly in the 'fair' band for definition quality.
Send a question to the user via Telegram and wait for their response. Optionally include inline buttons for quick replies. Times out after 10 minutes.
Send a notification message to the user via Telegram. Does not wait for a response.
Input schemas present in Zod but parameter descriptions are bare/minimal. 'message' param described only as 'The question or message to send' (40 chars) and 'The notification message to send' (32 chars). No format constraints, no guidance on max length, encoding expectations, or special character handling. LLMs cannot infer whether markdown is supported, whether newlines are safe, or what happens if message exceeds platform limits.
'buttons' parameter in ask_user has thin description ('Optional list of button labels for quick replies (inline keyboard)' = 65 chars) but lacks constraints: no min/max array length, no max label length, no guidance on what happens if a label is too long for Telegram's inline keyboard. LLMs may pass buttons=['very long button text that exceeds telegram limits'].
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 7 | - | v1 |
No documented output schema. Tool descriptions say what is returned ('Send a question to the user ... and wait for their response'), but nowhere does the schema define: what is the structure of the response? Is it always text? Can it be null if timeout occurs? What fields does the response object contain? The code returns {content: [{type: 'text', text: response}]}, but an LLM reading just the description has no way to predict this structure for downstream chaining.
Timeout handling is non-idempotent and breaks agent retry semantics. ask_user waits 10 minutes, then rejects. If an agent retries immediately after the first timeout, the second ask_user call will ALSO wait 10 minutes, creating a 20-minute hole before the agent can move on. No idempotency key, no retry guidance in error message. Pattern:idempotent-operation violated.
Error responses lack recovery guidance. The timeout error 'Timed out waiting for user response (10 min)' tells the LLM the operation failed but does NOT say: Should I retry? Should I ask the user to check Telegram? Is this a transient network issue or a user availability issue? Pattern:recovery-guide violations prevent agents from making intelligent next moves.
No error classification. Both tools can fail (Telegram API error, polling error, network timeout), but error responses don't categorize failures as retryable vs fatal. An LLM cannot distinguish between a transient Telegram server issue (retry) and a misconfigured CHAT_ID (fatal). Pattern:error-classification missing.
notify_user lacks destructive/write implications in description. The tool sends a message, a stateful operation with side effects. The description should explicitly say 'This is a write operation that sends a message to the user. It cannot be undone. Verify the message content before calling.' LLM needs to know this is not a dry-run or preview.
No confirmation or dry-run pattern for ask_user. This tool initiates a blocking wait for user input. An agent error (e.g. malformed question) could lead to the user receiving a confusing or nonsensical message and the agent waiting 10 minutes for a response. No preview step, no confirm-before-execute pattern.
Single-user design with race condition vulnerability. The global 'pendingResolve' variable assumes only one concurrent ask_user call. If an agent calls ask_user twice in rapid succession (or if multiple agents access the same MCP instance), the second call overwrites pendingResolve, causing the first ask_user to never resolve. No queue, no request IDs, no per-call context isolation.
HTML escaping in ask_user is applied to the message, but the description says 'parse_mode: HTML'. This implies the user can pass raw HTML markup. However, unescaped HTML in button labels or user input could lead to injection attacks or malformed Telegram messages. Input validation policy is unclear.