A Model Context Protocol server allows to interact with Twitter, enabling posting tweets and searching Twitter.
The server has 4 tools with explicit registration in src/mcp.ts. All tools have names, descriptions, and input schemas visible. Naming follows verb_noun convention (retweet, like_tweet, post_tweet, search_tweets), which is good. Descriptions are present but brief (10-50 chars), below the 194-char baseline for A+ tools. Input schemas are properly typed with JSON Schema constraints (minLength, maxLength, minimum, maximum), which is a strength. However, no output schemas are documented, responses are not specified, forcing LLMs to guess what fields to expect. Error handling guidance is absent. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite having both READ_ONLY and WRITE operations. Parameter descriptions are minimal ('The ID of the tweet to retweet' is functional but not LLM-optimized). The post_tweet tool has complex nested array constraints (images with maxLength 4), properly specified but could benefit from clearer documentation. No dry-run or confirmation pattern for destructive operations (retweet, like, post are state-modifying). Field naming is inconsistent across tools (some expect tweetId, but response field names are not visible). Output response schemas are completely absent from the code.
Like a tweet on Twitter/X
Post a new tweet to Twitter/X
Retweet a tweet on Twitter/X
Search for tweets on Twitter/X
No output schemas documented. Tools do not specify what fields they return, forcing LLMs to guess downstream field names and types. This breaks tool chaining and causes hallucination of non-existent fields.
Descriptions are too brief (10-50 chars vs. 194-char baseline). 'Like a tweet on Twitter/X' does not explain WHEN to use this vs other tools, what it returns, or consequences. Does it fail if already liked? Can it be retried? LLMs lack necessary context.
No tool annotations despite having destructive operations. retweet, like_tweet, and post_tweet are state-modifying (WRITE risk), but lack destructiveHint or idempotentHint annotations. This prevents agents from reasoning about consequences and retry safety.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 58 | - | v1 |
No error recovery guidance. If a tweet is not found or already retweeted, what should the agent do? Descriptions lack actionable error hints. No support for dry-run or confirmation patterns for destructive operations.
Parameter descriptions lack constraint clarity. search_tweets 'count' parameter specifies min/max in schema (10-100) but description does not state this range. LLMs may infer invalid values from similar tools or patterns they've seen. Format 'Number of tweets to retrieve (10 - 100)' is needed.
post_tweet images parameter has complex nested constraints (array with maxLength 4) but lacks clarity in description. Description says 'URLs of images to upload' but does not specify max 4 images, accepted formats, or size limits. Agents may attempt to upload 10 images and fail.
search_tweets lacks pagination support. No offset, cursor, or page parameters visible. Returning unbounded results could blow context window. If Twitter API returns 100+ tweets, tool should enforce limits and offer pagination guidance.