Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
The xiaohongshu_scraper MCP server has severe quality issues across naming, descriptions, and schemas. While 4 tools are defined with basic verb-starting names, descriptions are minimal (10-16 characters), parameters lack proper type definitions in most cases, and output schemas are completely undocumented. The server lacks input validation, error guidance, and follows few of the 54 agentic tool patterns. Parameter schemas for search_notes are partially present but inconsistent. The implementation mixes stateful global browser context with MCP's stateless protocol, creating architectural mismatches. Tool descriptions are generic Chinese text that provides no actionable guidance for LLM selection or use.
Missing input/output schemas for 3 of 4 tools (login, get_note_content, get_note_comments). Only search_notes has partial schema definition. Schema score must be 0 when no input schema is visible.
login tool has empty input schema ({}), violating the requirement that every tool must accept parameters or explicitly document zero parameters. No description of what happens during login flow or how long user should wait.
Expand tool descriptions to 100-200 characters in English, explaining WHAT the tool does, WHEN to use it (in preference to similar tools), and WHAT it returns. Example for search_notes: 'Search Xiaohongshu notes by keyword. Returns up to <limit> results with URL and title. Use this before get_note_content to discover notes. Limited to 100 results per call.'
Define explicit input and output JSON schemas for all four tools. For login: document that it returns a status string and blocks for up to 180 seconds. For search_notes: document that output is an array of {url: string, title: string} objects, sorted by relevance. For get_note_* tools: document that output includes {title, author, publish_time, content, comments} as structured fields, not free-form text.
Add parameter constraints: search_notes limit must be 1-100 (document this in schema and description). Provide enums where applicable (e.g., if note_type filtering were added, use enum not string).
Implement per-tool error handling that returns structured error objects: {error_code: string, message: string, recovery_hint: string, retryable: boolean}. Example: if login times out, return {error_code: 'LOGIN_TIMEOUT', message: 'User did not complete login within 180 seconds', recovery_hint: 'Ask the user to complete browser login manually or try again', retryable: true}.
Refactor login tool to either (a) use MRTR protocol with result: input_required to ask the client to prompt the user for manual browser action, or (b) remove it from the tool interface and handle login as a server startup task before accepting LLM requests.
Score history
Overall score trend
First recorded score · v2 rubric
39/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
39
2026-07-28+
v2
No output schemas documented for any tool. LLMs cannot determine what fields to expect or how to chain results into downstream calls. Returns are unstructured strings (search_notes returns formatted text, login returns status messages).
get_note_content and get_note_comments accept only 'url' parameter but source code shows logic attempts to extract metadata (title, author, timestamp) via fragile Playwright selectors. These fields are not documented in the schema, making the tool's actual return value unknown to LLMs.
Error handling is absent. Tools return raw exception strings (e.g., 'get_note_content时出错: {str(e)}') with no guidance on whether the error is retryable, user-fixable, or fatal. No recovery paths suggested to the LLM.
search_notes accepts 'limit' parameter with default=5, but no maximum bound is enforced or documented. LLM could pass arbitrary large values, potentially causing timeouts or excessive API load.
Stateful browser context stored in global variables (browser_context, main_page, is_logged_in) violates MCP stateless protocol. Each request should be independent; sharing browser state across calls risks race conditions, session conflicts, and failures when multiple LLM agents invoke tools concurrently.
login tool requires manual user interaction (3-minute wait for browser login) but is exposed as a tool the LLM can invoke. No clear mechanism for LLM to communicate 'waiting for user' to the client, violating single-turn tool execution model. Should use MRTR (result: input_required) or be excluded from LLM tools.
login
Eliminate global browser state. Refactor to either (a) spawn a new browser per request and clean up after, or (b) use a session token returned by login and passed to subsequent tools. Option (b) requires modifying search_notes, get_note_content, get_note_comments to accept an optional session_id parameter.
Document output pagination for search_notes. If the API can return >100 results, add offset/cursor and total_count fields to the response schema.
Add input validation in each tool. Example: validate that 'url' parameters in get_note_content and get_note_comments start with 'https://www.xiaohongshu.com' to prevent injection attacks.
Replace generic error messages ('搜索笔记时出错') with specific, actionable ones. Example: 'Search failed: network timeout after 60 seconds. This is retryable, try again in 10 seconds.' or 'Search failed: invalid keyword characters. Keywords must be alphanumeric or Chinese characters.'
Add tool annotations (toolAnnotations) to mark read-only vs. destructive operations. All four tools are marked READ_ONLY in the Risk field but this is not exposed in the MCP protocol. Use readOnlyHint: true in the tool definition so clients know these tools are safe to auto-invoke.