A WeChat automation bot server that monitors groups, handles user interactions, manages menus, sends scheduled content (news, images, holiday information), and provides administrative controls through a database-backed system
This server exhibits critical deficiencies across all definition quality dimensions. Of 51 tools, the vast majority lack proper input schemas (visible in code), have minimal descriptions (10-30 chars), and show severe naming and composition issues. Tool definitions appear inferred from code rather than explicitly registered with full MCP schemas. No evidence of structured output documentation, error handling guidance, or parameter validation. The server is a WeChat bot framework exposing database operations and text processing functions without MCP-native tooling patterns.
Generate an admin chat response
Calculate character similarity between a keyword and content
Check game response based on columns, content, and sender
Check reward eligibility
Check if sender matches the given token
Process complaint supervision for content
Retrieve conditional picture or holiday task content
Input schemas missing or inferred for 35+ tools. Cannot verify parameter types, constraints, or descriptions in actual code. Tools like fetch_latest_storename, get_access_token, stop_op, and quit_admin show no visible schema registration.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 31 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Delete menu items from database by similarity matching
Handle complaint processing
Handle protection processing
Extract full menu items from menu text
Extract menu items from menu text
Fetch the latest store name from the server
Fetch menu items from the database
Get an access token from the server, updates internal token state
Retrieve the latest image in a specified directory
Retrieve the server authentication token
Get user menu mode status
Get user operation mode status
Get user operation-holiday mode status
Get user setup mode status
Get user table mode status
Insert menu items into the database
Insert order/table record into database
Check if the system is in holiday mode
Generate a menu-related chat response
Check menu validity by filename
Retrieve Weibo trending news content
Get non-zero columns from the database
Generate an order/table response
Prevent screen sleep and reduce brightness to 30%
Generate a price-related chat response
Check protection result for content
Quit administrative mode
Process random number court interactions
Send weekly holiday information to admin
Set the 4-digit administrator authentication code in the database
Set the nickname of the monitoring bot in the group
Add a group name to the monitored groups list in the database
Set user to menu mode
Set user to operation mode
Set user to operation-holiday mode
Set user to setup mode
Set user to table mode
Check if a user should be stored in the database
Start operation mode
Stop operation mode
Stop setup mode
Calculate similarity between two strings
Update menu items in the database
Generate a usual reply to user content
Descriptions uniformly too short (10-25 chars). Examples: 'Fetch the latest store name from the server' (41 chars, minimal), 'Stop setup mode' (14 chars, no context). No descriptions explain WHEN to use the tool, what it modifies, or dependencies. Violates baseline of 194 char average and 50-200 char optimal range.
Naming ambiguity and low specificity. Tools like 'usually_reply_content', 'protectresult', 'conditional_pic_task', 'complaint_supervise' lack action verbs or clarity. 'protectresult' as a name does not indicate whether it checks, applies, or retrieves protection status. No consistency with verb_noun pattern (get_, set_, check_).
No documented output schemas for any tool. LLMs cannot plan downstream tool calls or extract required fields. Example: does fetch_menu_from_db return a list of menu items? What fields does each item contain? Is there pagination? Unknown.
No error handling guidance. Tools that write or delete data (insert_into_menu_db, delete_from_menu_db_by_similarity, do_with_complain) do not document recovery paths, retry strategies, or what caused failures. Agents have no way to self-correct on errors.
Multiple tools perform same or overlapping functions without clear distinction. Examples: set_user_setup_mode, set_user_menu_mode, set_user_table_mode, set_user_op_mode, set_user_opholiday_mode are all variants of state-setting. Also get_user_* tools are redundant reads. LLM reasoning overhead with minimal differentiation.
No input validation or constraint documentation. Parameters like admincode, groupname, nickname lack format specifications, length limits, or regex patterns. 'Set the 4-digit administrator code' is placed in description but not enforced in schema.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Critical for agent safety: delete_from_menu_db_by_similarity is destructive but unmarked. Agents cannot distinguish safe reads from risky writes without explicit hints.
No composition support for common workflows. Inserting menu items (insert_into_menu_db) requires an array, but no pagination or batch fetch tool exists. Complex multi-step operations (menu CRUD, user state management) lack orchestration tools.
Security issues: tools accept user state and database operations without apparent permission checks or audit logging. Example: set_admincode directly sets a 4-digit code without verification that the caller is authorized. No scope declarations or permission gates visible.