基于 LangGraph 的小红书运营 Agent (LangGraph-based Xiaohongshu operational agent)
Scoring was not performed
Tool names do not start with clear action verbs in English. 'generate_xhs_note' uses 'xhs' (Chinese platform abbreviation) instead of platform-agnostic naming. LLMs trained primarily on English may misinterpret 'xhs' without context.
Descriptions are entirely in Chinese. While this may be intentional for a Chinese platform tool, LLMs (especially English-centric models like GPT-4, Claude) have reduced comprehension and cannot properly evaluate when to call these tools. No English descriptions provided as fallback.
Parameter descriptions lack actionable constraints. 'topic' is described as 'content topic, e.g. how to make lattes at home' but does not specify length limits, forbidden characters, or format requirements. 'profession', 'personality', 'mood' lack enum constraints despite being categorical fields suitable for enums.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 22 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
No output schemas are documented. The descriptions claim outputs include 'title, body, tags' and 'lifestyle content with image generation prompts', but no structured output schema is visible. LLMs cannot plan downstream tool calls or extract required fields without knowing the response structure.
No error handling or recovery guidance. If the underlying Google Generative AI API fails, times out, or rate-limits the request, the tool definition provides no error classification (retryable vs. fatal) or recovery hints. LLMs will have no guidance on what to do next.
Optional parameters ('scene', 'content_type', 'topic_hint' in generate_lifestyle_content) lack clarity on how they influence output if omitted. No defaults documented; LLMs cannot reason about whether to include them.
No dependency hints between tools. generate_xhs_note takes only 'topic'; generate_lifestyle_content takes profession/age/gender/personality/mood. There is no guidance on when to use one vs. the other, forcing the LLM to infer intent.