LangChain-based agent with MCP tool servers for SSE and HTTP streaming, supporting RAG retrieval, file operations, shell execution, and basic utilities
The server exposes 21 tools with mixed definition quality. Naming is generally clear with action verbs (echo, build_uuid, get_date, create_file, delete_files, shell_exec). Descriptions vary widely: some tools have detailed docs (rag_retrieve, shell_exec, date_calculate), while others are minimal (build_uuid has one sentence, url_build is terse). Input schemas are present and properly typed for all tools, but parameter descriptions are inconsistent, some are excellent (rag_retrieve lists all 6 params with detailed guidance), others are sparse (build_uuid has no description for 'metadata' beyond 'optional metadata dictionary'). No output schemas are documented, which violates the pattern:tool-description and pattern:response-shaper rules. Error handling is absent from all tool descriptions, no guidance on what to do if a file doesn't exist, if an invalid expression is passed to calculator, or if shell_exec times out. Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are manually marked in the Risk field but NOT present in the MCP schema, suggesting they are inferred from the submission metadata rather than declared in the actual tool definitions. The server includes a few high-value tools (rag_retrieve, file operations, shell_exec) but lacks composition guidance for chaining them. Overall, this is a borderline C/D server, good names and schemas, weak descriptions and error handling.
向文件追加内容,可选不存在时创建,默认创建
Generate a random UUID4 string.
计算数学表达式,支持基本运算和常用数学函数,主要支持的是python语言的数值计算。 Args: expression (str): 要计算的数学表达式,支持以下运算: - 基本运算:+, -, *, /, ^, % - 括号:() - 数学函数:sqrt(x), sin(x), cos(x), tan(x), log(x), exp(x), abs(x) - 常量:pi, e - 支持角度符号:如 sin(30°) Returns: str: 计算结果的字符串形式,如果出错则返回错误信息
创建一个文件,可选择写入初始内容;默认不覆盖已存在文件
对日期进行加减计算(默认基于当前日期) Args: base_date: 基准日期(可选,格式如「2025-12-01」,不传入则是默认当前日期) days_diff: 加减天数(正数加,负数减,如 3=3天后,-1=昨天) Returns: str: 计算后的日期(中文格式) Example: date_calculate.run("2025-12-01", 3) → "2025年12月4日" date_calculate.run(days=-1) → "2025年11月30日"
使用该工具可删除一个或多个文件
No output schemas documented. Tools return results but no schema tells LLMs what fields to expect (e.g., does rag_retrieve return {contexts: [...], count: N} or {results: [...]}? Does echo return {text, metadata} or {output}?). This violates pattern:response-shaper and forces LLMs to guess field names for downstream tool chaining.
No error recovery guidance. Tool descriptions do not explain how LLMs should react to failures. E.g., read_file says 'reads files' but not what happens if the file doesn't exist or encoding fails. shell_exec documents return fields but not how to handle timeouts or permission errors. This violates pattern:recovery-guide.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Return the same text with optional metadata. 常用于:测试 MCP server 是否可用、agent 调试、链路对齐。
获取当前日期(年月日) Returns: str: 返回当前的日期信息(中文描述),格式如"2025年1月8日" Example: get_date() → "2025年1月8日" Note: - 此工具用于获取今天的日期 - 不要与 get_data(获取数据)混淆
获取当前时间 Returns: str: 当前时间(24小时制) Example: 例如:get_time.run() → "14:30"
执行知识库检索,返回参考相关文档片段及来源。 参数选择指南: - strategy: 'hybrid'(有明确术语/文件名/关键词)或 'vector'(概念性/语义理解问题) - top_k: 10-20(hybrid用15-20,vector用10) - alpha: 0.6(仅hybrid有效,越接近1越偏向向量) - use_rerank: 是否使用重排序 - rerank_top_n: 3(重排序后保留数量) 返回格式: {contexts: [文档内容...], count: 数量},contexts为空表示无相关内容。 如果检索失败,可先调用rag_rewrite_query重写query后再次检索。
将问题改写为更适合检索的query(不执行检索)。返回refined_query供后续使用。
这个工具读取文本文件的内容,支持 100MB 以内文件,默认 UTF-8 编码(若为 GBK 编码需手动说明)
删除对应路径的文件或目录(目录需 recursive=True)
Greet a user by username
按关键词搜索文本文件,返回匹配行及上下文
在指定的shell会话中安全地执行命令。 参数: command (str): 要执行的shell命令 timeout (int, optional): 命令执行超时时间(秒),默认30秒 work_dir (str, optional): 命令执行的工作目录,默认为当前目录 返回: dict: 包含以下字段: - success: bool, 命令是否成功执行 - stdout: str, 命令的标准输出 - stderr: str, 命令的标准错误 - return_code: 命令的返回码 - error: str, 如果发生错误,输出对应的错误信息 - os:操作系统 Example: shell_exec.run("ls -la") -> {'success': True, 'stdout': '...', 'stderr': '', 'returncode': 0}
Basic statistics for a numeric list.
对时间进行加减计算(默认基于当前时间) Args: base_time: 基准时间(可选,格式如「14:30」,不传入则是默认当前时间) hours_diff: 加减小时数(正数加,负数减,如 3=3小时后,-1=1小时前) minutes_diff: 加减分钟数(正数加,负数减,如 30=30分钟后,-15=15分钟前) Returns: str: 计算后的时间(中文格式) Example: time_calculate.run("14:30", hours=2, minutes=30) → "下午3:00"
Build url by merging params into base_url.
验证输入字符串是否为合法邮箱格式 Args: email: 待验证邮箱(如 "user@example.com") Returns: str: 验证结果(成功/失败原因) Example: validate_email.run("user@example.com") → "合法邮箱格式" validate_email.run("user.example.com") → "非法邮箱:缺少@符号"
使用这个工具可以将内容写入或者创建对应的文件
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are marked in the Risk metadata but NOT declared in the MCP tool schema. The MCP spec allows tools to declare these via outputMime and other hints. Currently, Risk is inferred from submission data, not from the actual MCP definition. LLMs cannot see these markers and must infer safety from descriptions alone.
Minimal descriptions for utility tools. build_uuid ('Generate a random UUID4 string'), url_build ('Build url by merging params into base_url'), stats ('Basic statistics for a numeric list'), say_hello ('Greet a user by username'), all under 50 chars and lack context for WHEN and WHY an LLM should use them. Per the rubric baseline (194 chars avg for A+ tools), these are too terse to drive confident tool selection.
Parameter descriptions are uneven. Some tools (rag_retrieve, date_calculate, time_calculate, shell_exec) have excellent param docs with constraints, ranges, and examples. Others (build_uuid 'optional metadata dictionary', url_build 'parameters to merge') are so generic that LLMs cannot infer valid values. Rubric baseline: 100% of A+ tools have param descriptions; this server is ~60%.
Destructive tools (delete_files, remove_path, shell_exec, write_file) lack confirmation/dry-run step. Per pattern:confirmation-request, irreversible operations should support a confirmation or dry-run mode to prevent agent mistakes. shell_exec especially, running arbitrary commands is high-risk and should require explicit user consent or a sandbox.