Multi-package agent trajectory processing system with MCP server implementations for sandbox, recorder, reward, and hub orchestration
The knowlyr-gym MCP server exposes 24 tools across 5 packages (sandbox, recorder, reward, hub). While the server demonstrates substantial functionality for agent trajectory processing, critical definition gaps limit production readiness. Most tools lack comprehensive parameter descriptions and have minimal error guidance. Schemas are present but often incomplete. Descriptions are present but frequently generic or terse. The server targets a specialized domain (agent training/evaluation) but does not compensate for definition gaps with domain clarity.
从多条轨迹构建偏好对 (用于 RLHF/DPO 训练)
将自动 Reward 与人工标注进行校准
将 Agent 日志转换为标准化轨迹格式
创建 Docker 沙箱执行环境
在沙箱中执行工具 (file_read, file_write, shell, search, git)
将轨迹数据导出为指定的训练格式 (SFT / DPO / Benchmark / HuggingFace)
读取文件内容
写入文件内容
Generic descriptions that lack context for LLM tool selection. Many tools use single-clause descriptions (e.g., 'get_schema' → no description provided in visible code, 'list_rubrics' → empty properties object). LLMs cannot determine WHEN to call these tools or what they return without rich descriptions.
Missing or incomplete parameter descriptions across all tools. Parameters like 'params' in execute_tool are described generically ('工具参数' / tool parameters) without explaining the expected structure, constraints, or valid options. LLMs cannot infer what to pass.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | 1.0+ | v1 |
返回标准化轨迹的 JSON Schema 定义
执行 Git 操作
列出可用的评估 Rubric 维度
查看 Pipeline 执行状态和进度
处理单个 Agent 日志文件,解析并评分生成标准轨迹
批量处理 Agent 日志目录,解析并评分生成标准轨迹
对比两条轨迹的差异 — 步骤数、工具使用、成功率等维度对比
在沙箱中重放 Agent 执行轨迹
重置沙箱到初始状态
从多条轨迹生成奖励排行榜 — 按 Reward 分数排序,对比不同模型/策略的表现
运行完整的 Agent 轨迹数据 Pipeline (Task -> Sandbox -> Recorder -> Reward -> Export)
保存沙箱当前状态快照(文件系统 diff + 环境信息)
对单条 Agent 轨迹计算过程级 Reward
搜索代码内容
执行 Shell 命令
验证日志文件是否为指定的 Agent 框架格式
No error handling guidance. Tools that write to filesystem, spawn Docker containers, execute shell commands, and modify agent state offer no description of failure modes, recovery paths, or what to do if a step fails. Error responses visible in source are minimal or absent.
No documented output schemas. Tools return JSON objects but descriptions do not specify field names, types, or structure. For example, 'recorder_diff' returns comparison data but nowhere states whether the response includes arrays, objects, numeric metrics, or text summaries. LLMs must guess.
Complex nested parameters with no documentation. 'run_pipeline' accepts 'agents' as an array of objects with 'framework' and 'model' fields, and 'trajectory' parameters are described as bare 'object' types without specifying required subfields, constraints, or examples. LLMs cannot construct valid calls.
Enum constraints present but incomplete. 'framework' parameter in convert_logs accepts ['openhands','swe-agent'] but process_log accepts ['openhands','sweagent','swe-agent'], inconsistent enum values across similar tools force LLMs to reason about subtle differences. Should standardize.
Destructive and reversible operations (file_write, shell, git, create_sandbox) lack confirmation or dry-run patterns. No indication that these operations have side effects, are irreversible, or require caution. Descriptions do not state 'This modifies state' or 'Cannot be undone.'
No pagination or result limits documented. Tools that process multiple files/trajectories (process_logs_batch, reward_leaderboard) do not specify max result count, pagination support, or how LLMs should handle large result sets. Can cause context window exhaustion.