Lightweight RunGuardian/SciPilot-style training run wrapper for experiment tracking, AI-driven analysis, and report generation
sciagent has critically weak tool definitions across all quality dimensions. All 8 tools lack proper input schemas with type information, parameter descriptions exist but schema definitions are either missing or incomplete. Tool names are generic and lack clear action verbs. Descriptions are present but superficial (10-50 chars), providing minimal guidance for LLM tool selection. No output schemas are documented. The server is missing fundamental MCP quality baselines.
Automatically track all local variables as parameters (decorator/context manager)
Log arbitrary metadata about the run (e.g., commit hash, git branch, environment info)
Log a single training metric
Log multiple training metrics at once
Log a single training parameter
Log training parameters/hyperparameters for tracking
Save all tracked parameters and metrics to metrics.json
All tools use generic names (log_*, auto_track, track) that do not clearly convey action + resource. Names like 'log_param' lack specificity, is this logging a training parameter, API parameter, or configuration? 'track' is vague without an object. The rubric baseline shows 90% of A+ tools start with action verbs like 'create_', 'get_', 'send_'. These names do not meet that bar.
Input schemas are severely incomplete. Tools like 'save' and 'auto_track' have empty input objects ({}) with no parameters defined, yet the tool descriptions suggest they accept data. For 'log_params' and 'log_metrics', the 'kwargs' parameter is defined as type 'object' with a description, but no JSON Schema is visible, no properties, no required fields, no type constraints. The rubric requires 'Every parameter needs a description explaining what it controls' AND formal type definitions. These tools fail both criteria.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 32 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Decorator to automatically track function parameters and execution
Descriptions are too brief (10-50 characters). Baseline for good descriptions: 194 chars (p10=34, p90=392). Examples: 'Log training parameters/hyperparameters for tracking' (52 chars) does not explain WHEN to use this vs log_param (singular), WHAT format the kwargs should take, or WHAT is returned. Descriptions must be 10 - 1024 chars and answer: What does it do? When should the LLM call it? What does it return?
No output schemas are documented. The rubric states '100% of A+ tools have documented return types.' For example, does 'save' return success/failure, a file path, a summary, or nothing? Does 'log_param' confirm the value was logged? Without documented outputs, LLMs cannot chain tools or extract data for downstream calls.
Parameter types are either missing or use generic 'object' declarations. JSON Schema requires 'type' (string, number, boolean, object, array, etc.) on every parameter. For 'log_param', the 'value' parameter is declared as type 'any', this is not a valid JSON Schema type and provides zero constraint. For 'kwargs' parameters, type is 'object' but with no 'properties' defined, making them useless to validate or document.
No error handling or recovery guidance. Tools have no documented error conditions, retry logic, or guidance for LLMs on what to do if a call fails. The rubric requires 'Error responses must tell the LLM what to do next.' For example, if 'log_metric' is called with an invalid metric name or non-numeric value, what is returned? How should the agent recover?
Composite vs singular tool distinction is unclear. The server offers both 'log_param' (singular) and 'log_params' (plural), but there is no guidance on when to use which. The rubric states 'When multiple tools operate on the same resource, their names must make the distinction obvious.' Current names are confusingly similar. Additionally, there is no batch variant that accepts arrays, if an LLM needs to log 10 metrics, it must call 'log_metric' 10 times instead of a single batch call.
'track' tool has a 'func' parameter of type 'callable', an invalid JSON Schema type. The description says 'Function to track' but does not explain how to pass a callable through MCP, which is a stateless text protocol. This parameter is not usable by LLMs.