A clean, powerful agent runtime with tools, MCP support, profiles, and streaming events.
This MCP server has severe definition quality issues across nearly all tools. The fundamental problem: 8 out of 16 tools (50%) have NO visible input schemas in the source code, only names and descriptions. Additionally, most tool descriptions are generic and lack actionable guidance for LLM tool selection. Tools like 'browser_use' (WRITE risk) lack error handling guidance, recovery steps, or confirmation patterns. Parameter naming is inconsistent (some tools expose internal concerns like 'compounds_per_year' without context). The examples show reasonable naming for simple math tools (haversine_distance, compound_interest, add_numbers) but these are outliers. Most critically, 'workspace_files', 'sub_agent', and other tools in shipit_agent/deep/deep_agent/toolset.py appear to be stubs or placeholders with no schema information visible, suggesting incomplete implementation.
Add two integers and return the result.
Drive a real browser to accomplish a goal that requires reading and interacting with web pages: clicking, typing, scrolling, navigation.
Calculate compound interest on a principal over a time period.
Create and evaluate decision matrices.
Synthesize evidence from multiple sources.
Calculate great-circle distance in km between two GPS coordinates.
Plan tasks and break down goals into steps.
50% of tools (8/16) have NO input schema, only empty {} definitions.
Generic, uninformative descriptions. 'Decompose complex thoughts into simpler components' (thought_decomposition), 'Interact with workspace files' (workspace_files), 'Delegate tasks to a sub-agent' (sub_agent), these lack guidance on WHEN to use, what inputs to provide, or what output to expect. LLMs cannot reliably select these tools.
No error handling or recovery guidance for destructive tool 'browser_use' (WRITE risk). No confirmation pattern, no dry-run option, no guidance on what errors might occur or how to retry. Agents calling this tool with incorrect goals could trigger unintended web interactions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 40 | 2026-07-28+ | v2 |
Return the local project context and execution style for this agent.
Delegate tasks to a sub-agent.
Decompose complex thoughts into simpler components.
Semantic tool discovery - find relevant tools by natural language query.
Verify and validate outputs.
Get weather information for a city.
Return local workspace conventions and delivery rules.
Interact with workspace files.
Return a short summary of how a workspace agent should operate.
Two WRITE-risk tools ('browser_use', 'workspace_files') lack destructiveHint or idempotentHint annotations. LLMs cannot determine retry safety or whether side effects are expected.
No documented output schemas. Tools return data but LLMs have no contract specifying which fields to expect, making downstream tool chaining error-prone. Example: does 'weather' return temperature as {temp: number} or {temperature: {value: number, unit: string}}?
Parameter descriptions are missing or minimal across most tools. 'city' in weather has a description, but many tools lack descriptions entirely (project_context, planner, sub_agent all show empty parameters in Input). LLMs cannot infer meaning from parameter names alone.
Naming inconsistency: 'browser_use' uses underscore, but function names in examples use camelCase (haversine_distance, compound_interest). No consistent verb-noun pattern. Some tools are nouns ('planner', 'verifier') rather than action verbs (should be 'plan_tasks', 'verify_output').
'sub_agent' and 'workspace_files' are vague names. Do they modify files or read them? Does sub_agent delegate or supervise? Ambiguous names force LLMs to hallucinate behavior.