Batteries-included agent harness for Python — tool-calling, sandboxed execution, multi-agent teams, and unlimited context on Pydantic AI
This MCP server presents significant definition quality gaps. While 20 tools are registered with input schemas and descriptions, the quality is inconsistent. Tool names are generally well-formed with action verbs (get_, list_, navigate, click, type_text, etc.), but descriptions vary widely in quality and completeness. Many parameter descriptions are present but lack detail about constraints, valid ranges, or format expectations. Output schemas are not documented in the provided source material. Error handling guidance is minimal or absent across most tools. The browser toolset (navigate, click, type_text, screenshot, scroll, execute_js) lacks parameter detail for common constraints (e.g., max URL length, CSS selector specificity, JavaScript security boundaries). The GitHub tools (github_list_repos, github_list_issues, etc.) describe parameters but omit pagination guidance despite being list-returning tools. Checkpointing tools (save_checkpoint, list_checkpoints, rewind_to) have minimal descriptions of checkpoint metadata structure and restoration semantics. Code analysis tool (analyze_code_complexity) lacks output schema documentation. Overall, this reads as a functional but unpolished implementation rather than a production-ready agent framework.
Analyze the complexity of a Python file.
Click an element on the current page. Args: selector: CSS selector (e.g. 'button#submit', 'a.nav-link') or pixel coordinates as 'x,y' (e.g. '450,300'). Returns updated page content after the click.
Execute a JavaScript expression in the browser and return the result. Args: script: JavaScript expression to evaluate. For example: 'document.title' or 'Array.from(document.querySelectorAll("h1")).map(e => e.innerText)'. Strings are returned as-is, objects and arrays are returned as JSON, and a null/undefined result is returned as the literal 'undefined'. Returns an error message if evaluation failed.
Get the current date and time.
Get the text content of the current page or a specific element. Args: selector: CSS selector to extract text from. Omit for full page text. Returns plain text content (Markdown for full page, raw text for element).
Get detailed statistics for a repository.
List tools lack pagination documentation. github_list_repos, github_list_issues, github_list_pull_requests have no page, offset, limit, or cursor parameters visible in input schemas; descriptions do not mention pagination strategy or result limits.
Output schemas are not documented in source material. Tools return complex structures (e.g., navigate returns page title, URL, and markdown content; github_get_repo_stats returns detailed statistics; checkpoint tools return metadata) but no formal output schema is visible in tool registration or descriptions.
Browser tools lack security boundary documentation. execute_js accepts arbitrary JavaScript with no mention of sandbox isolation, CSP, or what globals/APIs are available. click and type_text have no validation guidance for selector injection or input sanitization.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Get information about a GitHub user.
List issues in a repository.
List pull requests in a repository.
List GitHub repositories.
Navigate back in the browser history. Returns page content of the previous page.
Navigate forward in the browser history. Returns page content of the next page.
List all saved checkpoints with their labels and metadata. Use this to see available restore points before deciding to rewind.
Log a message to /logs/agent.log.
Navigate the browser to a URL and return the page content as Markdown. Args: url: Full URL to navigate to (e.g. https://example.com). Returns page title, current URL, and rendered page content as Markdown. Content is truncated to max_content_tokens if the page is very large.
Rewind the conversation to a previously saved checkpoint. This restores the conversation state to the checkpoint and discards all messages after it. Use this when the current approach isn't working and you want to try a different strategy from a known good state.
Save a named checkpoint of the current conversation state. Labels the most recent auto-checkpoint with the given name. Use this before risky operations (major refactors, destructive changes) so you can rewind later if things go wrong.
Take a screenshot of the current page. Args: full_page: If True, capture the full scrollable page (default False). Returns base64-encoded PNG image prefixed with the data URI scheme.
Scroll the page in a given direction. Args: direction: 'up', 'down', 'left', or 'right'. x: Optional x coordinate for localized scroll (ignored when None). y: Optional y coordinate for localized scroll (ignored when None). Returns updated page content after scrolling.
Type text into an input field on the current page. Args: selector: CSS selector for the target input element. text: Text to type. Replaces any existing value in the field. Returns updated page content after typing.
Missing error handling and recovery guidance. Descriptions do not explain what happens on invalid input (e.g., malformed URL in navigate, nonexistent GitHub repo in github_get_repo_stats, invalid CSS selector in click). No mention of retryable vs fatal errors or suggested corrective actions.
Checkpoint tool semantics undefined. save_checkpoint, list_checkpoints, and rewind_to lack formal documentation of checkpoint metadata (timestamps, conversation state format, size limits), retention policy, whether rewind discards intermediate messages, and what conversation state is actually preserved.
Parameter constraints missing or incomplete. navigate URL parameter has no max length, format, or protocol validation hint. scroll direction parameter has description but no enum constraint. github_list_repos sort_by has description but no explicit enum of valid values. github_list_issues state parameter description mentions 'open, closed, or all' but may not be a formal enum.
analyze_code_complexity lacks output schema and result format documentation. Description mentions 'analyze the complexity' but does not explain what metrics are returned (cyclomatic complexity? lines of code? function count?) or how results are structured for agent consumption.
go_back and go_forward have minimal descriptions. Both are under 50 characters ('Navigate back in the browser history.' / 'Navigate forward in the browser history.') and do not explain expected output, error conditions (e.g., no history to navigate), or when these tools are preferable to explicit navigation via navigate().