A FastAPI-based chat backend with multi-agent support, tool execution, document collaboration, and image generation capabilities
This MCP server exposes 15 tools spanning project sources, memory, images, collaborative documents, and code execution. While tool names follow verb-noun conventions (list_, search_, read_, store_, remove_, generate_, edit_, create_, update_, activate_, delete_, vary_), the implementation has critical gaps in schema documentation, parameter descriptions, and output schema visibility. Most tools lack visible input schema definitions in the provided source code, only parameter names and basic types are listed in the evaluation summary, but the actual JSON Schema implementations are not shown. Output schemas are entirely undocumented. Error handling and recovery guidance are absent. No tool uses risk annotations (readOnlyHint, destructiveHint, idempotentHint) even though the risk profile (READ_ONLY, WRITE, DESTRUCTIVE, IRREVERSIBLE) is known. The code_execution tool poses significant security and error-handling risks without clear sandboxing documentation or recovery patterns. For a production-grade agent toolkit serving 15 tools, this represents significant gaps against the 54 Agentic Tool Patterns baseline.
Activate a specific collaborative document in the current chat
Execute code in a sandboxed environment (Python, JavaScript, etc.)
Create a new collaborative document for the current chat
Delete a collaborative document from the current chat
Edit an existing image by removing or replacing parts with generated content
Generate an image from a text prompt using configured image generation models
List all collaborative documents in the current chat
No visible input schema definitions in source code. Evaluation summary lists parameter names/types but actual JSON Schema implementations (with type constraints, min/max bounds, enums, patterns) are not shown in the provided files. Cannot verify that schemas match the rubric baseline of properly typed, constrained parameters.
No output schema documentation visible. Tools return data but the structure (fields, types, pagination, next_cursor) is not documented in the provided source. LLMs cannot predict what fields are available for downstream tool calls or result chaining.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | 2026-07-28+ | v2 |
List all visible sources (uploaded files, URLs, text sources, and indexed chats) for an agent project
Read content from a specific project source with pagination support
Remove a stored memory entry by ID
Search through past chat conversations by title and content
Search through project sources using semantic search to find relevant passages
Store a memory entry for the current user
Update the content of an active collaborative document
Create variations of an existing image
Destructive and irreversible tools lack confirmation/dry-run patterns. remove_memory, delete_cowork_document, and especially code_execution (IRREVERSIBLE) can cause data loss or unintended side effects. No confirmation request (pattern:confirmation-request) or dry-run mode visible.
code_execution tool lacks security and error documentation. No visible sandboxing constraints (e.g. 'supports Python, JavaScript only'), timeout limits, resource caps, or recovery guidance. Agents cannot safely reason about execution boundaries. IRREVERSIBLE risk tag suggests unrecoverable failure modes but no error handling or compensation patterns are documented.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Risk profile is known (READ_ONLY, WRITE, DESTRUCTIVE, IRREVERSIBLE) but not encoded as MCP tool annotations. Agents cannot infer which tools are safe to retry or whether side effects are present.
Parameter descriptions are minimal. remove_memory has only 'UUID of the memory to remove (required)' (26 chars), under the 50-200 char LLM-optimized range. Many parameters lack context: does search_project_sources 'limit' have min/max bounds? Does generate_image 'prompt' accept free text or must it follow a schema? Do pagination parameters exist?
No recovery or error guidance. Tools do not document retryability, user-fixable errors, or actionable remediation. If code_execution fails, what should the agent do? If a source is not found, should it offer alternatives? No pattern:recovery-guide implementation visible.
Pagination not explicitly documented. Tools like list_project_sources, search_project_sources, and search_past_chats accept 'limit' parameters but no offset/page/next_cursor, total_count, or has_more fields are mentioned. Agents cannot safely iterate over large result sets.
read_project_source accepts both 'source_id' (numeric or UUID string) and manual offset/max_chars pagination. Mixing reference types (integer vs string UUID) without explicit validation or format description invites type confusion. Described as 'source ID (numeric from list or UUID string)' but no constraint guarantees LLM will use consistent types across calls.
create_cowork_document and update_cowork_document manage versioning and format but no enums or validation rules are documented for 'format' (markdown, code, text, json, csv, presentation) or 'language' parameters. LLM may pass invalid values.