TypeScript SDK for building multi-provider AI agents that combine LLM reasoning with MCP tools. Supports OpenAI, Anthropic, Mistral, Bedrock, and Vertex AI with automatic tool discovery, connection pooling, and OAuth authentication.
This is a framework/SDK for building MCP servers, not an MCP server itself. The evaluation focuses on the 25 example tools across multiple example servers (filesystem, tasks, weather, astro, favorites, gmail, and test servers). Most tools have minimal descriptions (10-40 chars), lack comprehensive input validation guidance, and exhibit inconsistent schema quality. Many parameters lack type specificity (e.g., 'status' in mark_item has no enum constraint). Error handling is largely absent from tool definitions. Tool naming is generally acceptable but some tools show redundancy (two 'get_weather' tools in different servers with slightly different schemas). Output schemas are not documented in the definitions provided. The framework demonstrates basic MCP support but the example tools fail to meet production-grade quality standards for LLM-driven agents.
Archive an item
Mark a task as completed
Create a new item
Create a new Gmail label or get existing label ID
Create a new task
Fetch user profile
Get full email body content
Duplicate tool names across servers: two 'get_weather' tools with different input schemas (one with 'unit' enum, one without). LLMs will conflate these during selection.
Most descriptions under 60 characters; 12 tools have descriptions <45 chars. Insufficient context for LLM tool selection.
Parameter 'status' in mark_item and 'action' in process_email lack enum constraints and descriptions. Free-form strings invite LLM hallucination of invalid values. No guidance on valid values provided.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 57 | - | v1 |
Get 3-day forecast
Get item details by ID
Return favorite food and drink given an astrological sign
Return astrological sign for a birthdate (YYYY-MM-DD)
Get current weather for a city
Get current weather for a city
List all unread emails from inbox (up to specified limit, good for bulk operations)
List files in a directory
List all tasks
List unread emails from inbox. Returns up to 100 emails per call. Use pageToken to get more.
Mark an email as spam and move to spam folder (keeps email as unread). Optionally add custom label.
Mark an item (simulates batch operation)
Process an email by ID
Process an item by ID (takes 200ms)
Read contents of a file
Search for text in files
Send a notification
Write content to a file
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract required fields (e.g., does create_task return task_id?).
Error handling absent from tool definitions. No recovery guidance, retryability hints, or actionable error messages. Agents cannot self-correct on failures.
Pagination not documented for list tools. list_directory, list_tasks, list_unread_emails, list_all_unread_emails lack guidance on result limits and cursor/offset mechanics. Large results will blow context windows.
Generic/vague parameter names: 'action' in process_email, 'status' in mark_item, 'pattern' in list_directory (no explanation of pattern syntax).
Destructive operations (write_file, complete_task, mark_as_spam, archive_item) lack confirmation/dry-run support.
No tool annotations present (readOnlyHint, destructiveHint, idempotentHint per spec 2026-07-28). LLMs cannot infer operation safety or retry semantics.
Gmail tools accept optional parameters with unclear defaults: maxResults defaults to 50 but can go to 100; maxTotal defaults to 30 but can go to 50. Defaults should be stated explicitly in descriptions.