Sample repository demonstrating advanced Strands Agents integration with Model Context Protocol (MCP) servers, including multi-agent patterns, memory management, custom hooks, and MCP tool integration
This MCP server has severely deficient tool definitions. Descriptions are present for all tools but are extremely brief (5-50 characters), well below the 10-1024 character best practice and too short to guide LLM tool selection. Most descriptions are single clauses (e.g. 'Performs arithmetic calculations', 'Read the contents of a file') that fail to explain WHEN to use the tool, what prerequisites exist, or what the output structure is. No tools show error handling guidance, no output schemas are documented, and parameters beyond mem0_memory are completely invisible in the source code provided. The server appears to be built on strands-agents framework which may have auto-wiring of tools, but if tool registration is inferred rather than explicitly visible in code, per the hard scoring rule, those tools are capped at 50 individually. The mem0_memory tool (only one with visible schema) scores higher but still has minimal description.
Performs arithmetic calculations
Create and edit Python tool files under cwd()/tools/*.py
Read the contents of a file
Write content to a file
Generate detailed cost analysis reports for AWS services
Get architecture patterns for AWS Bedrock
Get actual pricing data using service code, region, and filters
Get valid values for AWS pricing attributes
15 of 16 tools have NO visible input schemas in source code. Schema visibility is 0 for all except mem0_memory and websearch.
All tool descriptions are extremely brief (5-50 characters), far below the 10-1024 best practice range. Descriptions lack WHEN to use the tool, prerequisites, return type structure, or error cases. Examples: 'Performs arithmetic calculations' (32 chars), 'Read the contents of a file' (28 chars). These cannot adequately guide LLM tool selection.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Find filterable attributes for an AWS service
List all available AWS service codes
Make HTTP requests to external services
Store and retrieve persistent memories for users
Get full content of specific AWS documentation
Get related AWS documentation
Find relevant AWS documentation
Search the web for updated information using DuckDuckGo
No output schemas documented for any tool. LLMs cannot plan downstream tool calls or extract required data without knowing response structure. E.g., does file_read return {content: string, size: int, mtime: string} or just a raw string?
Tool definitions appear to be inferred from framework auto-wiring (strands-agents) rather than explicitly visible in provided source. Tool registration code not shown. This applies to 14 of 16 tools.
No error handling guidance documented. LLMs have no recovery path if a tool fails. E.g., what should the agent do if file_read fails on a non-existent file? If get_pricing fails due to invalid service code?
Some tool names are generic or ambiguous. 'recommend' lacks action verb and context, recommend what? 'editor' is vague, edit what? Does it edit Python files, AWS configs, or arbitrary text?
websearch and http_request accept string parameters without format constraints. websearch's 'region' param should be an enum (us-en, uk-en, ru-ru) not free-form. http_request's URL param needs a URI format constraint.
file_write tool has no description of what happens on existing files (overwrite vs append vs error). mem0_memory 'action' param description says 'store, retrieve, or list' but lacks clarity on what 'store' returns, when to use 'list' vs 'retrieve', or auth model.