Production-grade AWS MCP server with safety controls and cross-service intelligence
AWS Sage presents 17 tools with moderate definition quality. All tools have explicit names and descriptions visible in server.py, and input models are defined via Pydantic. However, critical gaps emerge: (1) most parameter descriptions are minimal or generic, e.g., 'AWS service (e.g., 's3', 'ec2')' provides no constraint or validation hint; (2) no output schemas are documented in the visible code, tools return results but the structure is not specified; (3) descriptions lack depth about when/why to use each tool, dependencies, or recovery guidance; (4) enums are not used where they should be (safety modes, incident types, environment names), leaving LLMs to hallucinate values; (5) error handling and validation rules are not visible; (6) critical tools like 'aws_execute' and 'assume_role' are high-risk but lack confirmation or guard descriptions; (7) tool naming is generally clear but some names are ambiguous (e.g., 'aws_query' vs 'aws_execute' distinction not obvious from names alone). Average per-tool score: ~48.
Analyze AWS costs and find optimization opportunities
Assume an IAM role for cross-account access
Execute an AWS operation
Execute a natural language AWS query
Compare configurations across AWS environments
Discover AWS resources across services by tags
Get current AWS account and session information
No output schemas documented in code. Tools return results but structure is not specified for LLM planning. E.g., list_profiles returns a list, but fields/structure unknown.
Parameter descriptions are minimal and lack validation/constraint hints. E.g., 'AWS service (e.g., 's3', 'ec2')' does not state valid services, format, or constraints. Should specify: enum values, format patterns, length limits, ranges.
No enum constraints for enumerated parameters. 'set_safety_mode' accepts 'read_only|standard|unrestricted' but parameter is a free-form string. 'investigate_incident' accepts 'incident_type' but no enum. LLMs will hallucinate invalid values.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Analyze the impact of deleting a resource
Investigate AWS incidents with guided troubleshooting
List all available AWS profiles
Map dependencies of an AWS resource
Query AWS knowledge base for best practices
Search AWS documentation
Select an AWS profile to use for operations
Set a resource alias for convenience
Change the safety mode for operations
Switch between production and LocalStack environments
Destructive operations (aws_execute, assume_role) lack explicit confirmation or guard descriptions. No error guidance for recovery. No mention of reversibility or side effects.
Tool descriptions lack WHEN/WHY guidance and dependency hints. E.g., 'map_dependencies' and 'impact_analysis' do not explain when to call them or relationship to each other. No hints like 'Call this before deleting a resource'.
No error handling or recovery guidance visible in tool definitions. No mention of what errors are retryable, what requires user action, or next steps on failure.
Ambiguous tool naming: 'aws_query' vs 'aws_execute' distinction unclear from names alone. Both operate on AWS services but difference not obvious without reading descriptions.
Parameters accept free-form strings where IDs/references are expected, but no guidance on resolution. E.g., 'assume_role' takes 'role_arn' but if user says 'assume the prod role', no discovery tool mentioned.