A multi-agent system that demonstrates fiscal self-awareness in autonomous cloud governance. Implements budget-aware agentic computing with circuit breaker mechanisms, cost prediction, and dynamic model routing to prevent cloud cost overruns.
This server demonstrates significant foundational issues across naming, description quality, and schema completeness. While 16 tools are present with basic intent, most lack proper LLM-optimized descriptions, parameter validation constraints, and output schema documentation. The codebase shows Python implementation with S3 and Ollama integration, but tool definitions are scattered across module files with inconsistent documentation patterns. Average tool description length is ~95 chars (baseline is 194 chars), and critical parameter details are missing from most tools. Error handling is minimal, tools return simple success/failure without recovery guidance. The server appears to be an academic project (CSCI-750 Cloud Computing) rather than production-grade tooling.
Analyze a research topic and generate a 3-point technical summary using LLM Brain with Researcher persona
Estimate token count and calculate simulated cost for a text string using formula: tokens ≈ len(text) / 4, cost = (tokens / 1000) * $0.015
Verify that the Ollama server is reachable and responsive by listing available models
Transform raw research notes into a polished executive summary using LLM Brain with Writer persona
Generate a response from the local Ollama LLM with cost simulation metadata. Sends prompt to Ollama, calculates simulated token usage and cost, updates fiscal ledger, and returns structured response with cost metadata.
Get the fiscal summary including tokens used, costs incurred, and cost per 1k tokens
Descriptions lack detail and LLM-optimization. Most descriptions are 40-60 characters; baseline is 194 chars. E.g., 'read_from_s3' description is only 'Read a text file from S3 bucket' (38 chars). LLMs need context on WHEN to use this vs similar tools, WHAT to do if it fails, and WHAT the output structure is.
Output schemas are not documented in any tool definition. Tools like 'get_ledger' and 'get_fiscal_summary' return structured data (financial state), but the response structure is never specified. LLMs cannot plan downstream calls or extract fields without documented schemas.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 42 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Return the current financial state including global budget, daily limit, current spend, total savings, and per-agent spend breakdown
Dynamic Model Routing rule: If spend exceeds 80% of daily limit, advise 'FRUGAL' local model to preserve funds, otherwise return 'PREMIUM'
Get the total simulated cost for the current session
Track dollars saved when routing actively avoids cloud costs by using local models
Execute the full writing workflow: read raw research notes from S3, transform into executive summary via LLM Brain, save polished report to S3, and return the executive summary
Read a text file from S3 bucket
The preemptive circuit breaker. Evaluates if an agent's task can proceed by checking if estimated cost would exceed daily limit. Returns True if approved, raises BudgetExceededException if denied.
Execute the full research workflow: read research topic from S3, analyze with LLM Brain, save research notes to S3, and return the summary
Feedback loop: Adjust the prediction multiplier for an agent based on the difference between predicted and actual cost. Increases factor if under-predicted, decreases if over-predicted.
Write text content to S3 bucket
Parameter descriptions are missing or generic. Many tools have parameters with no description text (e.g., 'bucket' in read_from_s3 has description 'The S3 bucket name' but no mention of format, constraints, or error cases). Parameter descriptions should be 50-100 chars on average, current average is ~35 chars.
No error handling or recovery guidance. Tools like 'request_funds' raise 'BudgetExceededException' but there is no documentation of what the exception contains, what the LLM should do next (retry? escalate?), or how to check remaining budget before calling. Error responses must tell LLMs what to do next.
Composite tools violate single-responsibility principle. 'research_and_summarize' reads from S3, calls LLM, writes to S3, and returns summary, four concerns in one. 'polish_and_publish' does the same. Split these into read_from_s3 + analyze_topic + write_to_s3 + summarize so agents can compose as needed.
Missing enum constraints on constrained parameters. 'model_override' in generate_response is a free-form string with no enum of supported models. 'tier' concept (FRUGAL vs PREMIUM) in get_recommended_tier is mentioned in description but not formalized as an enum. LLMs will hallucinate invalid values.
Tool names contain domain jargon without verb clarification. 'log_savings', 'request_funds', 'update_bias_factor' are clear, but 'get_recommended_tier' could be confused with 'get_recommended_model'. Additionally, several financial tools (ledger, savings, cost) operate on the same resource but naming distinctions are unclear (get_ledger vs get_fiscal_summary vs get_session_cost, what is the difference?).
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Tools are marked with Risk metadata (READ_ONLY, WRITE) in the server index but these are not exposed in the MCP protocol schema as tool annotations. LLMs cannot automatically detect safe-to-retry vs destructive operations.
No pagination or result limiting on list/summary tools. 'get_ledger' returns agent_spends as a full dict; 'get_fiscal_summary' returns unspecified structure. If there are thousands of agents, these could bloat context. No mention of limits, pagination, or truncation strategy.
Identifiers and references are not self-documenting. 'agent_name' parameter in request_funds and update_bias_factor assumes agents have a consistent naming scheme, but no description clarifies the format (string slug? UUID? email?). No guidance on how to discover valid agent names.