A template repository for building generative AI applications with LangGraph, LangChain, and marimo notebooks, featuring agent workflows, embeddings, vector databases, and multi-component orchestration
This server defines 2 tools with basic descriptions and input schemas, but falls significantly short of production quality. Both tools have descriptions that meet minimum length, but lack critical guidance on when to use them, what they return, error handling, and dependencies. The tools are simple read-only operations, but the schemas are minimal and parameter descriptions are too generic. The implementation is a LangGraph notebook example rather than a production MCP server, and the tools are defined inline in marimo/agent_example.py with no explicit MCP registration visible. No error handling, no output schema documentation, no parameter constraints, and no tool annotations. This lands in the poor category (D grade).
Add two numbers together.
Get the current weather for a specified city.
No output schema documentation. get_weather and calculate_sum do not document return types or structure. LLMs cannot plan downstream calls or extract data reliably.
Tool definitions are inferred rather than explicitly registered. Tools are defined as @tool decorators in a marimo notebook with no visible MCP server registration, schema declaration, or capability negotiation. Capping per-tool scores at 50 per HARD SCORING RULES.
Parameter descriptions are trivial ('The name of the city', 'First integer'). Descriptions lack actionable guidance: expected format, constraints, or when to use alternatives. Descriptions under 30 chars cannot exceed score 40; these are at that boundary.
No error handling or recovery guidance. Tools have no documented error responses, failure modes, or what the agent should do if a city name is invalid or math fails.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 15 | - | v1 |
No parameter constraints or validation hints. 'city' parameter lacks format guidance (e.g., accepted length, case sensitivity). 'a' and 'b' parameters have no numeric range documentation (int range, positive/negative allowed).
Tool descriptions do not indicate what data is returned or how to use results downstream. get_weather returns a string description; calculate_sum returns an integer. No hint on structure or downstream tool compatibility.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Both tools are read-only, but this is not explicitly declared. Lack of annotations prevents the agent from reasoning about safety and retry behavior.
No documentation of what each tool is for or when to prefer one over another. Tool descriptions answer 'what it does' minimally but not 'when to use it' or 'what makes it different from similar tools'. For get_weather, there is no hint on whether it works for all cities, what timezone it returns, or how current the data is.