A simple, reliable library for getting structured JSON output from LLMs with self-correction.
The server exposes a single tool, 'generate_structured', with a moderately documented description and partial parameter coverage. The tool definition is explicitly visible in the FastMCP decorator (structsure/mcp_server.py, lines 38 - 66), so no inference penalty applies. However, several definition-quality gaps prevent a higher score: the parameter descriptions exist but lack depth (no format/range guidance, no enum constraints); the output schema is not explicitly documented in the tool definition; no error handling guidance is provided; and the tool's responsibility blurs the line between 'generate' and 'validate' (accepting optional schema, defaulting to a minimal model if absent). The naming verb 'generate' is clear and action-oriented, meeting the naming baseline. Overall, the definition quality is fair but falls short of production grade.
Generate a structured JSON response matching the provided (optional) JSON Schema. If no schema is provided, a minimal model with a single `content` field is used.
Tool description lacks actionable detail on when to use this tool vs. alternatives, prerequisite setup, and failure modes. Current: 'Generate a structured JSON response matching the provided (optional) JSON Schema. If no schema is provided, a minimal model with a single `content` field is used.' This is functional but minimal (78 chars). Rubric baseline for description length is 194 chars (p10=34, p90=392). The description should explain: What LLM providers are supported and how to configure them? What happens if JSON Schema parsing fails? What is the expected format of schema_json (valid JSON string)? When should an agent call this vs. calling an LLM directly?
Parameter descriptions are present but lack constraint detail. 'schema_json' is described as 'Optional JSON Schema as a string to validate the output against', but does not specify: valid JSON required? If the schema is malformed, what error is returned? Can nested properties be used? 'max_retries' lacks guidance on typical values (rubric example: 'max_retries 1 - 10' would be actionable). 'provider' and 'model' descriptions omit the decision logic: if provider is None, how is it resolved (environment variable fallback)? If model is None, what defaults apply? These are critical for LLM planning.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No documented output schema. The tool returns 'str' (Pydantic model dumped as JSON string), but the tool definition does not state: What structure will the JSON have? Which fields are guaranteed? What happens if schema_json is provided, are the returned fields strictly those in the schema? If no schema, is the output always {"content": "<string>"}? Without a declared output schema, LLMs cannot infer downstream composition (e.g., whether to extract a field from the response or pass the entire JSON string to the next tool).
No error handling guidance in tool description. The code imports and calls generate() from structsure.core, but the tool definition does not document: What happens if the LLM provider (OpenAI/Ollama) is unreachable? What if max_retries is exhausted and the LLM cannot produce valid JSON? What if schema_json is syntactically valid JSON but semantically invalid as a JSON Schema? An LLM needs recovery guidance: should it retry, ask the user for clarification, or try a different schema?
Parameter 'provider' lacks an enum constraint. Currently described as 'The provider to use: 'openai' or 'ollama'' but not declared as an enum in the schema. This invites the LLM to hallucinate alternative values (e.g., 'anthropic', 'google'). Per the rubric: 'When a parameter accepts one of a known set of values, declare it as an enum.' The FastMCP tool decorator should specify allowed_values or equivalent.
No indication of whether this tool is idempotent or has side effects. If the same prompt and schema are passed twice, will the tool return the same output? Or does it call the LLM twice, potentially incurring costs and producing different results? Per the rubric: 'If the tool modifies state (creates, updates, deletes, sends), the description must say so.' This tool does not modify state, but should explicitly state: 'This tool is read-only and idempotent, multiple calls with the same parameters will produce the same result (modulo LLM randomness if temperature > 0).'