Thin protocol shim that exposes Misata's public API to AI agents over the Model Context Protocol. Generates realistic, referentially-intact multi-table synthetic data with no real data, no ML model, fully seeded. Provides tools for schema design, data generation, validation, and database seeding.
Misata MCP server has clear, domain-specific tool names and generally good descriptions that explain WHAT the tools do and WHY to use them. However, input schemas are not fully visible in the provided source code, only parameter names and brief type hints are shown. The source excerpt cuts off mid-description for tool 2, and complete schema definitions are not included. Based on what is visible: tool names follow verb_noun conventions (generate_from_schema, list_domains, validate_domain); descriptions range 60 - 200 characters and include context ('Look before you generate', 'Prove the result holds'); parameters are typed (dict, string, int, boolean) but lack detailed range, format, or enum constraints in descriptions. Error handling is not visible in the code excerpt. Output schemas are undocumented. The server exposes legitimate risks (WRITE, DESTRUCTIVE) but does not show confirmation or dry-run patterns for destructive operations. Strengths: clear intent, good naming, adequate descriptions. Weaknesses: incomplete schema visibility, missing output documentation, no visible error recovery guidance, no permission gates or audit logging shown.
Prove the result holds. Coherence: reader-visible contradictions, scored. Runs coherence audit on any folder of CSVs.
From one sentence; Misata's parser designs the schema. Quick one-sentence requests where Misata's own parser should design the schema (18 curated domains + structural composition for unknown ones).
Design a schema, get data. Primary tool: the agent designs the schema, Misata guarantees the math. Returns a per-relationship integrity verification.
Look before you generate. Inspect a schema's structure, column definitions, relationships, and constraints without generating data.
Look before you generate. Return the public list of domains with canonical sample stories and keywords.
Look before you generate. Show the agent Misata's interpretation of a one-sentence story before generating, so it can revise the schema before generation runs.
Input schemas incompletely documented. Provided source shows parameter type hints (dict, string, int, boolean) but no JSON Schema definitions, no format/pattern constraints, no min/max for numeric ranges, no enum declarations for constrained values (e.g., domain names in validate_domain). Parameter descriptions lack actionable format guidance.
Output schemas not documented. No visible tool descriptions explain what fields are returned, what structure the response has, or what the LLM should expect. This forces agents to blindly call tools without knowing what to extract or chain next.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Fill a real database. Plans by default; writes only on apply=true. Reads the schema from the database itself, inserts parents before children, and verifies every foreign key against the database afterwards.
Prove the result holds. Domain validation: flags values that are physiologically or financially impossible for a stated domain.
Prove the result holds. Schema validation: will this schema generate at all?
Destructive operation (seed_database with apply=true, truncate=true) lacks confirmation/dry-run pattern. Tool description notes 'Plans by default; writes only on apply=true' but no visible confirm-before-execute guard or recovery path if truncation succeeds but insertion fails.
No visible error handling guidance in tool descriptions. If generate_from_schema fails (invalid schema, out-of-memory on large row count), the LLM has no guidance on what to do next: retry? Call inspect_schema? Reduce rows? Users see vague errors.
seed_database accepts a database connection_string as a parameter. This is a credential-bearing string (postgresql://user:pass@host/db). Credentials should never be passed as tool parameters, they end up in agent traces and logs. Use server-side environment injection instead.
No visible permission gates or audit logging. seed_database is marked DESTRUCTIVE but there is no evidence of permission checks, user/agent logging, or audit trail showing who called the tool and what happened. For a tool that can truncate production tables, this is a compliance gap.
Tool annotation hints missing. Tools marked WRITE, DESTRUCTIVE, READ_ONLY but no visible readOnlyHint, destructiveHint, or idempotentHint fields in tool registration (as per MCP 2026-07-28 spec). This prevents clients from rendering appropriate UI warnings or enforcing safety policies.