MCP server that integrates SchemaCrawler with AI services, providing tools to analyze database schemas and generate insights
SchemaCrawler AI exhibits significant structural gaps in schema documentation and parameter descriptions across all 12 tools. While tool names are reasonably descriptive (verb-noun pattern: describe_*, list_*, lint, diagram, etc.), the submission provides only high-level tool descriptions without visible input parameter schemas, output schemas, or parameter-level descriptions. The source code provided shows only Dockerfile, pom.xml build artifacts, and dependency declarations, no actual Java tool implementation files that would contain FunctionDefinition classes with schema details. Based on the file paths cited (e.g., DescribeErRelationshipsFunctionDefinition.java), these implementations likely exist in the repo, but they are not visible in the submission. Per the hard scoring rule 'If you cannot see the actual tool definition in the source (only inferred): cap that tool's overall at 50,' all tools are capped at 50 due to missing visibility. The submission provides tool names and risk labels but lacks concrete schema evidence. Description lengths (averaging 80-150 chars in the visible list) are adequate but lack the actionable specificity required (e.g., no parameter constraints, dependency hints, or error guidance). No evidence of output schema documentation, pagination support, error handling patterns, or security scoping.
Provides database environment and server configuration metadata, including engine type and version, collation, encoding, parameters, capabilities, and platform details. Adapts output to the specific database (such as Oracle, SQL Server, PostgreSQL and so on) to support platform-aware SQL generation and schema analysis.
Describes entity-relationship model relationships and constraints in the database schema
Describes stored procedures, functions, and other database routines
Provides detailed information about database tables including columns, constraints, and metadata
Detects and identifies clusters or groups of related tables in the database schema
Generates visual diagrams of the database schema
Performs schema quality checks and linting to identify potential issues and best practice violations
No visible input schemas for any tool. The submission lists tool names and provides brief descriptions, but does not include the actual FunctionDefinition source code showing parameter types, constraints, or JSON Schema definitions. The about_database tool shows an empty input object '{}', but the other 11 tools lack any schema visibility.
No documented output schemas. Tool descriptions do not specify what fields are returned, what data types they are, or how results are structured. LLMs cannot plan chaining (e.g., using a returned table_id in table_sample) without explicit output documentation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 35 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 14 | - | v1 |
Lists database objects such as tables, views, and schemas
Lists columns and members of specific tables
Analyzes and rates the importance of tables in the database schema
Finds and describes relationship paths between tables in the database schema
Provides sample data from database tables
No parameter-level descriptions. The submission does not show what parameters each tool accepts, their types, valid ranges, enums, or constraints. Even tools with obvious parameters (e.g., list likely accepts schema_name, filter, limit) have no documented contract.
Tool descriptions lack actionable specificity. Descriptions are between 60 - 150 chars and state WHAT the tool does (e.g., 'Describes entity-relationship model relationships') but do not explain WHEN to use it, dependencies (e.g., 'Call list() first to discover table names'), or expected output structure.
Generic tool names reduce clarity. 'list' and 'lint' are vague, 'list' could mean list tables, views, schemas, or all objects; 'lint' does not hint at what it checks (foreign keys? naming conventions? performance issues?). Per naming patterns, tools should use descriptive verb_noun forms like 'list_tables', 'check_schema_compliance', or 'lint_constraints'.
No error handling guidance. The submission provides no information on what errors each tool can raise, how they are classified (retryable, user-fixable, fatal), or what recovery steps an LLM should take. This violates recovery-guide and error-classification patterns.
No tool composition chain documented. With 12 discovery and analysis tools, the submission does not explain which tools should be called first (list → describe_tables → table_sample?) or how outputs chain together. This forces LLMs to guess optimal call sequences.
No pagination or result limiting documented. Tools like 'list' and 'describe_tables' on large schemas could return hundreds or thousands of items. The submission does not specify if pagination is supported, what the default/max result limits are, or how to fetch subsequent pages.
No security scoping. The submission does not declare what permissions these tools require (read:schema? read:data? read:metadata?). Since they operate on live database connections, agents should understand the access scope and potential blast radius.