A web-based data lakehouse viewer and analyzer for Apache Iceberg catalogs with insights, scheduling, and data exploration capabilities
Lakevision is a Python/FastAPI MCP server for Iceberg catalog introspection with 17 read-only tools. The server has CONSISTENT and COMPLETE tool definitions across all tools, every tool has a name, description, and JSON Schema input parameters. However, several DEFINITION QUALITY issues prevent a higher score: (1) Parameter descriptions lack detail and constraint information (e.g., 'Iceberg Table object' is vague and unmappable to LLM context); (2) No output schemas are documented, forcing LLMs to guess what fields are returned; (3) Tool descriptions, while present, are generic and do not clearly signal WHEN to use one tool vs. another (e.g., get_tables vs get_all_table_names distinction is unclear); (4) Most parameters that accept 'Iceberg Table object' are not user-discoverable, there is no tool to CREATE or LOAD a table first, breaking the tool chain; (5) No error handling guidance in descriptions; (6) The 'rule_*' tools (tools 10-17) lack composition context, it is unclear if they return boolean pass/fail, scored findings, or structured diagnostics. The server is functionally complete for READ operations on Iceberg catalogs, but lacks the LLM-optimization patterns (constrained inputs, output schemas, recovery guidance, chaining IDs) that production-grade agent tools require.
Retrieves all tables across multiple namespaces
Retrieves data change history for a table with flattened summary statistics
Returns a paginated list of the table's data and delete files with metadata, newest spec first
Retrieves a list of namespaces from the Iceberg catalog, optionally including nested namespaces
Retrieves partition information for a table as a DataFrame
Executes a SQL query against table data and returns results with optional optimization analysis
NO OUTPUT SCHEMAS DOCUMENTED. All 17 tools lack documented return types, forcing LLMs to guess what fields are in responses.
VAGUE PARAMETER DESCRIPTIONS. Many parameters are described generically (e.g., 'Iceberg Table object') without explaining format, source, or how to obtain them. This breaks tool composition and forces LLMs to reason about unmappable types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | - | v1 |
Retrieves snapshot history for a table sorted by commit timestamp in descending order
Lists all tables in a specified namespace
Loads a table from the Iceberg catalog by its identifier
Insight rule that detects tables containing UUID columns
Insight rule that detects tables containing large files (>=1GB by default)
Insight rule that detects tables without a defined location property
Insight rule that detects empty tables with no rows
Insight rule that detects partitions with uneven data distribution (skewed partitions)
Insight rule that detects tables with many small files (>100 files averaging <100KB each)
Insight rule that detects large tables (>50GB) with small average file sizes
Insight rule that detects tables with excessive snapshot history (>500 snapshots by default)
UNCLEAR TOOL COMPOSITION. No tool to create/obtain an Iceberg Table object, yet 14 tools require it as input. The tool chain is broken, LLM cannot discover how to populate this parameter.
RULE_* TOOLS LACK OUTPUT STRUCTURE DOCUMENTATION. All 8 diagnostic tools (rule_small_files, rule_no_location, etc.) do not document whether they return boolean, list, score, or structured findings. This makes them unusable for composition.
OVERLAPPING TOOL NAMES AND PURPOSES. get_tables, get_all_table_names, and rule_* tools have overlapping or ambiguous purposes. Tool descriptions do not clearly explain WHEN to call one vs. another, forcing LLMs to guess.
VAGUE PARAMETER FORMAT GUIDANCE. Parameters like 'table_id' lack format documentation (e.g., 'expected format: namespace.table_name'). LLMs cannot reliably construct valid identifiers without explicit constraints.
NO ERROR HANDLING GUIDANCE. Tool descriptions do not explain what happens on error (e.g., if a namespace does not exist, if SQL is invalid, if a table is empty). LLM has no recovery strategy.
COMPOUND TOOL NAME (rule_skewed_or_largest_partitions_table). The 'and/or' in the name signals multiple concerns. This may violate single-responsibility principle, consider splitting into two tools or clarifying which condition is checked.