Provides tools for reading, querying, and modifying parquet files in a repository data directory
This server provides 16 well-intentioned tools for parquet file operations with reasonable naming conventions and mostly complete parameter schemas. However, significant gaps exist: output schemas are not documented anywhere in the source code; error handling is minimal with no recovery guidance; descriptions lack LLM optimization and are sometimes generic; several tools expose low-level implementation details that should be abstracted. The server follows basic tool structure but falls short of production-grade quality. Average per-tool score across 16 tools is 62, placing this in the 'Fair' range with noticeable gaps.
Perform aggregation operations (count, sum, avg, min, max, stddev) on parquet file columns, optionally grouped by one or more columns.
Create a full snapshot backup of a parquet file
Delete rows from a parquet file that match the given filters. Supports the same filter syntax as query_parquet.
Export filtered parquet data to a new parquet file or CSV format
Retrieve audit log entries for modifications made to parquet files
Get the schema (column names and types) of a parquet file
Get statistical summary of numeric columns in a parquet file (count, mean, std, min, max, quartiles)
Output schemas not documented. No tool describes what fields are returned, their types, or structure. LLMs cannot plan downstream composition or extract specific fields without trial-and-error.
Error handling lacks recovery guidance. No tool specifies what errors can occur, whether they are retryable, or what the LLM should do next. Stack traces and HTTP codes provide no actionable context.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Insert new rows into a parquet file. Creates the file if it doesn't exist.
List all available data types with their parquet file status
List all available snapshots for a given data type
Query a parquet file with filtering, sorting, and column selection. Supports advanced filtering with operators like $contains, $regex, $fuzzy, $gt, $lt, etc.
Read a parquet file and return its contents as rows
Restore a parquet file from a snapshot backup
Get a random sample of rows from a parquet file
Perform semantic search on a parquet file using OpenAI embeddings (requires OPENAI_API_KEY). Returns rows ranked by semantic similarity to the query text.
Update rows in a parquet file that match the given filters. Supports the same filter syntax as query_parquet.
Destructive operations (delete_rows, restore_from_snapshot, update_rows) lack confirmation/dry-run capability. An LLM could accidentally delete all rows in a table without safeguards.
No pagination guidance. Tools like list_data_types, list_snapshots, and get_audit_log accept 'limit' but do not document result count, total count, or next_cursor pattern. Large result sets risk context overflow.
Descriptions are generic and lack LLM-optimization. Many descriptions (e.g., 'Get the schema', 'Create a snapshot') are under 50 characters and do not explain WHEN to use the tool or what it returns. Baseline for A+ tools is 50-200 chars with clear context.
No permission checks or scope declarations. Tools expose write and destructive operations without verifying authorization. Audit logging exists but no gate prevents unauthorized agents from deleting data.
query_parquet and update_rows expose low-level filter syntax ($contains, $regex, $fuzzy, etc.) without explaining that these are custom operators. LLMs may not understand the syntax or may hallucinate operators.
semantic_search requires OPENAI_API_KEY but does not document this dependency or graceful degradation if the key is missing. LLMs could fail silently or produce confusing errors.