Programmatic data access layer for LLM agents - navigate context, don't memorize it. Provides schema discovery, querying, chunking, relationship graphs, and triplet extraction across multiple data sources.
The server provides 12 tools with clear naming and mostly complete schemas. All tool names follow verb-noun convention (cgc_add_source, cgc_list_sources, etc.), making intent obvious. Input schemas are present for all tools with proper type definitions and required field declarations. However, descriptions are inconsistent in quality: some are adequate (50-100 chars), others are sparse or lack context about when/why to use them. Parameter descriptions are present but often minimal (e.g., 'Source ID to discover' could explain what discovery entails). Output schemas are not documented, LLMs cannot plan downstream calls without knowing what each tool returns. Error handling is not evident in the tool definitions themselves. Overall: solid foundation (naming, schemas, parameters) but gaps in description depth and output documentation prevent higher scores. Baseline indicates avg tool description length 194 chars; this server averages ~80 chars, suggesting underdocumentation.
Add a data source to CGC. Supports: postgres, mysql, sqlite, filesystem, qdrant, pinecone, pgvector, mongodb
Chunk data for LLM processing. Returns data in manageable pieces.
Discover schema for a data source. Returns tables/files, fields, relationships.
Discover schemas for all connected sources
Find all records related to a specific value across sources
Get relationship graph showing how entities across sources are connected
List all connected data sources
Output schemas not documented for any tool. LLMs cannot plan multi-step workflows or extract chaining IDs without knowing return structure.
Descriptions are sparse and lack context. Average length ~80 chars vs baseline 194 chars. No explanation of WHEN to call each tool or its relationship to others (e.g., 'Call cgc_discover first to understand schema, then cgc_sample to preview data').
Destructive tool (cgc_remove_source) lacks confirmation/dry-run pattern. No indication that this operation is irreversible.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Remove a data source
Get sample data from an entity (table/file). Use this to understand data before querying.
Search for data using ILIKE pattern matching with trigram fallback
Execute a SQL query against a database source
Get a summary of all connected sources and their relationships (truncated in source)
cgc_sql tool accepts arbitrary SQL without validation guidance in description. No documentation of SQL injection prevention, read-only constraints, or what queries are unsafe.
No error handling guidance visible in tool definitions. When a source doesn't exist or a query fails, LLMs won't know what corrective action to take.
Enum constraint for 'source_type' in cgc_add_source is good, but description lacks format/validation rules for 'connection' parameter (what format for postgres vs filesystem?).
cgc_chunk strategy parameter lacks precise documentation. 'rows:N, tokens:N, or sections' is vague, what values of N are valid? What does 'sections' mean? How should LLM choose?