MCP server for AI-assisted Glean connector development
The Glean Connector MCP server demonstrates moderate definition quality with substantial gaps. All 19 tools are explicitly registered with descriptions and verb-based naming (get_, create_, update_, etc.), which is a strong foundation. However, input schemas are largely missing or invisible in the provided source code excerpt. Only tool #8 (confirm_mappings) has a visible, detailed schema with typed parameters. The remaining 18 tools reference schema constants (e.g., getStartedSchema, createConnectorSchema) that are imported but not shown, preventing verification that parameters have type definitions and descriptions. This is a critical gap: per HARD SCORING RULES, tools with no visible schema must score 0 for schema assessment. Descriptions are generally well-written (averaging 120-180 chars), clearly stating what tools do and when to use them, but several lack actionable error recovery guidance. The server lacks tool annotations (readOnlyHint, destructiveHint, idempotentHint), error classification, and confirmation patterns for irreversible operations, despite having high-risk tools like run_connector and build_connector. Overall quality is above-average for community servers but falls short of production-grade implementation.
Deep-dive on a single field from the current schema: samples, type details, and Glean mapping suggestions.
Generate Python connector files from schema + mappings + config. Use dry_run: true to preview the generated code without writing files.
Check that all required tools and credentials are installed and configured.
Save field mapping decisions to .glean/mappings.json. Merges with any existing mappings.
Scaffold a new Glean connector project using the standard template. Call this after get_started. Sets the active project directory for this session.
Read the current connector configuration from .glean/config.json.
18 of 19 tools have no visible input schemas in source code. Schemas are defined in separate files (imported as constants) but not shown for verification. Per HARD SCORING RULES, invisible schemas must score 0. Only confirm_mappings has a documented schema with typed parameters and descriptions.
High-risk operations lack confirmation patterns. run_connector (IRREVERSIBLE), build_connector (WRITE), and manage_recording (WRITE/DELETE) lack dry_run, confirmation, or explicit error recovery guidance. Users/agents can inadvertently destroy work or trigger costly external processes.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 13 | - | v1 |
Read data_client.py for a module — use before asking AI to implement real API calls.
Return the current source schema alongside the Glean entity model so you can decide which source field maps to which Glean field.
Read the current field schema from .glean/schema.json.
Entry point for building a Glean connector. Returns an opening prompt that orients the AI and asks the user what data source they want to connect. Call this before any other tool.
Parse a data file (.csv, .json, .ndjson) and return field analysis: detected types, null rates, cardinality, and sample values. Use this to understand the source data before defining mappings.
Check execution status and retrieve records. Returns status, records fetched, per-record validation results, and recent logs.
List all connector classes found in this project with their module paths.
Manage connector recordings. action: "record" saves fetched data, "replay" runs from a saved file, "list" shows available recordings, "delete" removes one.
Start async execution of the Python connector. Returns an execution_id immediately. Poll status with inspect_execution.
Write connector configuration to .glean/config.json. Merges with existing config. Code-generation keys (consumed by build_connector): name, display_name, datasource_category, url_regex, icon_url, connector_type. Runtime-only keys (used by the connector at execution time, not during code generation): auth_type, endpoint, api_key_header, page_size, rate_limit_rps.
Write a new data_client.py implementation (replaces the mock with real API calls).
Write field definitions to .glean/schema.json. Call this after infer_schema to save the agreed schema, or to make manual edits. Set merge: true to merge incoming fields with the existing schema (fields with the same name are replaced; new names are appended) instead of replacing the entire field list.
Check current mappings against Glean's entity model. Reports missing required fields and type mismatches.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are not present in visible code. These annotations help LLMs understand retry safety and side-effect severity. Tools like delete (in manage_recording) should be marked destructive; get_* tools should be marked read-only.
manage_recording uses a polymorphic 'action' parameter (record/replay/list/delete). This design forces the LLM to reason about action enum values. Better: split into record_connector_output, replay_connector_output, list_connector_recordings, delete_connector_recording, each with a single, clear responsibility.
No visible error recovery guidance in tool descriptions. Error handling pattern (pattern:recovery-guide) requires actionable recovery steps. E.g., build_connector should explain: 'If dry_run shows errors, call analyze_field on affected fields before retrying.'
Output schemas are not documented in visible code. Pattern:tool requires documented return types so LLMs know what fields to expect for downstream chaining. E.g., run_connector should document that it returns execution_id (needed for inspect_execution), status, and timestamps.