Use natural language to write and execute queries on your organizational data in DX Data Cloud
The server provides 10 tools with mostly clear naming and reasonable descriptions. However, there are significant gaps in parameter documentation, output schema specification, and error handling guidance. Tool names follow verb_noun conventions well (queryData, listEntities, getEntityDetails, etc.), and descriptions provide context for selection. The main weaknesses are: (1) parameter descriptions are inconsistent in quality and completeness, some parameters lack descriptions or are vague; (2) output schemas are not formally documented for most tools, forcing LLMs to guess response structure; (3) error handling returns generic error objects without recovery guidance; (4) no tool annotations (readonly hints, destructive hints) despite the server having both READ_ONLY and WRITE operations; (5) some descriptions are under 100 chars, missing important context about when to use each tool. The code shows competent implementation with proper validation in entity_tools.py and scorecard_tools.py, but the tool interface definition itself needs refinement for production LLM agent use.
Get comprehensive details about a specific entity including its information, tasks, and scorecards - we can use this to check operational readiness/health of an entity.
Get initiative details including both the initiative info and its progress report. Note: This calls two endpoints: - initiatives.info - initiatives.progressReport
Retrieve details about a specific scorecard, including its defined levels and checks.
Retrieve details for an individual team. Note that searching by team_emails will return things like the team name and members, where the search by team_id/reference_id will return more detailed information about the team structure.
List entities from the DX software catalog.
Output schemas not documented. Tools return dict/str but LLMs cannot infer the shape of response fields, forcing them to guess what keys/structure to expect. This breaks tool chaining and forces wasteful discovery calls. Example: listEntities returns entities and next_cursor but this is not formally declared, and getEntityDetails returns result['entity'], result['tasks'], result['scorecards'] with potential error keys, LLM cannot plan downstream calls without knowing which field contains what.
Parameter descriptions incomplete or missing for several tools. reviewTasks takes 'check_ids' as comma-separated string but no description states the format or constraint (is it comma-separated? space-separated? how many?). getTeamDetails has three optional parameters (team_id, reference_id, team_emails) with good descriptions, but other tools like listTeams have no parameters documented. This forces LLMs to infer what fields control behavior.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 57 | - | v1 |
Lists all initiatives with summary information.
List all active scorecards.
List all teams in DX.
Execute a SQL query against the DX Data Cloud PostgreSQL database. Always query from information_schema if you are uncertain about which tables and columns to look at.
Review/resolve/complete outstanding DX tasks (failing checks).
Error handling does not guide recovery. When WEB_API_TOKEN is missing, tools return {"error": "WEB_API_TOKEN environment variable is not set"}, this is a server-side configuration issue, not something the agent can recover from. When an API call fails (e.g., listEntities returns {"error": "API error: ..."} from the DX API), the response does not indicate whether the error is retryable, what caused it, or what the agent should do next. No distinction between user-fixable errors (wrong identifier) and fatal ones (service down).
No tool annotations despite mixed READ_ONLY and WRITE operations. The server declares reviewTasks as WRITE and all others as READ_ONLY in the evaluation metadata, but the tool definitions lack readOnlyHint/destructiveHint/idempotentHint annotations. LLMs cannot see which tools are side-effecting without reading descriptions carefully, risking accidental execution of irreversible operations.
Bare environment variable configuration creates runtime fragility. The server requires DB_URL, DX_API_HOST, and WEB_API_TOKEN to be set but does not validate them at startup. If these are missing or wrong, errors surface at tool invocation time with cryptic messages ("Database Error: ..." from psycopg). No health check or initialization validation.
Pagination not explicitly described in tool descriptions. listEntities, listInitiatives, listScorecards all accept cursor and limit parameters but the descriptions do not explain the pagination strategy or warn about hitting API limits. getInitiativeDetails passes pagination params to initiatives.progressReport but does not document this composition clearly.
Multi-step tools lack clear composition notes. getEntityDetails calls three endpoints (entities.info, entities.tasks, entities.scorecards) and returns all results in a single response, but the description does not state that failures in one step do not block others. The response includes separate _error keys (entity_error, tasks_error, scorecards_error), which is good, but LLMs may not know to check all of them.