Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server suffers from critical and widespread definition gaps. Of 19 tools, 13 have NO visible input schemas in the source code (tools 7-19), and most lack substantive parameter documentation. The 6 agent management tools (1-6) have explicit schemas with descriptions, but 13 Azure DevOps, GitHub, Wiki, and Slack tools show only empty {} input schemas with no visible parameter definitions. Tool names are reasonable (verb_noun pattern mostly followed), but without schemas and detailed parameter docs, LLMs cannot reliably invoke these tools. Descriptions exist but are often generic. The server represents a skeleton implementation where most tools lack the structured metadata needed for production agent use.
Tools (19)
bulkManageAgentswritesource verified53/100
Sends a batch of instructions to multiple agents in a single request. Can be used to launch, instruct, or shut down agents.
create_sprintwriteauth30/100
Create a new sprint in Azure DevOps.
create_work_itemswriteauth30/100
Create a new work item in Azure DevOps.
enrich_work_itemwriteauth30/100
Enrich a work item with additional information or metadata in Azure DevOps.
execute_wiqlread onlyauth25/100
Execute a WIQL query on Azure DevOps, returning the results.
getAgentStatusread onlysource verified73/100
Gets the status of a specific agent.
get_github_file_contentread onlyauth30/100
Retrieve the content of a file from a GitHub repository.
13 of 19 tools have NO visible input schemas (empty {} definitions). Tools 7-19 lack any parameter type or description metadata visible in source code.
Tool descriptions are generic and brevity-limited (50 chars or less for most tools 2-5). Descriptions like 'Lists all active agents' lack context on when to call, what to expect, and dependencies. Per pattern baseline, descriptions should be 50-200 chars with clear intent signals.
Add complete input schemas for all 13 Azure DevOps, GitHub, Wiki, and Slack tools. Each parameter must have: name, type (string|number|boolean|object|array), description (50-150 chars explaining what it does and constraints), required (true/false), and enum/pattern if applicable. Example: {"sprint_id": {"type": "string", "description": "The UUID of the sprint. Retrieve via get_sprints().", "required": true}}
Expand tool descriptions to 100-150 chars with WHEN to call, WHAT it does, and WHAT it returns. Example for getAgentStatus: 'Retrieve the current execution status and results of a running agent. Use after launchAgent() to monitor progress. Returns status (running|completed|failed), result, and iteration count.'
Add output schema documentation for every tool. Use a comment block in the tool definition or description field. Example: 'Returns: {agent_id: string, status: "running"|"completed"|"failed", iterations: number, result: string}'
Rename 'wiki' to 'update_wiki' or 'create_wiki_page' or split into separate tools (get_wiki, create_wiki, update_wiki). Generic names like 'manage' or 'wiki' are not actionable. Add description: 'Create or update a wiki page in the Azure DevOps project. Specify page_path (e.g., "Architecture/Design") and content in Markdown.'
Add enum constraints for multi-valued parameters. Example for bulkManageAgents: update the 'operations' parameter schema to include examples of valid JSON structures and document allowed 'action' values ('launch'|'instruct'|'shutdown'). Consider splitting into separate batch tools or requiring structured input validation.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 0 points across a rubric change (v1 → v2)
32/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
32
<=2025-11-25
v2
2026-03-09
F
32
-
v1
auth
35/100
Get all sprints in Azure DevOps, useful for the model to understand which sprints exist, especially when creating new sprints.
get_work_itemsread onlyauth30/100
Get the details of a work item in Azure DevOps.
instructAgentwritesource verified73/100
Sends a new instruction to a waiting agent.
launchAgentwriteauthsource verified81/100
Launches a new agent with a given system and user prompt.
listAgentsread onlysource verified75/100
Lists all active agents.
search_work_itemsread onlyauth30/100
Search for work items in Azure DevOps by keywords, abstracting away the WIQL query.
shutdownAgentdestructivesource verified73/100
Terminates a running agent.
slack_post_messagewriteauth30/100
Post a message to a Slack channel.
sprint_itemsread onlyauth25/100
Get the items in a sprint in Azure DevOps.
sprint_overviewread onlyauth25/100
Get the overview of a sprint in Azure DevOps.
update_work_itemswriteauth35/100
Update a work item in Azure DevOps. This should be capable of dealing with the full range of work item fields, including assignment, status, custom fields, sprint, relationships, comments, etc.
bulkManageAgents accepts a 'operations' parameter documented as a JSON string, but the schema definition shows it as a single string type with no enum or format constraint. This invites hallucinated JSON and parsing errors. Parameter docs must clarify format (JSON object array with required 'action' field) and valid action values ('launch', 'instruct', 'shutdown').
'wiki' tool has exceptionally vague name (generic, not verb-prefixed) and sparse description. Does it create, read, update, or delete wiki pages? The description 'Manage wiki pages in Azure DevOps' does not disambiguate. LLMs will struggle to choose this tool.
No error handling or recovery guidance visible. Tools lack descriptions of what errors might occur (e.g., 'agent_id not found', 'sprint does not exist') and what to do next. Per pattern baseline, errors should guide LLMs: 'Agent not found. Call listAgents() to see active agents.'
No documented output schemas. LLMs cannot predict what fields will be returned, preventing effective chaining and response parsing. E.g., does listAgents() return agent IDs or full agent objects? Are there pagination fields?
launchAgent requires 'temperature' and 'max_iterations' parameters but marks them as required with no defaults. Descriptions state 'Defaults to 1' and 'Defaults to 10' but parameters are required=true, creating ambiguity. LLMs may not know how to handle this contradiction.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in code. LLMs cannot infer which tools are safe to retry. shutdownAgent and other destructive tools should be marked clearly to prevent accidental re-execution.
Add error handling guidance in descriptions. Example: 'If the agent ID is not found, call listAgents() to retrieve valid IDs. If the sprint does not exist, call get_sprints() first.' Use the recovery-guide pattern.
Mark destructive tools with idempotentHint=false. For launchAgent, instructAgent, shutdownAgent, bulkManageAgents, create_sprint, create_work_items, update_work_items, wiki, and slack_post_message, add a note: 'This tool modifies state. Verify inputs before execution. Cannot be safely retried without manual review.'
Standardize tool naming convention. Choose either list_*/get_* consistently. Prefer 'list_' for discovery (list_agents, list_sprints, list_work_items) and 'get_' for retrieval by ID (get_agent_status, get_sprint_details). Avoid 'sprint_items', use 'list_sprint_items' or 'get_sprint_items'.
For parameters accepting human-friendly identifiers (e.g., agent names, sprint names, work item titles), add separate parameters or a lookup note. Example: 'agent_id can be a UUID or agent name. If you have a name, call listAgents() to find the ID.'
Add pagination support and result limits to list/search tools. Example for search_work_items: add 'limit' (max 100, default 20) and 'offset' (default 0) parameters. Document: 'Returns up to limit results. Use offset to retrieve subsequent pages.'
Document parameter relationships. For update_work_items, clarify which fields are updateable and which are read-only. For bulkManageAgents, provide a clear JSON schema example in the parameter description showing required and optional fields per action type.