The server defines 3 tools with basic schemas and descriptions, but has significant gaps in completeness and LLM-optimized guidance. Tool names follow verb_noun convention (ai_code, get_models, ask_question), which is positive. However, descriptions are minimal (under 100 chars for 2 of 3 tools), parameters lack validation constraints and format guidance, output schemas are not documented, and error handling is absent. The server appears to work but would not pass production code review for agent tool quality.
Tools (3)
ai_codewriteauthsource verified63/100
Run Aider to perform coding tasks.
ask_questionread onlyauthsource verified50/100
Ask a question using the specified model and return the response.
get_modelsread onlysource verified58/100
List available Aider models filtered by substring.
Tool descriptions are too brief and lack LLM-optimized guidance. 'ask_question' (35 chars) and 'get_models' (65 chars) provide minimal context for when/why to use them. Baseline is 194 chars average; these fall well below. No mention of prerequisites, return structure, or error conditions.
Output schemas are not documented. The code returns `str` for ai_code/ask_question and `list[str]` for get_models, but there is no schema showing what fields are in the response, what format strings use, or what structure LLMs should expect. LLMs cannot plan downstream calls without knowing return structure.
Parameter descriptions lack format guidance and constraints. 'ai_coding_prompt' has no guidance on length, structure, or examples. 'substring' for get_models has no note on case-sensitivity or partial match behavior. 'model' in ask_question is documented as 'optional' but the schema shows no enum of valid models or guidance on format.
Recommendations
Expand ai_code description from 63 chars to 150-200 chars. Include: (1) What it does: 'Run Aider AI to perform code edits on specified files.' (2) When to use: 'Use when the user asks for coding changes, refactoring, or bug fixes.' (3) Prerequisites: 'Requires editable_files list (files that can be modified) and optionally readonly_files (reference files). (4) What it returns: 'Returns a summary of changes made and success/failure status.'
Expand get_models description from 65 chars to 120-150 chars. Add: 'Returns a filtered list of available Aider models. Use this before calling ai_code to choose a model. The substring parameter filters by name (case-insensitive partial match). Returns array of model identifiers suitable for the model parameter in ai_code and ask_question.'
Expand ask_question description from 35 chars to 100-140 chars. Add: 'Ask a question and get a response from the specified model. Use this for general queries, analysis, or clarification. The model parameter is optional; if omitted, uses the server's default editor model. Returns the model's text response.'
Document settings parameter schema in ai_code. Add a description: 'Optional configuration object for the Aider session. Accepted fields: timeout_seconds (int, 30-3600), max_edits (int, 1-100), model_override (str, valid model name). Invalid fields are ignored. If omitted, uses Aider defaults.'
Add structured output schemas. For ai_code, return: {success: bool, changes_summary: str, files_modified: [str], error: str|null}. For get_models, document: [str] (array of model names). For ask_question, document: {response: str, model_used: str, tokens_used: int|null}.
No error handling or recovery guidance. If ai_code fails (file not found, model unavailable, API error), the tool returns a generic string. No error classification (retryable vs user-fixable), no actionable guidance for the LLM, no information about what went wrong. This forces the LLM to guess the next step.
Settings parameter in ai_code is a generic object with no schema. Consumers cannot understand what fields are valid, what types they accept, or what values are allowed. This invites hallucinated settings and silent failures when invalid keys are passed.
Tool composition: ai_code likely needs a model name, but get_models only returns names as strings. If get_models returns ['gpt-4', 'claude-3'], there is no way to know which model is best for ai_code or what constraints apply (max_tokens, pricing, capability). No documentation links these tools as a logical workflow.
ask_question and ai_code appear to do similar things (prompt + model → response) but have different signatures and return types. LLM will have to reason about which to use. No clarification on when to use ask_question vs ai_code, or whether they overlap.
ask_questionai_code
Add error handling guidance to ai_code description: 'On failure, returns {success: false, error: <reason>}. If error is "model_unavailable", try get_models() to verify the model exists. If error is "file_not_found", check the relative path and ensure files are in the working directory.'
Clarify the distinction between ask_question and ai_code. Suggest in descriptions: 'Use ask_question for general questions, research, or analysis. Use ai_code when the user wants to modify files in the codebase.'
Add per-file success/failure in ai_code return. Instead of a single 'success' boolean, return: {success: bool, file_results: {filename: {success: bool, changes: str, error: str|null}}} so agents know which files were modified and why others failed.
Document the default model behavior. If model parameter in ask_question is omitted, clarify: 'Uses the editor_model specified at server startup (see --editor-model flag). To use a different model, pass the model name explicitly.'
Add pagination guidance if get_models can return large lists. E.g., 'Returns all matching models up to a maximum of 100. If more exist, the description will say "(and N more)", use a more specific substring filter to narrow results.'