MCP server for querying other LLMs through a unified interface
The server defines 10 tools with complete input schemas and descriptions visible in src/ask_another/server.py. Naming follows verb_noun convention (query, list_models, research, get_research_job, etc.). However, several issues limit the score: (1) Output schemas are not documented, we can see input parameters clearly but no field definitions for return values. (2) Descriptions are present but generic; many lack detail on WHEN to call vs similar tools, what fields are returned, or error cases. (3) No pagination support despite tools like list_research_jobs and list_models that could return many results. (4) Error handling guidance is absent, error responses are not shown. (5) Tool annotations (readOnlyHint, destructiveHint) are declared in the type imports but appear unused. (6) The research task pattern uses job_id integers for correlation but no guidance on polling strategy, retry behavior, or job status values. Average tool score is 58/100, the server is functional but lacks production-grade polish.
Cancel a background research job
Get the top 5 favourite/most-used models based on usage annotations
Get the health/error status of all configured providers
Get models discovered in the last N days
Get status and result of a background research job
List all available models from configured providers
List all background research jobs
Query another LLM with a prompt and get a response
Output schemas not documented. Tools like list_models, list_research_jobs, get_favourites, and get_recent_models return data but no schema describes what fields are in the response. LLMs cannot plan downstream calls or extract specific fields without seeing the output structure.
Tool descriptions are generic and lack context. For example, 'List all background research jobs' does not explain job states (running/completed/failed), what fields are returned (job_id, status, result?), or when to call this vs get_research_job. Descriptions should answer: What does it do? When to use it? What does it return?
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 17 | - | v1 |
Start a background research task with another LLM
Send feedback about a model query result
No pagination for list tools. list_models, list_research_jobs, get_favourites, and get_recent_models lack page/limit/offset parameters. Without pagination, large result sets blow context windows. Tools should accept a limit parameter and return total_count or next_cursor.
Tool annotations imported but not used. The code imports ToolAnnotations but does not apply readOnlyHint, destructiveHint, or idempotentHint to tool definitions. cancel_research_job (REVERSIBLE) and send_feedback (WRITE) should be marked with appropriate annotations.
No error handling guidance in tool descriptions. If a model is not found in query tool, if a research job fails, or if a provider is unhealthy, users cannot see what error messages to expect or how to recover. Error responses should guide the LLM toward recovery (e.g., 'Model not found. Call list_models to see available models.').
Async job polling pattern undocumented. research() returns a job_id but no guidance on polling strategy (interval, timeout, max retries). get_research_job() retrieves status but no documentation of valid status values (running, completed, failed, etc.). LLMs must guess the polling behavior.
Natural-language model identifiers not documented. query() and research() accept 'model' parameter with examples like 'openai/gpt-4' and 'anthropic/claude-3-sonnet' but no regex pattern or format constraint. Should specify format (provider/model) and reference list_models for discovery.
Unclear composition with list_models vs get_provider_status. Both tools expose provider-level information but it is unclear when to use each or if they return overlapping data. Descriptions should clarify: list_models shows available models by provider; get_provider_status shows provider health. These are complementary, not redundant.