An MCP server implementation that provides tools for agent-based task execution, evidence submission, answer generation, and integration with various external services like arXiv, AskNews, and ACI APIs.
Anemoi provides 22 tools with varying levels of definition quality. Strengths: most tools have descriptions (though quality varies), input schemas are present for all tools, and tool names follow verb-noun patterns. Critical weaknesses: many parameter descriptions are thin or missing context; output schemas are not documented; error handling lacks recovery guidance; no tool annotations (readOnlyHint/destructiveHint/idempotentHint) are present; several tools have overly broad or vague descriptions that don't disambiguate similar tools; no evidence of LLM-optimized descriptions (50-200 chars). The toolkit mixes user-facing operations (submit_evidence, send_answer) with infrastructure operations (configure_app, link_account) without clear separation. Parameter descriptions often fail to explain constraints, allowed values, or prerequisite relationships. Output structure is undocumented, agents cannot reliably extract IDs needed for tool chaining.
Configure an app with specified authentication type.
Delete an app configuration.
Downloads PDFs of academic papers from arXiv based on the provided query.
Enable a linked account.
Provides an assessment of the potential risk associated with a function.
Get app configuration by app name.
Get details of an app.
Output schemas not documented. Tools like send_answer, get_news, get_app_details, search_papers return data but no schema is specified for downstream tool chaining. LLMs cannot reliably extract needed fields (e.g., paper_id for download_papers, app_id for configure_app flows).
Many parameter descriptions lack actionable constraints. 'Name of the app to configure' doesn't explain: is it case-sensitive? What if the app doesn't exist? What are valid app names? Descriptions for app_name, asset, metric need concrete constraints or error recovery hints.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 42 | - | v1 |
List all linked accounts for a specific app.
Fetch news or stories based on a user query.
Fetch stories based on the provided parameters.
Perform a live web search based on the given queries.
Sends a message to the server indicating that the agent is giving up on the task. Requires at least one justification to be submitted first. Will include all submitted evidence and justifications.
Force ignores the risk associated with named function. This ONLY ignores the RISK for the NEXT Function Call.
Link an account to a configured app.
List all configured apps.
Query financial data and metrics for cryptocurrency and stock assets.
Searches for academic papers on arXiv using a query string and optional paper IDs.
Search Reddit based on the provided keywords.
Search for apps based on intent.
Sends the final answer to the server. Requires at least one justification to be submitted first. Will include all submitted evidence and justifications.
Submits a piece of evidence, quote, source, calculation, or other supporting material. Call this once per piece of evidence to build up a collection of supporting materials.
Submits a justification for the answer. Can be called multiple times to build up the complete justification. Multiple justifications will be combined together when sending the final answer. Requires at least one piece of evidence to be submitted first using submit_evidence().
No tool annotations. Tools like delete_app, send_answer, ignore_risk carry destructive or side-effecting semantics, but input schemas lack destructiveHint, readOnlyHint, or idempotentHint annotations. LLMs cannot assess risk without explicit hints.
Vague or thin descriptions for infrastructure tools. 'List all configured apps', 'Get app configuration', 'Get details of an app' are similar and lack differentiation. LLMs may conflate list_configured_apps, get_app_details, and get_app_configuration. Each needs to clarify: what fields are returned, when to use this vs a sibling tool?
Dependency relationships undocumented. submit_justification requires submit_evidence first (runtime check exists in code), but this constraint is not formalized in schema. Similarly, send_answer depends on submit_justification. No tool documents prerequisite steps or suggests how to handle violations (e.g., 'Call submit_evidence first.').
Error handling lacks recovery guidance. The code (SendAnswerTool.py) raises ValueError if evidence is missing, but the error message is not documented in the tool schema. LLMs do not know whether to retry, call submit_evidence, or ask the user for clarification.
Generic descriptions for search/query tools. 'Fetch news or stories based on a user query' and 'Perform a live web search' do not clarify: What sources are searched? What is the ranking/freshness? Are results deduplicated? Does get_news include opinion or only factual reporting? Agents waste calls trying to disambiguate.
Parameter descriptions under 50 chars for several infrastructure tools. 'Name of the app to configure', 'ID of the linked account to enable', these offer minimal context and force LLMs to infer intent. Baseline for A+ tools: 72 chars average per param.
Result limits not documented. search_papers, search_news, search_reddit accept max_results or similar params, but no tool description mentions the cap or default. If max_results defaults to 5 for search_papers but the description doesn't state this, agents may expect arbitrary-sized results and be surprised by truncation.
Pagination documented but inconsistently. search_tool, list_configured_apps expose limit/offset, but descriptions don't clarify: what is the max limit? Does offset work if not all results fit in memory? No total_count or next_cursor documented in output schema.