A barebones library for agents that write Python code to call tools or orchestrate other agents.
This smolagents MCP server exhibits significant gaps in definition quality. Of 17 tools analyzed, naming is generally clear with action verbs (get_, search_, display_), but descriptions vary widely in quality. Most tools (12/17) have basic descriptions, but 5 lack meaningful detail. Parameter descriptions are present but often generic. Critical issue: tools 16 and 17 are duplicates (both 'web_search'), indicating incomplete server design. Schema quality is mixed, tools like get_weather, convert_currency, and python_interpreter have well-formed input schemas with typed parameters and descriptions. However, output schemas are not documented for any tool, making downstream composition error-prone. No tool includes error recovery guidance. The python_interpreter tool (tool 13) is marked IRREVERSIBLE but has no confirmation step, dry-run, or recovery path. Display tools (9-12) have generic descriptions ('Display a pandas DataFrame to the user in a formatted way') that lack guidance on WHEN to use them vs alternatives. inspect_file_as_text (tool 8) has an overly long, rambling description (275+ chars) that buries the key constraint. No security scoping, audit hints, or permission declarations. The server lacks any tool annotation hints (readonly, destructive, idempotent), forcing the client to infer consequences from tool names alone.
Converts a specified amount from one currency to another using the ExchangeRate-API.
Display a chart visualization to the user. This tool creates a text-based description of chart data since direct image rendering is not available in text outputs.
Display a pandas DataFrame to the user in a formatted way. This tool renders a DataFrame as a readable markdown table, suitable for display in agent outputs.
Display JSON data to the user in a formatted way.
Display a summary of data to the user. Provides statistical summary for numerical data and value counts for categorical data.
Provides a final answer to the given problem.
Duplicate tool names: Two tools named 'web_search' (one for DuckDuckGo, one for Google) create ambiguity. LLMs cannot disambiguate which to call without additional context. Should be renamed to 'search_web_duckduckgo' and 'search_web_google' or consolidated into a single 'web_search' with a provider parameter.
Missing output schemas: No tool documents its return structure. LLMs cannot plan downstream calls or extract specific fields without knowing what data is returned. For example, get_weather returns a string, but does search_wikipedia return structured fields or plaintext? This forces LLMs to make assumptions and causes composition failures.
Irreversible tool without confirmation or dry-run: python_interpreter (tool 13) can execute arbitrary code but has no safety guardrail. No dry-run option, no confirmation step, no undo mechanism. An agent mistake could delete files, crash systems, or leak data. Should implement confirmation-before-execute pattern.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Fetches a random joke from the JokeAPI. This function sends a GET request to the JokeAPI to retrieve a random joke. It handles both single jokes and two-part jokes (setup and delivery). If the request fails or the response does not contain a joke, an error message is returned.
Fetches the top news headlines from the News API for the United States. This function makes a GET request to the News API to retrieve the top news headlines for the United States. It returns the titles and sources of the top 5 articles as a formatted string. If no articles are available, it returns a message indicating that no news is available. In case of a request error, it returns an error message.
Fetches a random fact from the "uselessfacts.jsph.pl" API.
Fetches the current time for a given location using the World Time API.
Get the current weather at the given location using the WeatherStack API.
You cannot load files yourself: instead call this tool to read a file as markdown text and ask questions about it. This tool handles the following file extensions: [".html", ".htm", ".xlsx", ".pptx", ".wav", ".mp3", ".m4a", ".flac", ".pdf", ".docx"], and all other types of text files. IT DOES NOT HANDLE IMAGES.
This is a tool that evaluates python code. It can be used to perform calculations.
Fetches a summary of a Wikipedia page for a given query.
Asks for user's input on a specific question
Performs a google web search for your query then returns a string of the top search results.
Performs a duckduckgo web search based on your query (think a Google search) then returns the top search results.
Generic, uninformative descriptions for display tools: Tools 9-12 (display_dataframe_to_user, display_chart_to_user, etc.) have descriptions like 'Display a pandas DataFrame to the user in a formatted way.' These lack guidance on WHEN to use each tool (when is a chart better than a table?), what formats are supported, or what data structures work. LLMs cannot differentiate between similar display tools.
Overly long, rambling description: inspect_file_as_text (tool 8) has a 275+ character description with verbose explanations and all-caps warnings ('DO NOT use this tool for an HTML webpage'). This buries key info and wastes tokens. Should be condensed to 10-100 chars with critical constraints moved to parameter descriptions.
No error recovery guidance: Tools that call external APIs (get_weather, convert_currency, get_news_headlines, etc.) can fail, but none document error categories (retryable, user-fixable, fatal) or recovery steps. Descriptions should say 'If API key is missing, try again' or 'If the location is invalid, ask user to clarify.'
No tool annotations: No tool declares whether it is read-only, destructive, or idempotent. The MCP spec supports tool annotations (readOnlyHint, destructiveHint, idempotentHint). python_interpreter should be marked destructive; web_search should be marked read-only. Without these hints, clients cannot optimize caching or retry logic.
No security scoping or permission declarations: Tools that call external APIs or execute code (python_interpreter, web_search, get_weather) have no declared permissions or scope requirements. The server should declare read:web, write:code, read:weather etc. This enables least-privilege agent configuration and audit trails.
Missing parameter constraints and validation guidance: Many parameters accept freeform strings without enums or format constraints. For example, get_time_in_timezone's 'location' parameter accepts any string but expects 'Region/City' format, this should be enforced with a regex pattern or enum of known zones in the parameter description. Display tools accept 'chart_type' as a string with no enum (should be 'line'|'bar'|'scatter'|'pie').
No documentation of API key requirements: Tools like get_weather, convert_currency, get_news_headlines, and web_search call external APIs that require keys, but descriptions do not state 'API key required' or provide setup guidance. The actual code comments show placeholder API keys, which should never appear in example code.