An MCP server that provides tools for weather forecasting, web search, Wikipedia queries, document analysis, Google Calendar integration, and a GenAI knowledge base with RAG capabilities
SammelsuriumMCP is a multi-tool server with significant structural issues. All 9 tools are explicitly registered with descriptions and basic input schemas, but quality is inconsistent. Naming follows verb_noun convention well (get_*, search_*, ask_*), which is correct. However, parameter descriptions are minimal or missing entirely, most parameters lack the detail needed for LLMs to infer correct usage. Output schemas are entirely undocumented; there is no specification of what fields the tools return, forcing LLMs to guess structure. Error handling is present but generic ('Something went wrong: {e}'), it does not guide recovery or categorize errors. The server mixes READ-ONLY tools (9/9 are safe) but lacks annotations to declare this. Several tools accept free-form string parameters without constraints (e.g., search_web query, location in get_current_weather), no enums or validation ranges. Parameter relationships are undocumented (e.g., search_wikipedia's topic and query parameters could clarify their interplay). The codebase shows try-catch blocks but returns unstructured error text rather than actionable recovery guidance. Compared to baseline metrics: avg tool description is ~100 chars (below 194 baseline), avg params per tool is 1.4 (well below 4 baseline), and 0% of tools have documented return schemas (vs 100% in A+ tools). This server is functional for basic tasks but lacks polish needed for production agent deployment.
Uses the GenAI knowledge base to answer the user's question. The knowledge base contains detailed information about GenAI in general, Retrieval Augmented Generation and the implemention in Python.
Returns a comma separated list of all available calendars.
Returns all entries of a specific calendar for the next N days.
Returns the current date
Returns the current location of the user
Returns the current weather forecast for a given location
No documented output schemas for any tool. LLMs cannot infer the structure of returned data (field names, types, nested objects). This forces agents to parse responses speculatively and wastes tokens on confirmation queries.
Generic error handling: all tools return 'Something went wrong: {e}' or similar. Errors do not classify as retryable/user-fixable/fatal, nor do they include the invalid value or suggest recovery steps. LLMs cannot learn from failures.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Load a file from the local computer to answer the user's question
To query unknown or recent data this tool can be used to search the internet
Get more information about a specific topic from Wikipedia to answer the user query
Free-form string parameters with no constraints or validation hints. 'search_web(query: str)' and 'get_current_weather(location: str)' accept any string with no guidance on format, length, or valid values. No enums, no regex patterns, no examples in constraints.
Parameter descriptions are minimal. 'The location for which to get the weather forecast' is generic; it does not specify format (city name, lat/long?), examples, or constraints. 'The search query' says nothing about length, syntax, or intended domain. Baselines show avg param description ~72 chars; most here are <40.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). All 9 tools are READ_ONLY but this is not declared in the tool schema. Agents cannot optimize for safe retry or concurrent execution.
Undocumented tool behavior: search_wikipedia returns a string but the internal logic calls query_wikipedia(topic, query, 'de') with a hardcoded language. This is invisible to the LLM, the agent cannot control language selection or understand why results might be in German.
No pagination or limit controls. Tools like get_calendar_entries return all events for N days with no batching. If a user has 1000 calendar events, the response could exceed context limits. No total_count, no next_cursor, no cap on results.