Production-ready multi-tenant AI agent SaaS platform with MCP server integration, WhatsApp/Telegram/Slack channels, OpenRouter model provider, and skill execution
This MCP server exhibits significant quality gaps across naming, descriptions, parameters, and error handling. While 10 tools are defined with basic schemas, most lack sufficient clarity for reliable LLM operation. Tool names are inconsistent (some use verbs like 'send', others are nouns like 'chat', 'health_check'). Descriptions are present but often generic and under-optimized for LLM decision-making. Parameter schemas lack type information in many cases, and descriptions are minimal. No output schemas are documented. Error handling is absent, there is no indication of recovery guidance, error classification, or actionable error messages. The server mixes read-only and write operations without clear permission boundaries or audit capability. Security considerations for the chat and channel_send tools are not evident.
Send message through a specific channel
Get channel statuses from the channel manager
Process chat message via OpenRouter with model selection and response generation
Read local files
Get system health status
List available models from OpenRouter
Execute a skill through MCP
Get available skills from the skills platform
Tool name 'chat' is generic and does not follow verb_noun convention. It violates the pattern that tool names should be specific action verbs (e.g., 'send_message', 'generate_response', 'call_model'). An LLM cannot infer intent from 'chat', it could mean joining a chat, sending a message, retrieving history, or analyzing sentiment.
No output schemas documented for any tool. The rubric requires documented return types for all tools (100% of A+ tools have documented return types). Without knowing what fields the agent receives, LLMs cannot plan multi-step chains or extract required data. For example, does 'skills_list' return { skills: [{id, name, description}] } or { data: [...], meta: {...} }? This ambiguity prevents downstream tool composition.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 45 | <=2025-11-25 | v2 |
Get system statistics including uptime, memory usage, active sessions, and user tier information
Search the live web
Parameter descriptions are minimal and lack actionable constraints. For example, 'skill_execute' has a parameter 'skillId' with description 'ID of the skill to execute', no format, no example range, no guidance on how to obtain a skill ID. The rubric requires descriptions to include expected format, range, and constraints (e.g., 'UUID format, obtainable via skills_list').
Tool 'chat' accepts a 'model' parameter with a default but no enum constraint. The description says 'Model to use (defaults to openrouter/anthropic/claude-3.5-sonnet)', but does 'model' accept arbitrary strings, or only a known set? An LLM will hallucinate invalid model names. Define an enum of valid models or fetch the list dynamically via 'models_list'. Additionally, 'temperature' and 'maxTokens' have no documented bounds; an LLM could pass temperature=999 or maxTokens=999999, causing API failures.
No error handling or recovery guidance documented. None of the tools describe what happens on failure, what errors might occur, or what the LLM should do next. For example, if 'web_search' hits a rate limit or network error, should it retry? Wait? Call a fallback? The rubric requires error responses to 'tell the LLM what to do next' and 'categorize errors as retryable, user-fixable, or fatal'.
Tool 'channel_send' is a destructive (WRITE) operation but has no confirmation step or dry-run capability. The rubric states 'Irreversible operations (delete, send, publish) should support a dry-run or confirmation step. Agents make mistakes, a confirm_before_execute pattern prevents catastrophic errors.' Sending a message to the wrong channel could cause business damage.
Parameters lack type information in schema. For example, 'chat' accepts 'temperature' and 'maxTokens' but the schema type is 'number' without min/max bounds. The rubric requires 'Specify minimum and maximum for numeric parameters (e.g. page_size 1 - 100, days 1 - 365). Unbounded numbers let LLMs pass absurd values that break APIs or cause timeouts.' This server provides no validation or bounds.
No permission gates or scope declarations. Tools like 'channel_send' and 'skill_execute' perform write operations, but there is no indication of what permissions are required, what audit trail is kept, or whether the calling user/agent is authorized. The rubric requires 'Each tool should declare what permissions it requires (e.g. 'read:email', 'write:calendar') ... for least-privilege agent configurations and clear audit trails.'
Tool 'skills_list' and 'channels_list' do not document pagination. If a system has hundreds of skills or channels, the full list would blow the context window. The rubric requires 'Tools returning lists should accept page/offset and limit parameters and return a total count or next_cursor.' These tools show empty input schema, no pagination parameters are visible.
Tool 'file_read' accepts a 'path' parameter with no format constraints or security guidance. No mention of path traversal prevention, file size limits, or allowed directories. The rubric requires 'Treat all agent-provided input as untrusted. Sanitize against SQL injection, command injection, and path traversal. LLMs can be tricked via prompt injection into passing malicious payloads.'
Tool naming is inconsistent. Some tools use verb_noun (web_search, file_read, skill_execute, channel_send, health_check), while 'chat' is a bare noun (should be something like 'send_message' or 'call_llm'). 'skills_list', 'channels_list', 'models_list', 'system_stats' are noun-based (should be list_skills, list_channels, list_models, get_system_stats). This inconsistency forces LLMs to reason about naming conventions rather than relying on a single pattern.