Backend service for managing MCP (Model Context Protocol) servers, tools, endpoints, and access policies with audit logging and RBAC
This MCP server has 23 tools split across mock utilities (tools 1-9, 16-23) and a production backend (tools 10-15). The mock tools have basic descriptions and schemas but lack depth and error handling guidance. The backend tools are better documented but expose raw API patterns without agent-friendly abstraction. Overall: naming is verb-heavy (get_, list_, create_, etc.) which is good, but descriptions are often generic and lack 'when to use' context. Most tools lack output schema documentation and error recovery guidance. Parameter descriptions exist but are minimal. No tool uses enums for constrained inputs (e.g., 'operation' in calculate accepts 'add|subtract|multiply|divide' as free-form string). No toolAnnotations (readOnlyHint, destructiveHint, idempotentHint) are present despite clear risk classifications. The server mixes mock tools (simple utilities) with real backend CRUD operations, which muddies the purpose and makes composition difficult.
Add two numbers and return the sum.
Convert bullet points into a polished paragraph.
Simple calculator. operation: add | subtract | multiply | divide.
Code-review prompt for a given snippet and language.
Create a new note. author_id must be an existing user ID.
Create a tool and write initial version metadata. Source: backend/app/routers/tools.py
Soft-delete or hard-delete a tool by id. Source: backend/app/routers/tools.py
No enum constraints on free-form string parameters. Tools like 'calculate' accept 'operation' as a free string ('add|subtract|multiply|divide') instead of an enum. LLMs frequently hallucinate values like 'modulo' or 'power' not in the list.
No output schemas documented. Tools return responses but LLMs have no schema to plan downstream calls. E.g., 'list_tools' returns a dict with 'tools' array, but the schema for each tool object is not documented.
Generic descriptions without 'when to use' context. E.g., 'resource_all_users' is described as 'All users in the system.' Does the LLM call this to populate a dropdown? To verify access? To audit? To validate a user exists before sending them a note? Unclear, forces LLM guessing.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
Discover tools from MCP server URL without registering it. Source: backend/app/routers/servers.py
Return the current UTC timestamp as an ISO-8601 string.
Look up a user by ID (try: u1, u2, u3).
Generate a greeting for a named user in formal or casual tone.
List all non-deleted tools from the registry. Source: backend/app/routers/tools.py
Generate a random integer between low and high (inclusive).
Create or update an MCP server registration after connectivity probe. Source: backend/app/routers/servers.py
All notes in the system.
All users in the system.
User role definitions and their permissions.
Server metadata and capability list.
Reverse the given string.
Ask the model to summarise a block of text within a word limit.
Transform text case. mode: upper | lower | title.
Update tool metadata/state and optionally append a new version. Source: backend/app/routers/tools.py
Count words and characters in the supplied text.
No error handling guidance. Tools like 'delete_tool' (DESTRUCTIVE risk) and 'register_server' (WRITE risk) have no recovery guidance. If delete_tool fails because the tool is in use, what should the LLM do? No actionable error messages shown.
Missing tool annotations (toolAnnotations feature absent). Tools marked with Risk=DESTRUCTIVE (e.g., delete_tool) and Risk=WRITE (e.g., create_tool) should carry readOnlyHint, destructiveHint, or idempotentHint to inform agent behavior. Without annotations, agents cannot reason about side effects.
Resource tools conflate listing and discovery. 'resource_all_users', 'resource_all_notes', etc. are named as resources but behave like tools (no input). Unclear distinction between tool and resource. Resources should expose data streams, not parameterless queries.
No pagination on list_* tools. 'list_tools' returns all tools with no limit, offset, or cursor. If 1000 tools exist, LLM receives 1000 items, blowing context window and degrading reasoning.
Inconsistent parameter naming and types. Some params like 'tone' in 'greeting_prompt' default to 'formal' (string), but no enum constraint. Others like 'user_id' in 'get_user' accept string but no format or length guidance. This forces LLMs to guess valid formats.
Prompts and resources mixed with tools in the same registry. Tools 1-9, 16-19, 20-23 are really prompts or resources, not executable tools. Exposing them as tools confuses agents about what is stateful, what requires parameters, and what is read-only.
Backend tools expose raw API patterns without agent-friendly abstraction. E.g., 'create_tool' requires 'owner_id' and 'source_type' as explicit parameters, but agents rarely know internal IDs or source type enums. Should accept natural identifiers and infer metadata server-side.