MCP server for the Academy research platform, facilitating multi-agent AI conversations with PostgreSQL backend for session management, message tracking, analysis snapshots, and experiment runs
The academy-mcp-server defines 28 tools with basic structure, but exhibits significant quality gaps across naming, descriptions, and schema clarity. While all tools have names and most have descriptions, the descriptions are often generic and fall short of production standards. Parameter schemas are visible but lack depth, many parameters have minimal documentation (1-2 word descriptions), and there is no evidence of output schema documentation. No tool annotations (readOnlyHint/destructiveHint/idempotentHint) are present despite clear risk categories being assigned (READ_ONLY, WRITE, DESTRUCTIVE). Error handling is not evident in the source code provided. The tool set mixes session/chat management with direct LLM API calls (claude_chat, openai_chat, etc.), which violates single-responsibility principle. Naming is mostly consistent with verb_noun conventions but lacks disambiguation for similar tools (e.g., five separate chat tools for different providers).
Adds a new message to a session from a participant
Adds a new participant (AI or human) to a session
Performs analysis on a conversation and saves the snapshot
Calls Claude API directly for chat completions
Clears all analysis snapshots for a session
Calls Cohere API directly for chat completions
Creates a new experiment configuration
Eight direct LLM provider chat tools (claude_chat, openai_chat, grok_chat, gemini_chat, ollama_chat, deepseek_chat, mistral_chat, cohere_chat) violate single-responsibility principle. These should be consolidated into one parameterized tool or removed entirely, agents should not directly call LLM APIs.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are completely absent. Risk categories are assigned (READ_ONLY, WRITE, DESTRUCTIVE) but not surfaced in the MCP protocol, preventing clients from understanding tool safety characteristics.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 12 | - | v1 |
Creates a new chat session with specified name, description, and template
Calls DeepSeek API directly for chat completions
Deletes a session and all associated data
Exports all data from a session (messages, analysis, participants)
Calls Google Gemini API directly for chat completions
Retrieves all analysis snapshots for a session
Returns the ID of the currently active session
Retrieves results and data from an experiment run
Retrieves messages from a session, optionally filtered by participant or time range
Retrieves all participants in a session
Retrieves all chat sessions from the database
Calls Grok API directly for chat completions
Calls Mistral API directly for chat completions
Calls Ollama local model for chat completions
Calls OpenAI API directly for chat completions
Removes a participant from a session
Saves an analysis snapshot of the conversation state at a point in time
Starts a new run of an experiment
Updates the analysis configuration for a session (provider, model, window size)
Updates chat configuration for a session (context window, system prompt)
Updates session metadata such as name, description, and status
No output schema documentation visible in tool definitions. LLMs cannot predict what fields will be returned, forcing them to guess at follow-up tool parameters and increasing error rates.
Parameter descriptions are generic and insufficient. Examples: 'Optional message metadata' (9 chars), 'Participant settings' (20 chars), 'Analysis data' (12 chars). These provide minimal guidance for LLM parameter selection and fail the 50-200 character production baseline.
No evidence of error handling, recovery guidance, or error categorization in source code. Tools provide no actionable next steps when failures occur, leaving agents unable to self-correct.
Naming ambiguity in chat tools. Eight tools with identical parameter sets and near-identical descriptions (claude_chat, openai_chat, etc.) force LLMs to reason about provider selection repeatedly. Either consolidate into a single parameterized tool with 'provider' enum or implement provider-agnostic wrapper.
Parameters 'metadata' (add_message, add_participant) and 'settings' (add_participant) accept free-form objects with no schema constraints. LLMs cannot determine what keys/structure are valid, leading to malformed calls.
Enum constraint for 'template' parameter in create_session lists only 'blank', incomplete. Either document all valid templates or provide a discovery tool (list_session_templates).
Tools that modify state (create_session, add_message, delete_session) lack idempotency guidance. No indication whether repeated calls with same parameters are safe or will create duplicates.