Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
Agent Tiers exposes 12 tools via FastMCP HTTP transport, but suffers from significant definition quality gaps. Tool names are generally verb-based and clear (health_check, chat, upload_document, etc.), but parameter descriptions are minimal or missing context. Most critically, 8 of 12 tools have vague or generic descriptions that fail to distinguish them from one another or explain prerequisites. Input schemas are present but sparse, many parameters lack type constraints, enums, or validation guidance. Output schemas are not documented. Error handling is absent from tool definitions. The server reads as a thin MCP wrapper over existing FastAPI endpoints rather than a tool-optimized interface designed for LLM consumption.
Tools (12)
add_messagewrite50/100
Add message to session.
chatwritesource verified65/100
Non-streaming chat endpoint.
chat_streamwritesource verified63/100
Streaming chat endpoint using Server-Sent Events.
create_sessionwritesource verified68/100
Create new session.
delete_sessiondestructivesource verified67/100
Delete session.
get_user_sessionsread onlysource verified70/100
Get all sessions for a user.
health_checkread onlysource verified78/100
Health check endpoint.
hybrid_searchread onlysource verified65/100
Perform hybrid (vector + BM25 text) search on documents.
Descriptions lack context distinguishing similar tools. chat vs chat_stream, vector_search vs hybrid_search, and upload_document vs ingest_documents are not clearly differentiated. An LLM cannot determine which to call without trial and error.
Parameter descriptions are minimal. 'search_type' in chat/chat_stream lacks enum values or valid options. 'metadata' is described only as 'Optional metadata' without clarifying its structure or purpose. 'config' in ingest_documents is a bare object with no schema constraints.
No output schemas are documented. Tools like chat, list_messages, vector_search, and hybrid_search return responses, but the caller cannot see what fields to expect or how to chain them into downstream tool calls.
chat
Recommendations
Expand tool descriptions to 100 - 200 characters, following the LLM-optimized baseline. Each description must state: What does it do? When should I call it instead of the similar tool? What does it return? Example: 'chat: Non-streaming synchronous chat endpoint, use for single-turn interactions when you want the full response before proceeding. Returns message text and session ID. For continuous conversations with multiple turns, use chat_stream instead.'
Define enums for all constrained parameters. chat/chat_stream 'search_type' should list valid values (e.g. 'vector', 'hybrid', 'bm25', 'none'). ingest_documents 'config' should be a documented object with explicit fields (e.g. chunk_size: integer, strategy: enum('semantic'|'fixed'), etc.).
Document output schemas for every tool. Define the response shape as JSON Schema. Example for chat: { message: string, session_id: string, timestamp: ISO8601, tool_calls?: [{ name: string, result: any }] }. This enables the LLM to extract required fields and chain results into downstream calls.
Add pagination guidance to list_messages and get_user_sessions. Include total_count in the response so the LLM knows if results are truncated. Document the behavior of 'limit' (e.g. 'Default 20, max 100'). If cursor-based pagination is used, return a next_cursor field.
Add error recovery hints to all tools. Example: 'If chat fails with 'session not found', try create_session first. If vector_search returns empty results, try hybrid_search with lower text_weight.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 57 points across a rubric change (v1 → v2)
57/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
57
<=2025-11-25
v2
2026-03-09
F
0
-
v1
ingest_documentswritesource verified62/100
Ingest documents from folder with semantic chunking and extraction.
list_messagesread only50/100
List messages in a session.
upload_documentwritesource verified55/100
Upload and process documents with optional ingestion configuration.
No pagination hints in list tools. get_user_sessions and list_messages accept optional 'limit' but no offset, cursor, or total count guidance. An LLM cannot determine if results are truncated or how to fetch the next page.
No error recovery guidance in tool definitions. If chat fails, upload fails, or search returns no results, the LLM has no instruction on what to do next or whether to retry.
Destructive tools lack confirmation or dry-run support. delete_session irreversibly removes a session, but the description does not warn of consequences or suggest a confirmation pattern.
Parameters mix required and optional without clear defaults. upload_document and ingest_documents accept optional configuration objects but do not specify what happens if omitted. Will sensible defaults apply, or will the call fail?
Metadata and config parameters are overly generic. LLMs cannot determine what fields to pass inside a bare 'object' type. Structure these as explicit typed objects with documented fields.
Add confirmation or dry-run pattern to delete_session. Either require a confirm_delete: true flag, or offer a separate tool like describe_session_before_delete to let the LLM verify before destruction.
Clarify metadata and config parameter structures. Replace object type with explicit typed subobjects. Example for chat 'metadata': { conversation_context?: string, user_context?: string, preferences?: { language?: string, tone?: string } }.
Add dependency hints to tool descriptions. Example: 'To chat with a specific session, first call get_user_sessions to find the session_id, then pass it to chat.'
Document idempotency guarantees. Are create_session, add_message, upload_document idempotent? Can LLMs safely retry them without side effects?
Clarify the distinction between upload_document (single file, interactive?) vs ingest_documents (batch, folder-based). Explain when to use each and whether they populate the same vector store.