Multi-agent AI system with various specialized agents (Stock Agent, Data Agent, Web Search, Vector Search, IT Support, Invoice Processing) built with Agno framework. Contains UI components, FastAPI servers, and Docker deployments.
This is a collection of tool definitions inferred from requirements.txt and Dockerfile references, NOT an actual MCP server implementation. No source code shows explicit tool registration, schema definitions, or parameter validation. All 23 tools appear to be referenced in dependency lists rather than formally defined. The repository structure shows FastAPI agents using the 'agno' framework (v2.0.7), but no MCP-compliant server code is visible. Tool descriptions exist at a high level (1-10 word summaries) but lack the detail required for LLM reasoning. Input schemas are minimal, most tools show only parameter names and types without validation constraints, defaults, or error guidance. This is a collection of AI agent experiments, not a production-grade MCP server.
Tools (23)
anthropic_claude_inferenceread onlyauth42/100
Anthropic Claude LLM inference tool
azure_openai_inferenceread onlyauth43/100
Azure OpenAI inference tool for LLM completions
cerebras_inferenceread onlyauth42/100
Cerebras cloud SDK inference tool
dash_visualizationread only42/100
Dash interactive visualization tool for creating dashboards
docling_document_processorread only40/100
Docling tool for processing and converting documents (PDF, DOCX, etc.)
duckdb_queryread only48/100
DuckDB query tool for local analytical SQL queries
duckduckgo_searchread only50/100
DuckDuckGo search tool for web search functionality
No explicit MCP server implementation visible. All tools are inferred from requirements.txt and Docker files, not from actual MCP server code. Tool registration, schema definitions, and error handling code not present in provided source.
Tool descriptions are minimal (5-15 words) and lack context for LLM selection. None explain WHEN to use the tool vs. similar alternatives, WHAT the tool returns, or prerequisites. Baseline for A+ tools is 50-200 chars with clear intent signaling.
Input schemas show only parameter names and types without validation constraints. No minimum/maximum for numeric params, no enum declarations for categorical params, no pattern constraints for string params. This invites hallucinated values from LLMs.
Recommendations
Implement an actual MCP server using the official Python SDK (mcp package). Register tools with explicit Tool definitions including name, description, inputSchema, and outputSchema fields. The current setup relies on FastAPI agents, not MCP protocol.
Expand tool descriptions to 50-150 characters minimum. Each description must answer: What does this tool do? When would you call it instead of a similar tool? What does it return? Example: 'Search the web using DuckDuckGo. Returns title, snippet, and URL. Use for real-time information unavailable in your knowledge cutoff. Cannot handle images or PDFs.'
Add validation constraints to all input parameters: numeric ranges (min/max), enum values for categorical params, string length limits, and regex patterns. Document constraints in parameter descriptions so LLMs understand what values are valid.
Define and document output schemas for every tool. Include field names, types, and descriptions. Example for duckduckgo_search: { results: [{ title: string, snippet: string, url: string, rank: integer }], total_results: integer, next_page: string | null }.
Add error handling descriptions to all tools. Specify: What errors can occur? (e.g., 'API rate limit exceeded', 'File not found', 'Invalid query format'). What should the agent do? (e.g., 'Retry after 60 seconds', 'Try a shorter query', 'Check file path and permissions').
For inference tools, never expose API keys or prompts as schema fields. Implement server-side secret injection via environment variables. Document that the tool uses credentials from secure storage, not user parameters.
Score history
Overall score trend
↓ 3 points across a rubric change (v1 → v2)
29/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
29
2026-07-28+
v2
2026-03-09
F
32
-
v1
faiss_similarity_search
read only47/100
FAISS tool for efficient similarity search on embeddings
image_processingread only42/100
Image processing tool using OpenCV and Pillow
langchain_ragread only45/100
LangChain Retrieval-Augmented Generation tool for document-based QA
langgraph_workflowread only42/100
LangGraph workflow orchestration tool
newspaper_fetchread only45/100
Newspaper4k tool for fetching and parsing news articles
office_document_conversionread only45/100
Office document conversion tool (DOCX, PPTX to other formats)
ollama_local_inferenceread only45/100
Ollama local LLM inference tool
openai_gpt_inferenceread onlyauth45/100
OpenAI GPT model inference tool
pdf_extractionread only40/100
PDF extraction tool using pdfminer and pypdf
pgvector_similarity_searchread onlyauth50/100
PostgreSQL pgvector tool for semantic similarity search in vectors
qdrant_vector_storeread only47/100
Qdrant vector database tool for storing and retrieving embeddings
sentence_transformer_embeddingread only45/100
Sentence Transformers tool for generating embeddings locally
weaviate_knowledge_graphread onlyauth43/100
Weaviate knowledge graph tool for semantic search
whisper_speech_to_textread only45/100
Faster-Whisper tool for speech-to-text conversion
yfinance_toolread only43/100
Yahoo Finance tool for retrieving stock market data
No documented output schemas visible for any tool. LLMs cannot plan downstream tool calls or extract required data without knowing what fields to expect in responses.
LLM inference tool descriptions (azure_openai_inference, openai_gpt_inference, anthropic_claude_inference, cerebras_inference, ollama_local_inference) expose prompts as direct parameters. No mention of secret injection, token handling, or API key management in tool definitions.
Tool names ending in '_tool' or generic suffixes (e.g., 'yfinance_tool') are ambiguous. Names like 'get_stock_data' or 'search_financial_data' would clarify intent before description reading.
No error handling or recovery guidance documented. LLMs have no way to know if failures are retryable, user-fixable, or fatal. No guidance on what to do if a search returns no results or a file cannot be processed.
Multiple inference tools (anthropic_claude_inference, openai_gpt_inference, cerebras_inference, ollama_local_inference) perform similar functions. No clarity on when to use one vs. another, or how they differ in cost, latency, or model capability.
Vector search tools (pgvector_similarity_search, faiss_similarity_search, qdrant_vector_store) accept 'query_vector' as raw arrays with no guidance on dimensionality, normalization, or supported distance metrics. LLMs cannot infer this without extensive documentation.
Composite tools like 'langchain_rag' and 'langgraph_workflow' accept untyped 'documents' and 'workflow' object parameters. No schema, no field documentation, no guidance on structure. These are too vague for reliable LLM use.
langchain_raglanggraph_workflow
Split composite tools. 'langchain_rag' and 'langgraph_workflow' should be broken into smaller tools: search_documents, retrieve_context, rank_results, execute_step. Each with a single responsibility and clear schemas.
Consolidate inference tool names and descriptions. If all inference tools (Claude, GPT, Cerebras, Ollama) serve the same purpose, use one tool 'infer_with_model' with a model_name enum. If they differ in capability/cost, explicitly document the trade-offs in each description.
For vector search tools, document vector dimensionality, distance metrics supported, and any pre-processing expectations. Example: 'pgvector_similarity_search: Expects 1536-dimensional embeddings from OpenAI's text-embedding-3-small. Supports cosine, L2, and inner product distance. Returns top-k results ranked by similarity score.'
Add natural-identifier support. Tools like yfinance_tool and duckdb_query should accept human-friendly inputs (stock symbols, table names) but also return IDs for chaining. Example: yfinance returns both ticker (user-facing) and ticker_id (internal reference for follow-up tools).
Implement pagination and result limits. Tools returning lists (search_results, vector_results, document_results) must accept limit and offset/cursor parameters and return a total_count or has_more field. Cap default results at 20-50 to avoid context window overload.
Add tool annotations. Mark destructive tools (e.g., any delete or write operations if added in future) with 'destructive: true'. Mark read-only tools with 'readOnly: true'. Mark idempotent operations with 'idempotent: true'. These hints let agents reason about retry safety.
Publish an MCP server endpoint (HTTP or stdio). Current setup is FastAPI-based agents, not MCP. To work with MCP clients, implement the MCP protocol on top of the existing agent logic, or wrap agents as tool implementations behind an MCP server interface.