Multi-tool MCP server integrating weather, database, and math tools with LangChain
This server has 6 tools across 3 separate FastMCP instances. Tool definitions are present with basic schemas and descriptions, but quality is inconsistent. Parameter descriptions are often missing or minimal. One tool (mathserver.py's 'multiple') uses a verb that is not an action verb. The math server runs on STDIO, which violates remoteness principles. Output schemas are not documented. Error handling is absent, no recovery guidance, no categorization, no invalid value feedback. Security issues: hardcoded DB credentials in plaintext code, no input validation, no rate limiting. Tool composition is reasonable (separate tools for distinct operations) but lacks integration details.
Add 2 numbers
Fetch all students
Fetch the courses with their respective teachers
Fetch courses for a student by name
Get the weather location
Multiply 2 numbers
Critical: Tool descriptions far below rubric baseline. Descriptions are 10-25 characters; rubric requires 10-1024 chars with meaningful content. Examples: 'Get the weather location', 'Fetch all students', 'Multiply 2 numbers' do not explain WHEN to use, WHAT is returned, or dependencies. Scoring rule: descriptions under 20 chars → max 0-20 points.
Critical: Parameter descriptions missing or absent. JSON schemas show input types but no descriptions for 'location', 'course_name', 'student_name', 'a', 'b'. Rubric rule: every parameter MUST have a non-empty description. LLMs cannot infer meaning from parameter names alone.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Critical: No output schema documentation. Tools return structured JSON (e.g., weather returns city/country/temp/feels_like/description; database tools return rows) but output structure is not documented in tool definitions. Agents cannot plan downstream calls or extract fields without knowing return format.
Critical: No error handling or recovery guidance. Code returns generic responses (e.g., 'Sorry, I couldn't fetch the weather for {location}' in WeatherServer.py) with no actionable next steps. Database tools have no try/catch; connection failures will crash with no guidance. Rubric requires: error classification (retryable/user-fixable/fatal), recovery suggestions, invalid value details.
Critical: Hardcoded database credentials. db_Server.py lines 18-22 hardcode DB_PASSWORD='123456', DB_HOST='localhost', DB_USER='postgres'. These should be injected via environment variables or vault per pattern:secret-injection. Credentials in source code leak to git history and logs.
Medium: Verb naming issue. Tool 'multiple' should be 'multiply' to follow verb_noun convention. Baseline: 90% of A+ tools start with action verb. 'multiple' is an adjective, not a verb.
Medium: No input validation or constraint documentation. Parameters accept free-form strings (location, course_name, student_name) with no enums, patterns, min/max lengths, or format guidance. Agents can pass invalid/empty strings. Rubric: declare enums for known sets; specify format/range in descriptions.
Medium: No pagination or result limiting. Database tools (get_all_students_name, get_courses) fetch all rows from database with no LIMIT, offset, or cursor. A large student table could return thousands of rows, blowin context window. Rubric: accept page/limit, return total count.
Medium: No permission/scope declarations. Tools access databases and external APIs with no documented permission requirements (e.g., 'read:students', 'read:weather'). Agents have no visibility into what authority is required. Rubric: each tool should declare required permissions.
Low: Inconsistent async/sync implementations. WeatherServer and db_Server use async (async def fetch_weather, async def get_connection) while mathserver uses sync (def add, def multiple). Mixing async/sync in a single agent can cause event loop issues.