A collection of MCP (Model Context Protocol) servers including Math, PII Handler, Psychologist, Weather, and Terminal servers with LangChain integration
This repository contains 8 tools across multiple servers with highly inconsistent quality. The math tools (add, multiply) have minimal but adequate definitions. The PII server tools are well-documented but expose critical security risks. The psychologist tool has an overly complex nested schema that would confuse LLMs. The weather and shell tools have minimal documentation. Overall: naming is decent (verb_noun pattern mostly followed), but parameter descriptions are sparse, output schemas are undocumented, and error handling is absent. Security practices are poor, no validation guidance, no recovery instructions. Most tools lack the LLM-optimized descriptions (50-200 chars) that production tools require.
Add two integers together.
Conversational animal personality quiz tool for MCP/Claude integration that guides users through three animal selection questions and provides psychological analysis using LLM.
Clear stored PII mappings from the database for a specific session or all sessions
Get weather information for a specified location.
Multiply two integers together.
Process text to detect, mask, and handle personally identifiable information (PII) including credit cards, SSNs, emails, phone numbers, names, addresses, and city/state/zip combinations
Restore original PII values from masked text using stored mappings
animal_personality_conversation has deeply nested, overly complex input schema with object-in-object-in-object structure. LLMs struggle to construct such nested payloads correctly. Schema should flatten to top-level parameters or use simpler structured input.
get_weather has NO documented output schema. LLMs cannot plan downstream tool calls or extract the right fields from results. Must document: fields returned (temp, conditions, forecast), data types, and structure.
run_command is a destructive tool (SHELL execution) with minimal validation guidance and no error recovery instructions. Description does not warn about injection risks or state that LLMs should sanitize input. No mention of timeout, output limits, or what happens on failure.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Execute a terminal command and return its output.
PII tools (process_with_pii, restore_pii, clear_pii_mappings) store and retrieve sensitive data via session_id parameter, but descriptions do NOT explain the risks of session_id exposure, collision handling, or data retention policy. No guidance on how an LLM should manage session IDs securely.
add and multiply have trivial descriptions (under 30 chars). While adequate for toy examples, production tools need 50-200 char descriptions explaining context, when to use, and what to expect. E.g. 'Returns integer sum. Use for numeric calculations only, does not support floating-point values or overflow checking.'
No tool has error handling guidance. Descriptions do not explain what happens on invalid input (e.g. non-integer values for add/multiply, invalid location for get_weather, bad shell commands). LLMs need explicit recovery paths: 'If command fails, check syntax with validate_shell_command() first.'
get_weather location parameter has no format constraints. Description does not specify: is this a city name, lat/lon, postal code, or coordinate string? What happens if ambiguous (e.g. 'Springfield' exists in multiple states)? LLMs will guess and fail.
run_command has no documented output schema or example return value. Does it return stdout, stderr, exit code, all three? Are multiline outputs formatted as a single string or array of lines? LLMs cannot parse unknown output formats reliably.
animal_personality_conversation description references 'three animal selection questions' but the schema shows first/second/third are optional (null is allowed). Unclear if all three must be answered, whether order matters, or how null values are handled in the psychological analysis.