MCP server for pharmaceutical compound similarity analysis using PCA embeddings and MOA-based queries
The server provides 4 well-intentioned tools with full input schemas and descriptions, but falls short of production-grade quality in several areas. Tool naming is clear and verb-driven (query_, get_, list_). Descriptions are present but vary in depth, some are detailed (query_compound_similarity at ~180 chars) while others are minimal (get_compound_details at ~55 chars). Input schemas are properly typed with JSON Schema, but parameter descriptions lack specificity about ranges, formats, and constraints. Output schemas are documented via dataclass definitions but are not explicitly returned in tool metadata. Error handling is basic, missing compounds return a dict with an error_message field, which is good, but there's no guidance on recovery or alternative actions. The server lacks tool annotations (readOnlyHint, destructiveHint) that would help LLMs understand tool safety.
Look up detailed metadata for a single compound.
List compound names in the dataset with optional substring filtering.
Global embedding similarity: top-N compounds most similar to the given compound by PCA cosine similarity with **no filters or scoping**. Use this for all-compound nearest neighbors; use compound_neighbors for scoped neighbors.
MOA-centric analysis: summarize intra-class similarity and return the closest compounds that sit **outside** the MOA (via centroid embedding similarity). Use this when the intent is "find near neighbors not in this MOA."
Parameter descriptions lack specificity on ranges, formats, and valid values. E.g., top_n accepts an integer with no stated min/max; compound_name accepts any string with no guidance on partial vs. full match behavior. This invites LLMs to pass unreasonable values (top_n=999999, compound_name with special chars).
Output schemas are documented via dataclass/asdict but are not formally declared in tool metadata. The MCP spec expects tool definitions to include an explicit 'result' schema. Without this, LLMs cannot predict what fields to expect, forcing ad-hoc parsing of responses.
Missing tool annotations. All four tools are read-only but lack the readOnlyHint annotation. This forces LLMs to reason about tool safety without explicit guidance, increasing the risk of misuse.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2026-07-28+ | v2 |
get_compound_details description is only 55 characters and lacks context on when to use it vs. query_compound_similarity. The rubric baseline is 194 chars average; this is significantly under-documented. An LLM cannot distinguish the intent of this tool from the longer query_compound_similarity without more detail.
Error handling is basic. When a compound is not found, the response includes an error_message field but provides no recovery guidance. E.g., 'Compound "XYZ" not found. Try list_available_compounds() to browse the dataset.' would be more actionable than the current message.
list_available_compounds returns an unbounded number of results (controlled only by 'limit' default=50). The description states 'Maximum number of names to return' but the limit has no stated min/max. If an LLM passes limit=100000, the response could exhaust the context window.