Exposes STRING database functionality as an MCP server. Provides tools for resolving protein identifiers, retrieving interactions, functional annotations, enrichment analysis, and generating network visualizations.
The STRING MCP server defines 3 tools with reasonable naming (verb prefixes: string_resolve, string_interactions, string_help), but suffers from incomplete schemas, missing output documentation, and inadequate error guidance. Tool 1 (string_resolve_proteins) has the most complete input schema with typed parameters and descriptions. Tool 2 (string_interactions_query_set) has no visible input schema definition in the provided code. Tool 3 (string_help) has a properly constrained enum parameter. However, NONE of the tools include output schemas or field documentation, forcing LLMs to infer response structure. Error handling is present in the implementation (timeout, HTTP status errors, hints) but not formally documented as part of tool contracts. Descriptions are moderately detailed (70-150 chars) but lack clarity on return types and dependencies. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are declared.
Provides help documentation on STRING topics
Get interactions within query set
Maps one or more protein identifiers to their corresponding STRING metadata, including: gene symbol, description, sequence, domains, species, and internal STRING ID. This method is useful for translating raw identifiers into readable, annotated protein entries. Example input: "TP53%0dSMO"
Tool 2 (string_interactions_query_set) has an extremely minimal description ('Get interactions within query set') that does not explain when to use it, what it returns, prerequisites, or how it differs from other interaction-related tools. Violates pattern:tool-description requirement of 10-1024 chars with actionable context.
No input schema visible for string_interactions_query_set. The tool name and description imply it accepts protein identifiers or a query set, but no parameters are documented. This blocks LLMs from knowing how to invoke it or what constraints apply.
No output schemas documented for ANY tool. LLMs cannot know what fields to expect from responses (e.g., does string_resolve_proteins return protein objects with {id, symbol, description, sequence, domains, species} or something else?). No downstream tool chaining is possible without guessing field names. Violates pattern:tool (document output schema) and mxe:include-chaining-ids.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 27 | - | v1 |
string_resolve_proteins description mentions 'Example input: TP53%0dSMO', a concrete example value. Remove the example and use parameter constraints (enum/pattern) instead. The description currently violates best practice by embedding specific input strings.
Error handling (timeout, HTTP 400/404) is implemented in server.py with hints and diagnostics, but not formally documented in tool descriptions. LLMs invoking these tools have no published contract for what errors to expect, how to interpret them, or what recovery steps are available. Violates pattern:recovery-guide and pattern:error-classification.
No tool annotations declared (readOnlyHint, destructiveHint, idempotentHint). All three tools appear to be read-only and idempotent (Risk field says 'READ_ONLY'), but this is not communicated to the LLM via formal tool metadata. Agents cannot reliably determine which tools are safe to call repeatedly or if calling a tool will mutate state.
string_help tool accepts an enum of 11 topics, which is good, but the description 'Provides help documentation on STRING topics' is generic and does not explain when to call it instead of other discovery or documentation mechanisms. Does not state whether results are text, markdown, or structured. Lacks dependency hints.
Parameter 'show_sequence' in string_resolve_proteins is defined as enum ["0", "1"] (string values) with description 'Use only if user requests sequence data.' This mismatch is awkward, boolean flags should be type 'boolean' with enum [true, false] or default false. The string enum with advisory description forces LLMs to reason about type coercion.