MCP server for datacenter GPU liquid cooling thermal analysis
This is a specialized scientific/engineering MCP server for thermal analysis. Tool definitions are well-structured with complete JSON schemas, good parameter descriptions, and clear domain-specific documentation. All 5 tools have explicit registration with @mcp.tool decorators, proper input validation via Pydantic, and structured output via model_dump(). However, descriptions are verbose (500-1000+ chars), error handling lacks recovery guidance, and there is no documentation of output schemas for the LLM. The server implements domain logic correctly but falls short of LLM-optimized tool design patterns. STDIO transport is a hard constraint limiting remote accessibility.
Calculate junction temperature, thermal resistances, and pressure drop for a liquid-cooled cold plate. Uses a 1D thermal resistance network (junction -> case -> TIM -> base -> convection) with Dittus-Boelert convection and Darcy-Weisbach pressure drop. Supports water and 50/50 glycol coolants. Returns model-applicability notices and low-Reynolds-number warnings. Temperature acceptance requires a caller's component-specific limit or the decision-report tool. Set sensitivity=True to include finite-difference partial derivatives and illustrativperturbations: ∂Tj/∂Q, ∂Tj/∂R_tim, ±20% R_jc variation, and Tj rise when R_tim is doubled. These are not measured uncertainty or lifetime estimates.
Rack-level thermal analysis for N identical GPU cold plates. Models steady-state heat removal across a full rack using either series or parallel plumbing topology. Series: coolant flows through each cold plate in sequence. Each GPU's inlet temperature equals the previous GPU's outlet. Pressure drop accumulates (ΔP_total = N × ΔP_per_plate). Hottest GPU is always the last in the chain. Parallel: coolant splits equally to all cold plates. All GPUs share the same inlet temperature. Flow per GPU = total_flow_lpm / gpu_count. System ΔP equals per-plate ΔP (not cumulative). Assumptions: identical GPUs, uniform flow distribution, no manifold losses. Ambient temperature is optional; if omitted, rack analysis defaults ambient reference to cdu_supply_temp_c.
Compare thermal and hydraulic performance of water vs 50/50 glycol under identical conditions. Runs analyze_coldplate for each coolant and returns side-by-side results including junction temperature, pressure drop, and pump power for each.
Tool descriptions are excessively verbose (500-1200 chars), exceeding the 194-char baseline for production tools. Lengthy descriptions waste tokens, dilute signal, and force LLMs to parse domain detail that belongs in inline parameter docs or error messages. Example: analyze_coldplate description is 550+ chars explaining sensitivity analysis, thermal networks, and model applicability, this belongs in parameter descriptions or error guidance, not the tool summary.
Output schemas are not documented. Tools return model_dump() output (e.g. AnalyzeColdplateOutput, OptimizeFlowRateOutput) but the LLM has no declared schema showing what fields to expect, their types, and their meaning. This forces the LLM to infer structure from test calls or assume fields exist, risking downstream tool chaining failures. Example: optimize_flow_rate returns 'analysis_at_minimum_flow' with nested structure, but no schema doc declares it's optional or shows the nested field types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Generate a decision-quality thermal screening report for a GPU deployment scenario.
Find the minimum coolant flow rate that keeps junction temperature at or below a target. Uses binary search between flow_min_lpm and flow_max_lpm. Returns the minimum flow rate, whether the target was met, and the full thermal analysis at that operating point. margin_c: Optional safety margin in °C. The optimizer targets (max_junction_temp_c - margin_c) as the effective ceiling. Choose the guardband from component and operating evidence; the model does not establish a universal margin or service-life forecast.
generate_decision_report has a minimal description ('Generate a decision-quality thermal screening report for a GPU deployment scenario.') that does not explain WHEN to call it vs other tools, what a 'decision-quality' report includes, or what the LLM should do with the output. This violates the pattern requirement: 'A tool description must answer: What does it do? When should the LLM call it instead of a similar tool? What does it return?'
Error handling provides no recovery guidance. The implementations catch ValidationError and return {'error': exc.errors()}, a raw Pydantic error structure the LLM cannot parse. Example: if an LLM passes an invalid coolant, it receives a ValidationError JSON dump with field names and constraint violations, but no actionable message like 'Invalid coolant: must be water or glycol50. Please retry with a valid option.' This violates pattern:recovery-guide.
Parameters with constrained values (coolant, topology) use enums in the schema, which is correct; however, parameter descriptions do not explicitly state the valid options inline for easy LLM reference. Example: coolant parameter says 'water or glycol50' in the description, but enumerating the full list ('Coolant type. Valid options: water, glycol50') would reinforce the constraint and reduce hallucination risk.
Missing parameter descriptions: geometry parameter is described as 'Cold plate geometry parameters (optional)' with no detail on what sub-fields it accepts (e.g. length_mm, width_mm, channel_diameter_mm are likely expected, but not documented). This forces the LLM to guess the structure. The Geometry Pydantic model is defined in schemas.py but not exposed in the tool docs.
Tool composition: generate_decision_report appears to duplicate functionality (runs analyze_rack or analyze_coldplate + optimize_flow internally). The LLM does not know when to call generate_decision_report vs. composing the base tools. No description clarifies 'This is a high-level convenience tool that wraps multiple analyses, call this for end-to-end thermal screening; call analyze_coldplate for single-point analysis.' This risks decision-tree confusion.
No dry-run or confirmation mechanism for high-stakes outputs (e.g., generate_decision_report generates a report that may influence production deployment decisions). While these are read-only tools, the absence of a 'validate this scenario before I apply it' step means the LLM cannot ask the user to confirm assumptions before committing to analysis. Not critical but good practice for decision-support tools.