Dual-backend PBPK Model Context Protocol server for qualification, verification, and reporting
This server has severe definition quality issues across nearly all dimensions. Tool descriptions are generic and lack LLM-optimized guidance. Parameter schemas are not visible in the provided source code, making it impossible to verify type definitions, constraints, or descriptions. The Makefile shows tool paths exist (src/mcp/tools/*.py) but the actual tool definitions are not included in the source excerpt. Output schemas are undocumented. Error handling patterns are absent. The few descriptions provided are under 100 characters and lack 'when to use' context. No evidence of input validation guidance, parameter constraints (enums, ranges), or idempotent operation markers. Security patterns (secret injection, permission gates) are not apparent. Tool naming follows basic verb_noun conventions but descriptions do not disambiguate similar tools (e.g., 'get_job_status' vs 'get_results' vs 'load_simulation' could confuse LLM selection). Without access to actual schemas, this server cannot be scored higher than 40 on definition quality alone.
Discover available PBPK models in the catalog
Export PBPK qualification data in OECD NGRA format
Retrieve the status of a submitted job
Retrieve simulation results
Ingest and validate external PBPK model bundles
Load a PBPK simulation model for analysis
Execute a population-level PBPK simulation
Execute verification and qualification checks on PBPK models
No input schemas visible in provided source. Cannot verify parameter types, required/optional status, constraints (enums, ranges, patterns), or descriptions. Per hard scoring rules, schema score must be 0 for all tools.
Tool descriptions are generic and under 100 characters. Descriptions lack 'when to use' context, prerequisites, or LLM-optimized guidance. Example: 'Discover available PBPK models in the catalog' does not explain when to call this (before load_simulation?) or what structure is returned.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 38 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Validate PBPK model manifest against schema
Validate a simulation request against schema constraints
Similar tools not disambiguated by description. 'get_job_status', 'get_results', and 'load_simulation' could confuse LLM selection. Which is called first? Do they depend on each other? Descriptions do not clarify workflow.
No documented output schemas. LLMs cannot know what fields to expect from tools like 'discover_models' or 'get_results'. This prevents proper chaining and forces agents to guess output structure.
No error handling guidance visible. No recovery steps (e.g., 'if validation fails, call validate_model_manifest first'). No error categorization (retryable vs. user-fixable vs. fatal). No indication of timeouts or partial failure modes.
Destructive tools (ingest_external_pbpk_bundle, run_population_simulation) lack confirmation or dry-run guidance. No indication whether these are idempotent or have side effects. Agents may not know if retry is safe.
No parameter descriptions visible in provided source. Cannot verify if parameters have constraints (enums for model IDs, ranges for limits), examples are absent, or type hints are present.
No pagination or result limits documented. For discovery tools like 'discover_models', agents need to know: how many results? Is pagination required? What is the default limit? Without this, large result sets could exhaust context windows.