FastMCP-based tool for writing prompts against data in the NMDC (National Microbiome Data Collaborative) database. Provides access to microbiome data, biosamples, studies, and functional annotations including genomic data objects and PFAM domain analysis.
NMDC MCP defines 16 tools with reasonable naming and documented schemas. Most tools follow verb_noun conventions (get_*, search_*) and include multi-sentence descriptions with context. However, critical gaps exist: (1) NO ERROR HANDLING guidance, tools return raw API results without recovery hints; (2) MINIMAL PARAMETER VALIDATION, numeric ranges (elevation, lat/lon) lack min/max constraints in schema; (3) INCOMPLETE OUTPUT SCHEMAS, most tools show input schemas but do NOT document return types or field structures; (4) INCONSISTENT DESCRIPTION QUALITY, some descriptions are specific and action-oriented (get_samples_by_annotation, fetch_and_filter_gff_by_pfam_domains) while others are generic (get_collection_stats, get_collection_names); (5) COMPOSITION ISSUES, tools like get_entity_by_id_with_projection and get_entities_by_ids_with_projection are nearly identical, forcing LLM disambiguation. On positives: naming is consistently clear, parameter descriptions exist for most inputs, and the schema structure is valid JSON Schema. The Makefile reveals test coverage and integration patterns, suggesting operational maturity, but the tool definitions themselves lack the rigor expected for production-grade agent integration.
Use this tool to analyze GFF files for PFAM domains and show genomic locations. Use AFTER get_samples_by_annotation to get the actual gene coordinates and genomic context. Takes a data_object_id from the previous search results and downloads the entire GFF file by default to show specific gene annotations containing the PFAM domains. Only use sample_bytes parameter if you specifically need to limit download size.
Use this tool to get lists of IDs from NMDC collections in batches. Useful for sampling or systematic analysis of large datasets.
Use this tool to find all biosamples associated with a specific study. Returns a list of biosample IDs and metadata.
Use this tool to discover what types of data are available in the NMDC database. Returns a list of collection names like 'biosample_set', 'study_set', etc.
Use this tool to get statistics about NMDC collections including document counts. Helps understand the size and scope of available data.
NO OUTPUT SCHEMAS DOCUMENTED, LLMs cannot know the structure of returned data. All 16 tools lack documented return types, field names, and nesting. Forces LLMs to guess structure and plan downstream calls blindly.
NO ERROR HANDLING GUIDANCE, tools do not specify what errors can occur or how LLMs should recover. No messages like 'Entity not found, use get_collection_ids to discover valid IDs' or 'Request timeout, try reducing sample_bytes'. Raw API failures leave agents without recovery paths.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Use this tool to find biosamples that contain specific PFAM protein domains. Returns structured data about biosamples, their activities, and data objects. Perfect for finding samples with particular functional capabilities.
Use this tool to retrieve specific fields from multiple NMDC entities by their IDs. Efficient for batch operations with field filtering.
Use this tool to retrieve any NMDC entity by its ID. Works with biosamples, studies, data objects, and other NMDC entities.
Use this tool to retrieve specific fields from an NMDC entity by ID. More efficient than getting full entities when you only need certain fields.
Use this tool to find biosamples with specific functional annotations. Returns COMPLETE biosample records including all data objects (GFF files, protein files, etc.) with their IDs and URLs. ALWAYS set max_records to match the user's request (e.g., if they ask for '1 sample' or 'a sample', set max_records=1). Use max_records, NOT limit, to control how many samples to return. Required formats: PFAM domains use 'PFAM:PF04183', KEGG use 'KEGG.ORTHOLOGY:K00001', COG use 'COG:COG0001', GO use 'GO:GO0000001'. When users want genomic locations of domains, use this first to find samples, then use fetch_and_filter_gff_by_pfam_domains with a GFF data_object_id from the results.
Use this tool to find biosamples from specific ecosystem types, categories, or subtypes. Perfect for studying particular environments like soil, marine, or host-associated microbiomes.
Use this tool to find biosamples collected within a specific elevation range. Perfect for studying altitude-related microbial communities.
Use this tool to find biosamples collected within a specific geographic bounding box defined by latitude and longitude coordinates.
Use this tool to get detailed DOI information for a study, including publication DOIs, dataset DOIs, and award DOIs.
Use this tool to find the study associated with a specific biosample. Returns study information and metadata.
Use this tool to search for studies based on DOI criteria like provider, category, or DOI value patterns. Perfect for finding published studies or datasets from specific sources.
MISSING NUMERIC CONSTRAINTS, elevation (get_samples_in_elevation_range) and lat/lon (get_samples_within_lat_lon_bounding_box) parameters lack min/max bounds. LLM can pass invalid values (lat=200, elevation=-10000). Should enforce realistic ranges: lat ±90, lon ±180, elevation -500 to 8500 (typical Earth range).
MISSING ENUMS FOR KNOWN VALUES, ecosystem_category/type/subtype (get_samples_by_ecosystem), doi_provider/category (search_studies_by_doi_criteria), and annotation prefix formats (get_samples_by_annotation) should be enums, not free-form strings. Current state invites hallucinated values like 'swamp' if not in valid set.
REDUNDANT TOOL NAMES, get_entity_by_id_with_projection and get_entities_by_ids_with_projection are nearly identical (single vs batch) but names are long and do not clearly signal the distinction. Consider shorter aliases or a single parameterized tool (entity_ids accepting 1 or N). Forces LLM to reason about which to call.
GENERIC DESCRIPTIONS, get_collection_names ('discover what types of data are available'), get_collection_stats ('get statistics about NMDC collections'), get_all_collection_ids ('get lists of IDs from NMDC collections') are vague and under 50 characters. Do not explain WHEN to call them or how results enable downstream decisions. Should be: 'Call first to list available collection types (biosample_set, study_set, etc.). Use collection name in get_all_collection_ids() to sample data.'
AMBIGUOUS PROJECTION PARAMETER, get_entity_by_id_with_projection and get_entities_by_ids_with_projection accept projection as 'array | string' (comma-separated). Description does not clarify format preferences, whether field order matters, or what happens if invalid fields are requested. Specify: 'Array of field names, e.g. ["id", "name", "url"]. Or comma-separated string: "id,name,url". Invalid fields are silently omitted.'
MISSING PAGINATION GUIDANCE, tools returning lists (get_all_collection_ids, get_biosamples_for_study, search_studies_by_doi_criteria) accept batch_size but do not specify default, min, or max. Should state: 'batch_size default=20, min=1, max=1000. For next batch, use offset from last response or cursor parameter if provided.' No documented way to iterate through large result sets.