MCP server for AIND metadata access and retrieval from MongoDB database and NWB files
The server provides 6 tools for MongoDB data access and NWB file handling. Most tools have descriptions and basic schemas, but quality is inconsistent. Tool naming is clear (verb-first: get_, identify_), but parameter descriptions vary significantly in completeness. No output schemas are documented. Error handling returns raw exception messages rather than actionable guidance. The server reads from external databases and file systems but lacks input validation, type constraints, and recovery instructions.
Executes a MongoDB aggregation pipeline for complex data transformations and analysis.
Retrieves number of documents from MongoDB database using a simple MongoDB filter
Retrieves documents from MongoDB database using simple filters and projections.
Get an LLM-generated summary for a data asset, based on the _id field
Searches the /data directory in a code ocean repository for a folder and subfolder containing the subject_id and date, and loads the corresponding NWB file.
Identifies NWB folder in the given S3 link and opens it as a NWBZarrIO object.
No documented output schemas for any tool. LLMs cannot plan downstream operations without knowing what fields are returned.
Error responses return raw exception messages (e.g. 'An exception of type {type} occurred. Arguments: {ex.args}') instead of actionable recovery guidance. LLMs cannot determine whether to retry, ask the user, or abandon the plan.
Parameter 'filter' and 'projection' in get_records, and 'agg_pipeline' in aggregation_retrieval accept arbitrary dicts with no validation or constraint. LLMs cannot infer valid MongoDB syntax without examples or schema validation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
get_summary has only a 27-character description ('Get an LLM-generated summary for a data asset, based on the _id field'). Does not explain WHEN to call it vs get_records, what context it requires, or what the summary contains.
The 'limit' parameter in get_records defaults to 5 but the description says 'try to not exceed 100'. This is confusing, is 5 a safe default, or should agents request larger limits? The constraint is soft guidance, not a hard enum.
No pagination support documented. If get_records returns 100+ items, the response will bloat the context window. No mention of total_count, next_cursor, or offset-based pagination.
get_records includes a docstring explaining WHEN NOT to use the tool, but no corresponding guidance on when TO use it (beyond the docstring). Tool selection would benefit from explicit use-case triggers in the description field visible to schema consumers.