MCP server with 8 tools for NGSS standards: complete 3D framework filtering (SEP, CCC, DCI) and intelligent unit planning. Supports Smithery.ai deployment.
NGSS MCP Server demonstrates solid definition quality with consistent naming, comprehensive parameter schemas, and good descriptions across all 8 tools. All tools follow verb_noun naming convention (get_, search_, filter_). Input schemas are complete with proper type definitions, enums, and constraints. Descriptions are generally adequate (100-200 chars), though some lack pedagogical context on WHEN to use each tool vs. others. Output schemas are not explicitly documented in the source code, which is a notable gap. Error handling is present but minimal (generic error responses without recovery guidance). All parameters have descriptions and proper constraints. The server targets an educational domain (NGSS standards) with well-defined enumerated values (SEP, CCC, DCI), which naturally constrains the input space. Minor issue: detail_level parameter is repeated across 7 of 8 tools and could be extracted as a shared pattern to reduce duplication and improve consistency.
Filter NGSS standards by a specific Crosscutting Concept (CCC)
Filter NGSS standards by a specific Disciplinary Core Idea (DCI) code
Filter NGSS standards by a specific Science and Engineering Practice (SEP)
Find standards that share common 3D framework components (SEP, CCC, DCI) with a given anchor standard for coherent unit planning
Extract the three-dimensional learning components (SEP: Science and Engineering Practices, DCI: Disciplinary Core Ideas, CCC: Crosscutting Concepts) for a specific standard (e.g., MS-PS1-1, MS-LS2-3, MS-ESS3-1)
Retrieve a specific NGSS standard by its code identifier (e.g., MS-PS1-1, MS-LS2-3, MS-ESS3-1)
Find all NGSS standards in a specific science domain (Physical Science, Life Science, or Earth and Space Science)
Output schemas not documented. Tool descriptions state what is RETURNED (e.g., 'Response detail level: minimal, summary, full') but the actual response structure is not formally defined in the schema or description. LLMs cannot reliably extract downstream fields without knowing the response shape.
Error handling is generic and lacks recovery guidance. Error responses return JSON with error/message/code fields, but do not guide the LLM on what to do next (e.g., 'Standard not found. Try search_standards() with keywords instead.'). Per pattern:recovery-guide, errors must tell the agent what to do next.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Perform full-text search across all NGSS standard content including performance expectations, topics, and keywords (e.g., "energy transfer", "ecosystems", "chemical reactions", "climate change")
Descriptions do not clarify WHEN to use each tool vs. similar tools. E.g., search_standards, filter_by_sep, filter_by_ccc, and filter_by_dci all retrieve standards but along different dimensions. LLMs need guidance on which tool to call for a given user intent. Currently, a user saying 'show me all energy standards' could be interpreted as search_standards(query='energy') or filter_by_ccc('Energy and Matter'), and the descriptions do not disambiguate.
detail_level parameter is duplicated across 7 of 8 tools. This increases cognitive load on LLMs and the risk of parameter value mismatches. Consider extracting as a global server-level setting or documented shared pattern to improve consistency.
Pagination design could be clearer. Tools use offset/limit, but do not document whether results are sorted consistently or what happens at boundaries (e.g., if limit=10 and only 5 results remain, is that an error?). Add clarification: 'Results are sorted by [field]; if fewer than limit results are available, all are returned and the LLM can assume no more records exist beyond the limit.'
find_compatible_standards returns a compatibility_score but the breakdown of that score (domain_match, shared_seps, shared_cccs, shared_dcis) is computed in code but not documented as part of the response schema. LLMs cannot predict what data is available for downstream reasoning.