Provides documentation search and examples for Babylon.js development
Babylon MCP Server provides 6 well-structured documentation and source-code search tools for Babylon.js developers. All tools have descriptions, clear naming conventions (all verb_noun format: search_*, get_*), and explicit input schemas with typed parameters. However, there are significant gaps in parameter descriptions, output schema documentation, and error handling guidance. Most tool descriptions are adequate (100-150 chars) but lack depth about WHEN to use vs alternatives. Parameters like 'category', 'package', and 'filePath' have minimal descriptions. No output schemas are documented in the source code, making it impossible for LLMs to understand return structure and plan downstream calls. Error handling is absent, no recovery guidance, no categorization of retryable vs fatal errors. All tools are READ_ONLY (good for safety), but security and audit context is missing. The server demonstrates solid foundational tool design but lacks the depth expected of production-grade agent tooling.
Retrieve full content of a specific Babylon.js documentation page
Retrieve full Babylon.js source code file or specific line range
Search Babylon.js API documentation (classes, methods, properties)
Search Babylon.js documentation for API references, guides, and tutorials
Search Babylon.js Editor documentation for tool usage, workflows, and features
Search Babylon.js source code files
Output schemas are not documented. LLMs cannot predict return structure, field names, or types (e.g., search_babylon_docs returns what fields? An array of objects with 'title', 'url', 'snippet'?). Without documented schemas, agents cannot reliably chain tools or extract required IDs for downstream calls.
Parameter descriptions are minimal or missing for 'category' (search_babylon_docs, search_babylon_api, search_babylon_editor_docs), 'package' (search_babylon_source), 'startLine'/'endLine' (get_babylon_source). LLMs cannot infer valid values, ranges, or format expectations. E.g., for 'category', is it 'api' or 'api_reference'? Is it case-sensitive? Enum?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 58 | - | v1 |
No error handling guidance. If search_babylon_docs returns no results, what should the LLM do next? Try a broader query? Call a different tool? If get_babylon_doc fails with 'path not found', should it retry or call search_babylon_docs to discover the correct path? Recovery paths are undocumented.
No tool distinguishes between similar search operations. 'search_babylon_docs', 'search_babylon_api', and 'search_babylon_source' have overlapping purposes (all search Babylon.js knowledge). Descriptions do not clearly explain WHEN to call search_babylon_docs vs search_babylon_api. LLMs will waste reasoning cycles choosing between them or call all three redundantly.
get_babylon_doc and get_babylon_source accept opaque identifiers ('path', 'filePath') without guidance on format. Should 'path' be 'scene-setup' or 'en/features/scene-setup'? Should 'filePath' be 'scene.ts' or 'packages/dev/core/src/scene.ts'? Users (and LLMs) will not have these IDs naturally, they must be discovered via search tools first. Documentation should state: 'Use search_babylon_docs to discover the path value before calling this tool.'
No pagination guidance documented. search_babylon_docs, search_babylon_api, search_babylon_editor_docs accept 'limit' (default 5), but no 'offset' or 'cursor' for pagination. If an LLM needs all results, can it call with limit=1000? Are results sorted consistently? Will repeated calls with limit=5 fetch new results or the same top 5? Undocumented pagination invites misuse.
Tool descriptions lack WHEN context. E.g., search_babylon_docs says it searches 'API references, guides, and tutorials' but does not explain: Should I call this for high-level concepts or low-level API details? Is this for Babylon.js beginners or advanced users? When would I pick this over search_babylon_api? LLMs need explicit decision criteria to avoid wrong tool selection.