MCP Server for aggregating Bible translation resources from Door43
Translation Helps MCP demonstrates weak definition quality across most tools. While 7 of 10 tools have descriptions and explicit schemas are partially visible, the descriptions are often generic, parameters lack depth, and error handling guidance is absent. The server mixes production tools (fetch_scripture, fetch_translation_notes) with three experimental tools marked with 🧪 that return mock data, a significant red flag for production use. Schema details are inferred from test files rather than visible in a canonical tool registration block, limiting confidence in completeness. The few tools with detailed schemas (fetch_scripture) show competent structure, but most lack output schema documentation, parameter validation guidance, and error recovery hints. This is a typical C-grade community MCP server.
🧪 EXPERIMENTAL: AI-powered translation quality assessment (currently returns mock data)
🧪 EXPERIMENTAL: AI-powered content summarization for Bible references (currently returns mock data)
🧪 EXPERIMENTAL: Advanced cache performance analytics and optimization recommendations
Fetch scripture text for a given reference and language
Fetch Translation Academy articles for a given Bible reference
Fetch translation notes for a given Bible reference
Fetch translation questions for a given Bible reference
Three tools (ai_summarize_content, ai_quality_check, smart_recommendations) marked EXPERIMENTAL and return mock data. This violates the principle that tools should be production-ready or clearly gated. Agents cannot reliably use tools that hallucinate responses.
Output schemas are not documented for any tool. LLMs cannot predict what fields to extract from responses, forcing them to reason about unstructured output. LLMs need to know what fields to expect so they can plan downstream tool calls.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | 2024-11-05+ | v1 |
Fetch translation word links for a given Bible reference
Get translation word definition and information
🧪 EXPERIMENTAL: Context-aware resource recommendations based on user role and task
Tool descriptions for fetch_translation_* family are generic (all ~50 chars: 'Fetch translation [notes|questions|word links|academy] for a given Bible reference'). These lack WHEN to use it, WHAT it returns, and dependencies on other tools.
Parameter descriptions lack detail. 'ISO language code (default: en)' and 'Resource organization (default: unfoldingWord')' are minimal. State WHAT the tool does, WHEN to use it, and any prerequisites.' No guidance on valid values, ranges, or error cases.
No error handling guidance. Rubric requires: 'Error responses must tell the LLM what to do next... A raw error code or stack trace gives the agent nothing to act on.' No documented fallbacks, retry strategies, or what to do if a reference is invalid.
Experimental tools (ai_summarize_content, ai_quality_check) explicitly state 'currently returns mock data'. This is acceptable for development but MUST NOT be exposed in production APIs without a feature flag or gating mechanism. Agents will trust and act on mock responses, causing silent failures.
Parameter 'contentType' in ai_summarize_content uses enum with valid values, but missing context on what each enum value returns or costs. No guidance on whether 'all' is recommended or if individual selections are more efficient.
Parameter 'difficulty' in smart_recommendations has enum ['easy','moderate','difficult','auto'] but no explanation of how each affects recommendations or what 'auto' inference is based on. Ambiguous parameter values force agent guessing.
Tool names for experimental tools (ai_summarize_content, ai_quality_check, smart_recommendations) lack clear verb prefixes suggesting state-changing operations.
Tool registration inferred from test files (debug_tool_response.py, test-mcp-tools-direct.mjs) rather than visible in a canonical tool registry.