A multi-interface (REST and MCP) server for natural language inference with support for multiple backend providers (Ollama, HuggingFace, OpenRouter)
Omni-NLI provides 2 tools with clear verb-based naming and documented schemas. Both tools have functional descriptions and complete input parameter schemas with types and descriptions. However, output schemas are not explicitly documented, parameter descriptions could be more detailed with constraints, and error handling patterns are absent. The server follows basic tool design patterns but lacks production-grade robustness. Tool names are appropriately specific (evaluate_nli, list_providers) and start with action verbs. Input schemas are well-formed JSON Schema with proper type declarations and field descriptions. The main gaps are: (1) no documented output schemas visible in the source, (2) limited error recovery guidance, (3) missing enum constraints for some parameters that could validate early, and (4) no indication of result pagination or scaling behavior for the analyze tool.
Analyzes a premise and hypothesis to determine their logical relationship: entailment (hypothesis follows from premise), contradiction (hypothesis conflicts with premise), or neutral (neither).
Lists available NLI backend providers (Ollama, HuggingFace, or OpenRouter) and their configuration status.
Output schemas not documented. LLMs cannot plan downstream tool calls or know what fields to expect from tool responses. Without documented return types, agents must guess the structure of results.
Missing error handling guidance. No recovery patterns visible, errors should explain what to do next (retry, call a discovery tool, adjust parameters). Raw error codes are useless to LLMs.
Backend and model parameters could use enum constraints. 'backend' explicitly has enum ['ollama','huggingface','openrouter'], which is good. However, 'model' is a free-form string that LLMs could hallucinate. Consider constraining valid models per backend or providing discovery guidance.
No idempotency or confirmation pattern for evaluate_nli. If the tool makes external API calls, retries could cause duplicate analysis attempts. State whether the tool is idempotent or add dry-run capability.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 68 | - | v1 |
Parameter descriptions lack specifics on constraints and ranges. For example, 'premise' and 'hypothesis' have minLength/maxLength in schema (1/4096) but the description doesn't reinforce this or explain what content is acceptable. Descriptions should state: 'Natural language statement (1-4096 chars)'.