Multi-core Tokio-native orchestration for LLM inference pipelines: dedup, circuit breakers, backpressure, MCP, and autonomous self-improvement
Server provides 4 tools with basic descriptions and parameter schemas, but significant gaps in schema completeness, parameter descriptions, and error handling reduce quality. Tool definitions are visible in source code (src/bin/mcp.rs, src/tui/app.rs), so they are not inferred. Naming follows action-verb conventions (infer, batch_infer, pipeline_status, configure_pipeline). Descriptions are present but brief (10-80 chars), lacking WHEN/WHY context and action guidance needed for LLM tool selection. Input schemas show parameter names and basic types but lack rich descriptions for most parameters, particularly the configure_pipeline tool which accepts critical infrastructure parameters (circuit_breaker_threshold, rate_limit_rps) with minimal validation hints. No documented output schemas visible. No pagination support shown for batch_infer results. Error handling and recovery guidance are absent from descriptions.
Submit multiple prompts in a batch for non-blocking parallel processing through the pipeline
Modify pipeline configuration: select worker backend, set retry attempts, circuit breaker thresholds, and rate limits
Send a single prompt through the inference pipeline and receive tokenized output
Query real-time pipeline metrics: request throughput, deduplication savings, error rates, circuit breaker states, and per-stage latencies
configure_pipeline lacks parameter constraint documentation. Integer parameters (retry_attempts, circuit_breaker_threshold, rate_limit_rps) have type hints but no explicit min/max ranges or validation rules in descriptions. LLMs cannot infer that retry_attempts must be 0-10 without reading schema; the description should state this explicitly as '0-10 inclusive' or 'positive integer'.
Tool descriptions are too brief (40-80 chars) and lack actionable context. 'Submit multiple prompts in a batch for non-blocking parallel processing' does not explain WHEN to use batch_infer vs infer, whether results are eventually-consistent, retry behavior on failure, or how to track status. Descriptions should be 50-200 chars with dependency hints and use cases per pattern:tool-description baseline.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 37 | - | v1 |
infer, batch_infer, and configure_pipeline have input parameters with missing or minimal descriptions. 'session_id: Optional session identifier for grouping related requests' is vague, does session_id enable cross-request deduplication, affect rate limiting, or enable session replay? How is it generated? Required or auto-generated? Parameter descriptions must be explicit per pattern:tool-description.
No documented output schemas for any tool. LLMs need to know what fields infer returns (token_id, output_text, latency?), what batch_infer returns (job_id, queue_position?), what pipeline_status returns (throughput, latency percentiles, error_rate values and units), and what configure_pipeline returns (confirmation, new settings, validation warnings). Without output schemas, agents cannot plan downstream calls or extract the correct fields. See pattern:response-shaper.
configure_pipeline description does not indicate it is a destructive/configuration-altering operation. If rate_limit_rps is reduced mid-request, does it affect queued requests or only new ones? Can changes be undone? Description must state 'This tool modifies pipeline configuration, all changes take effect immediately on the running pipeline' to signal non-idempotent behavior.
batch_infer accepts 'prompts' as an array but provides no guidance on result handling. Does it return results as a list in the same order? Does it return a job_id for async polling? What is the maximum batch size? Are failed prompts retried? Lack of clarity forces agents to guess and makes error recovery impossible.
No error handling guidance in any tool description. What errors can infer raise? Is 'prompt too long' retryable or permanent? Does configure_pipeline validate input ranges or fail silently? Descriptions lack recovery guidance (e.g., 'If circuit_breaker_threshold is invalid, the server returns a validation error listing allowed range') per pattern:recovery-guide.
configure_pipeline 'worker' parameter accepts enum values (openai|anthropic|llama|echo) via type hint, but description does not state these as the only valid options. Description should explicitly list 'Valid values: openai, anthropic, llama, echo' or use enum constraint to prevent LLM hallucination of unsupported backends.