LLM-driven constraint solving (SAT, MaxSAT, SMT, CP, ASP) on top of agentic-python-coder
This server exhibits significant quality gaps that prevent it from being reliably used in production. Of the 10 tools, only 5 have properly documented schemas (select_backend, python_exec, python_interrupt, python_reset, submit_code). The arithmetic tools (add, subtract, multiply, divide) have complete schemas but lack depth in error handling and return type documentation. The read_resource tool has no visible schema definition in the provided source. Tool descriptions vary widely in quality: solver tools have detailed, actionable descriptions with context about backends and workflow; arithmetic tools have minimal descriptions (under 50 chars) that lack guidance on when to use them; read_resource lacks any visible description. The server conflates two distinct concerns (constraint solving + basic arithmetic) under one MCP interface, violating the single-responsibility principle. Error handling is minimal, divide() includes basic validation but lacks recovery guidance. Output schemas are not formally documented for any tool, forcing LLMs to infer structure. The arithmetic tools appear to be included as scaffolding for testing (minion/mcp-minion agent) but pollute the solver server's interface.
Add two numbers together. Use this when you need to calculate a sum.
Divide the first number by the second.
Multiply two numbers together.
Execute Python code in the persistent solving kernel and return the result
Interrupt the currently running Python code in the persistent solving kernel
Reset the Python kernel (clear state, optionally create a new independent kernel)
Read an MCP resource by its URI — one listed in the Resources section of your instructions, or received as a [resource_link] in a tool result.
read_resource tool has no visible description or schema in source code. The tool is referenced in minion/src/mcp_minion/agent.py but neither the description nor input schema are provided. This violates pattern:tool-description and prevents LLMs from understanding when/how to use it.
Arithmetic tools (add, subtract, multiply, divide) are included in a constraint-solving MCP server, violating single-responsibility principle (pattern:tool). These tools belong in a separate math helper server or should not be exposed at all. They dilute the solver server's interface and confuse agents about the server's purpose.
Arithmetic tool descriptions are trivially short (under 50 chars: 'Add two numbers together', 'Subtract the second number from the first') and lack guidance on when to use them. No mention of parameter constraints or return type.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Start solving a constraint problem: call this FIRST, once per problem. Sets up a persistent IPython kernel preloaded with the backend's solver library and helper functions, and returns the modeling instructions for that backend. Calling it again recycles the kernel for a new problem (previous session state is cleared). Backends: - pysat: Boolean satisfiability — feasibility, combinatorial search - maxsat: weighted optimization over Boolean constraints - z3: SMT — integers/reals/bitvectors, proofs, program verification - cpmpy: finite-domain constraint programming — scheduling, assignment, puzzles - clingo: Answer Set Programming — logic rules, defaults, reachability - didp: dynamic programming (state-space search) — routing, sequencing, packing Unsure which backend fits? Read the resource mcp-solver://guide first. Then iterate with python_exec (write, run, verify against the problem statement), and finish by calling submit_code with the final, verified, self-contained program. UNSAT / "no solution exists" is a valid outcome.
Submit the final, verified, self-contained solver program. Returns success/failure with optional resource link to the submission.
Subtract the second number from the first.
python_exec and python_interrupt lack sufficient description detail. python_exec (60 chars) does not explain the execution context (persistent kernel), available libraries, timeout behavior, or expected return format. python_interrupt (36 chars) does not explain whether interruption is graceful or forceful.
No output schemas documented for any tool. LLMs cannot infer what fields python_exec returns, what structure submit_code produces on success, or what shape divide() error messages take. Per pattern:tool, 'Document the output schema' is critical for tool composition.
divide() includes basic zero-check ('if b == 0: raise ValueError') but no recovery guidance in description. Per pattern:recovery-guide, error responses must tell the LLM what to do next. Currently the LLM sees only 'Cannot divide by zero' with no suggestions.
python_reset parameter 'kernel_id' is undocumented in terms of expected format, behavior when kernel_id does not exist, or side effects of omitting it. Description says 'Optional kernel ID to reset' but does not explain whether omitting it creates a new kernel or resets an existing shared kernel.
python_exec timeout parameter (integer seconds) has no min/max bounds documented. Unbounded timeout invites LLMs to pass absurd values (0, -1, 999999).
No error classification or recovery guidance for select_backend when solver backend is invalid or unavailable. If LLM passes 'solver: gibberish', the error response does not guide next steps (try list_available_backends, use a different backend, etc.).
python_exec and submit_code descriptions do not explain the expected format of code strings (Python syntax, imports, encoding) or what constitutes a 'self-contained' program. This ambiguity forces LLMs to guess.