A Python-based Model Context Protocol server exposing mathematical tools, resource management, prompts, and client-LLM sampling capabilities with Azure Entra ID OAuth2 authentication support.
PrynAI MCP defines 9 tools with consistent HTTP transport (Streamable HTTP via FastMCP/Uvicorn). All tools have basic naming (verb_noun style: add, multiply, divide, echo, etc.) and descriptions present. However, several critical gaps limit enterprise readiness: (1) Parameter descriptions are minimal or absent for most tools, 'First integer' is too terse and doesn't explain use-cases or constraints. (2) Output schemas are not documented, callers cannot infer return structure from tool definitions alone. (3) Error handling lacks recovery guidance, divide by zero 'emits a warning' but doesn't specify whether the call fails, succeeds with null, or returns an error object. (4) Tool composition gaps: slow_square and long_task both report progress but lack idempotency hints and error classifications. (5) Security: set_counter and bump_counter are WRITE operations but lack permission gate descriptions or audit trail documentation. (6) No tool annotations (readOnlyHint/destructiveHint) are declared, despite the fact that risk levels are known (set_counter and bump_counter are WRITE, others READ_ONLY). Average tool description is ~60 characters; baseline is 194 chars. Parameter descriptions average ~30 chars; baseline is 72 chars. This server is above baseline for naming consistency but significantly below production standard for parameter clarity and output documentation.
Add two integers.
Increment the server counter by a step value and notify subscribers.
Divide a by b. Emits a warning on b=0.
Echo text and emit an info log notification.
Demonstrate progress notifications.
Multiply two integers.
Set the server counter to an exact integer value and notify subscribers.
Parameter descriptions are generic and below production baseline. Examples: 'First integer', 'Number of steps', 'Amount to increment by'. These lack context for LLM selection and usage. Baseline is 72 chars; most here are <40 chars. No constraints (min/max, format, enum) documented.
Output schemas are not documented in tool definitions. Callers cannot infer what add() returns (integer? object with 'result' field?). This breaks tool chaining and forces LLMs to guess structure. Pattern: 'Document return types and fields.'
Error handling lacks recovery guidance. divide() 'emits a warning on b=0', but does it throw, return null, or succeed? set_counter and bump_counter lack error classification (retryable, user-fixable, fatal). Agents cannot plan recovery without explicit guidance.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 57 | - | v1 |
Square an integer while reporting progress in steps.
Ask the client LLM to summarize; fall back if unsupported.
WRITE operations (set_counter, bump_counter) lack permission declarations and destructive annotations. No description states 'This modifies server state' or 'Requires admin:write scope'. Missing tool annotations (readOnlyHint/destructiveHint) mean clients cannot infer safety classification.
Tool descriptions are too brief. Baselines: 194 chars (p10=34, p90=392). Observed: 'Add two integers' (18 chars), 'Multiply two integers' (21 chars), 'Demonstrate progress notifications' (33 chars). At the low end of acceptable; most are <50 chars. Missing WHEN to use, prerequisites, or failure modes.
slow_square and long_task both report progress but neither declares idempotency nor explains retry behavior. If an agent retries slow_square, does it re-execute the computation, or return cached result? This ambiguity risks duplicate side effects.
Tool names could be more specific. 'long_task' is vague, what does it do? Baseline: verb_noun naming. Suggested: 'run_long_task', 'wait_for_steps', or rename to reflect actual operation.