Spring Boot microservices platform for AI agents with RAG (Retrieval-Augmented Generation), diagram extraction, security posture analysis, and web search capabilities. Includes an MCP server (posture-service) for security posture tool exposure.
This MCP server exhibits significant quality gaps across definition, naming, and schema documentation. Of 8 tools, 6 have minimal descriptions, 2 tools lack descriptions entirely, and tool naming is inconsistent (mixing verb_noun with hyphenated forms). Input schemas are present but lack detail, most parameters have type strings with minimal descriptions. Output schemas are undocumented. The server mixes domain-specific security tools (posture, diagram_extract, rag_query, web_search) with generic demo tools (getOrderStatus, getTicketStatus, get-account-balance) suggesting this is a proof-of-concept rather than production-ready. Security tools have better descriptions but still lack error handling guidance and output schema documentation. No tool annotations (readOnlyHint/destructiveHint) are evident. The write-capable tools (getOrderStatus, getTicketStatus) lack confirmation patterns or dry-run options.
Extract components/edges from an uploaded architecture diagram (image, Draw.io PNG export, or screenshot). Returns a normalized JSON graph.
Get the current account balance for a given account ID
Query internal security policies and checklists with RAG; return concise quotes + citations.
Returns the posture of a service by service id and environment.
Get security posture for a service (internetFacing, data classes, TLS, vulnerabilities, secrets).
Two tools named 'security_posture' with different signatures (one with env param, one without). LLM cannot disambiguate; will select wrong one.
Tools getOrderStatus and getTicketStatus have ZERO descriptions. LLM cannot determine when to use them.
Inconsistent tool naming conventions: getOrderStatus, getTicketStatus use camelCase; diagram_extract, rag_query use snake_case; get-account-balance uses hyphens. LLM parsing and tool selection suffers from inconsistency.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 51 | <=2025-11-25 | v2 |
Fetch OWASP/NIST/CWE guidance (allowlisted domains only). Returns title/url/snippet.
No output schemas documented for any tool. LLM cannot infer what fields to expect, forcing parsing of unstructured responses and multi-step discovery.
Write-capable tools (getOrderStatus, getTicketStatus) lack confirmation/dry-run patterns. LLM could accidentally create duplicate orders or tickets on retry.
Parameters like topK in rag_query and web_search lack numeric bounds (e.g., min=1, max=100). LLM could pass absurd values (topK=999999) causing timeout or excessive token consumption.
No error handling guidance documented. Tools lack recovery hints (e.g., 'If serviceId not found, try search_services() first'). LLM cannot self-correct on failure.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. MCP client cannot determine which tools are safe to retry or which modify state.
Parameter descriptions are sparse. Most lack format constraints, ranges, examples, or enum values. E.g., fileName in diagram_extract, what formats are supported? Max size?