A multi-agent system for legal document processing and compliance advisory using LangGraph, MongoDB, and Google Generative AI. Consists of ADK-A (document processor) and ADK-B (compliance advisor) working in tandem to analyze legal documents, extract clauses, assess risks, and generate compliance suggestions.
This server exhibits severe gaps across all dimensions of definition quality. Of 9 tools, only 4 have visible input schemas in the provided source code, and most descriptions are generic or minimal. The server appears to be a Flask-based multi-document processing pipeline (ADK-A and ADK-B) but lacks LLM-optimized tooling. Tool naming is acceptable (verb-noun pattern present), but schemas, parameter descriptions, and error handling are critically underdeveloped. The codebase shows MongoDB integration and inter-service communication, but no evidence of structured output documentation, pagination support, or error recovery guidance. Most tools would not guide an LLM effectively on when to use them or how to handle failures.
Receive communication from ADK-B
Generate compliance suggestions based on ADK-A analysis results.
Generate compliance suggestions based on ADK-A results
Retrieve Markdown report for a document
Retrieve compliance report for a document
Retrieve processing result for a document
Health check endpoint
Missing input schemas for poll_for_work. No visible parameter documentation; tool name suggests polling operation but no mechanism documented to the LLM. Prevents LLM from understanding call constraints.
Generic, under-specified descriptions across most tools. E.g., 'Poll for new work from ADK-A' (30 chars) does not explain: What constitutes 'work'? What does the LLM do if none is found? When should it call poll_for_work vs generate_suggestions? Descriptions should be 100-200 chars and answer WHAT, WHEN, and NEXT STEPS.
No output schema documentation visible. Tools like generate_suggestions, get_report, and get_markdown_report have no documented return structure. LLMs cannot plan downstream calls or extract needed fields without schema clarity.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 33 | - | v1 |
Poll for new work from ADK-A
Process a legal document through ADK-A pipeline
Error handling is missing or inadequate. Flask endpoints return generic {'error': str(e)} (500 responses) with stack traces. No actionable recovery guidance. LLM cannot know: Should I retry? Ask the user? Try an alternative tool? All error responses require recovery instructions.
Parameter descriptions lack format constraints and validation rules. E.g., document_id accepted as string but no length, format, or naming convention documented. Parameter 'document_type' defaults to 'contract' but no enum of valid types provided. Constraints must be explicit in descriptions or schema enums.
Ambiguous tool composition. Multiple tools operate on documents (process_document, get_result, get_report) but interdependencies are undocumented. What state transitions occur? Does process_document always succeed? Must get_result be called after? No guidance for multi-step workflows.
No pagination or result limiting documented. Tools like get_report and get_markdown_report return full documents from MongoDB with no mention of size limits, pagination, or streaming. Large reports could exhaust context windows.
communicate tool accepts arbitrary 'data' object with no schema. This parameter is unvalidated and could accept any payload. No bounds on size, structure, or required fields. Schema must declare expected fields and types.
Idempotency not addressed. Tools like generate_suggestions and process_document modify state (write to MongoDB, call external services) but no indication whether repeated calls are safe. Agents retry on ambiguous failures, non-idempotent tools risk duplicates.
No permission or scope declarations. Tools perform state-altering operations (document processing, suggestion generation) with no visibility into authorization model. Who can call these? What audit trail exists?