Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server has 15 tools with HTTP transport and basic schema documentation, but significant gaps in description quality, parameter validation, and error handling. Most tool descriptions are terse (10-50 chars), and many parameters lack descriptions or type constraints. Schema definitions are present for most tools but incomplete. The server follows a RAG-Fusion domain pattern well but does not optimize for LLM tool selection and reasoning. Average tool score: 52/100.
Tools (15)
build_indexwritesource verified68/100
Trigger an index build job to ingest PDFs, extract text and images, chunk documents, generate embeddings, and build FAISS and BM25 indexes
classifyread onlysource verified50/100
Classify query intent
delete_jobreversiblesource verified57/100
Delete a job from the job history
get_chunk_detailsread onlysource verified57/100
Get detailed information about a specific chunk by ID
get_chunk_detailsread onlysource verified57/100
Get detailed information about specific chunks by ID
Duplicate tools: search_manual, get_chunk_details exist twice in the tool set (tools 1 & 11, tools 2 & 12) with slightly different schemas and descriptions. LLMs will struggle to disambiguate and may invoke the wrong variant.
List tools (list_jobs, list_pdfs, list_indexed_documents, list_indexes) have minimal descriptions (20-25 chars). These lack actionable context for when/why an LLM should call them, what structure they return, or dependencies.
classify tool description is 30 chars ('Classify query intent'), too terse. No explanation of what classification means, what intents exist, when to call this vs direct retrieve, or output structure. LLMs cannot reason about when to use it.
classify
Recommendations
Merge duplicate tools: consolidate search_manual (tools 1 & 11) into one canonical definition; consolidate get_chunk_details (tools 2 & 12). Use a single, well-documented version with all features.
Expand all list_* descriptions to 80-150 chars explaining what they return, what structure/schema is provided, and when to call them. Example: 'List all indexed documents and retrieve metadata including chunk counts, embedding models, and last-modified timestamps. Call before search to discover available indexes.'
Expand classify description: explain what intents are supported (e.g., 'query_intent will be one of: general_info, technical_requirements, how_to, troubleshooting, or unknown'), whether intent is auto-classified if not provided, and how intent affects retrieve results.
Document all parameter descriptions. For list_* tools, add 'Input schema: none required.' For get_job_status and delete_job, add: 'job_id: string. The unique job identifier returned by build_index or list_jobs.'
Add output schema to every tool. For search_manual, document: 'Output: { chunks: [{chunk_id, text, page, score, source}], metadata: {intent, index_version}, context: string }'. For build_index, document: 'Output: { job_id, status, progress_percentage, started_at, estimated_completion }.'
Add error handling guidance. For build_index: 'If the job fails, check get_job_status for error details. Common issues: invalid PDFs (unsupported format), insufficient disk space (free space required), or OCR timeout (set ocr_languages to primary language only). Retry with smaller chunk_size if indexing fails.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 4 points across a rubric change (v1 → v2)
51/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
51
<=2025-11-25
v2
2026-03-09
F
47
2024-11-05+
v1
read onlysource verified38/100
List available indexes and their metadata
list_jobsread onlysource verified37/100
List all index build jobs
list_pdfsread onlysource verified37/100
List all PDF files in the PDF directory
open_pdf_pageread onlysource verified62/100
Open a PDF document at a specific page
retrieveread only50/100
Retrieve and rerank chunks using RAG-Fusion pipeline with multi-strategy retrieval, reciprocal rank fusion, cross-encoder reranking, and context synthesis
search_manualread onlysource verified68/100
Search the CI360 manual using RAG-Fusion multi-strategy retrieval
search_manualread onlysource verified68/100
Search the indexed manual with optional examples and confidence filtering
upload_pdfwritesource verified68/100
Upload a PDF file for indexing. The file will be saved to the PDF directory and can be included in the next index build.
No error handling guidance in any tool description. Tools like delete_job (reversible) and build_index (long-running) should indicate retry behavior, expected failures, and recovery steps. Per pattern:recovery-guide, error responses must tell LLMs what to do next.
No confirmation or dry-run for destructive operations. delete_job is marked REVERSIBLE but has no dry-run parameter or confirmation step. upload_pdf and build_index (WRITE operations) lack warnings. Per pattern:confirmation-request, irreversible operations should support preview/confirmation.
Output schemas not documented. No tool describes what fields it returns, data types, or example structures. LLMs cannot plan downstream calls or extract values if output structure is opaque.
Parameter naming inconsistency and gaps. 'top_k' is clear but 'intent' (classify tool) is undefined, should describe valid intents, whether auto-classification is available, and what intent values the retrieve tool expects. Mutually exclusive params (search_manual: min_confidence vs intent) lack mutual-exclusivity documentation.
No permission checks or scope declarations visible in code. Tools like delete_job and build_index (state-modifying) should declare required permissions (e.g., 'admin:index', 'write:pdf'). No gate or audit trail evident.
Long-running operations (build_index) have no progress callback or status polling guidance. Job status must be polled manually via get_job_status. No indication of timeout, expected duration, or cancellation support. Agents will timeout waiting for completion.
build_index
Add dry-run or confirmation for destructive ops. For delete_job, add optional dry_run: bool (default false) that returns 'This will delete job X with Y completed chunks. Confirm by calling with dry_run=false.' For upload_pdf, add optional validate_only: bool to check PDF before storage.
Clarify parameter relationships. For search_manual: 'If intent is provided, min_confidence is relative to that intent's baseline; if intent is None, auto-classification occurs and min_confidence applies to the auto-classified intent.' Document mutually exclusive params explicitly.
Add capability discovery. Include a 'get_server_info' or 'describe_indexes' tool that returns available indexes, embedding models, supported OCR languages, current capacity, and rate limits. This prevents invalid requests and guides LLM planning.
Document required permissions. Annotate delete_job, build_index, upload_pdf with '@permission admin:index' or '@scope write:pdf'. Add tool descriptions: 'Requires admin:index permission. Contact your administrator if you lack access.'
Add pagination guidance. For list_jobs, list_pdfs, list_indexed_documents, add optional limit (default 20, max 100) and offset (default 0) params. Return total_count so LLMs know when to paginate.
Set timeout expectations. For build_index, document expected runtime ('Typical runtime: 30 seconds per 10 MB of PDF content; timeout: 5 minutes. For large batches, monitor via get_job_status and retry if timeout occurs.').
Add idempotency declarations. For upload_pdf, clarify: 'Idempotent: uploading the same file twice with the same filename is safe; the second call overwrites the first without duplication.' For build_index, clarify: 'Non-idempotent: each invocation creates a new index; calling twice creates two separate index versions.'
Reduce result verbosity. If search_manual returns 50+ fields per chunk, document which are core (text, page, score) and which are optional (embedding, metadata, provenance). Cap default limit to 10 results unless explicitly requested.