A Model Context Protocol server that provides semantic search (RAG) and file access tools over a repository index. Builds an in-memory semantic index using embeddings and chunk-based vector search.
The server has well-structured tool definitions with complete schemas and descriptions for all three tools (rag_query, read_file, list_files). All tools follow verb-noun naming conventions and have explicit JSON Schema input definitions with type constraints and descriptions. However, there are several medium-severity issues: (1) Output schemas are described in prose but not formally documented in JSON Schema format, which limits LLM ability to parse response structures reliably. (2) Error handling is minimal, no error recovery guidance or actionable error messages visible in the definitions. (3) Some parameter descriptions lack explicit range/format constraints (e.g., 'query' in rag_query has no guidance on length, special characters, or query format best practices). (4) Tool descriptions mention interpolated variables like '${FOLDER_INFO_NAME}' which are environment-dependent and could confuse LLMs about what folders are actually accessible. (5) The read_file tool lacks explicit validation guidance for path traversal attacks or clarification of what 'relative to folder' means for security. Overall, definitions are above-average for community MCP servers but fall short of production-grade clarity.
List files (and subdirectories) within a directory under '${FOLDER_INFO_NAME}' folder. Supports recursive listing with optional depth control and extension filtering.
Semantically search files under '${FOLDER_INFO_NAME}' folder and return relevant chunks. Returns an array of match objects with properties: 'path' (string, file path), 'score' (number, similarity score 0-1), 'snippet' (string, matching text chunk), 'totalLines' (number, original file line count), 'fileSize' (number, file size in bytes). Results are sorted by relevance score descending.
Read a specific file under '${FOLDER_INFO_NAME}' folder (optionally a line range). Returns the file content as a string. If startLine and/or endLine are specified, returns only the requested 1-based inclusive line range; otherwise returns the entire file content.
Output schemas not formally documented in JSON Schema format. Tool descriptions include prose descriptions of return fields (path, score, snippet, totalLines, fileSize for rag_query) but no machine-readable schema for the response structure. This forces LLMs to infer response structure from text, risking misparse and context waste.
Error handling and recovery guidance absent. Tool descriptions do not explain what errors can occur (e.g., file not found, permission denied, query too short) or what the LLM should do next (retry, ask user for clarification, try a different tool). Per pattern:recovery-guide, error responses must tell the LLM what to do next.
Security concern: read_file tool lacks explicit path traversal validation guidance. Description does not clarify whether paths like '../../../etc/passwd' are rejected, sanitized, or permitted. No mention of permission checks or boundary enforcement. Per pattern:tool-gateway, all agent-provided input (including file paths) must be validated against injection attacks.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | C | 61 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 24 | - | v1 |
Environment-dependent tool descriptions. Tool descriptions use '${FOLDER_INFO_NAME}' placeholder which is resolved at runtime. LLMs may not understand that this variable is substituted, leading to confusion about what folders are actually accessible. Consider moving folder name into the description as a concrete value at startup, or documenting this pattern in the tool definition.
Inconsistent relative path documentation. read_file describes paths as 'relative to FOLDER_INFO_NAME folder', but list_files says 'relative to ROOT'. These should be clarified as identical or explained if they differ. Per pattern:tool-chain, ensure tool outputs and parameter references align so agents can chain calls without confusion.
query parameter in rag_query lacks format guidance. Description says 'Natural language search query' with advice to 'Use concise, specific terms' but does not specify: maximum length, forbidden characters, minimum length, or examples of good vs bad queries. This forces LLMs to guess optimal query structure.
Pagination and result limits not clearly explained. rag_query accepts top_k (1-50) but description does not explain what happens if the query matches fewer results (does it return < top_k items?). list_files has a 'limit' parameter (default 500, no stated max) but description does not explain sorting/ordering when results are truncated or what 'total' count means for continuation.
No parameter interdependency documentation. list_files has several optional parameters (recursive, maxDepth, includeExtensions) but description does not clarify: if recursive=false, is maxDepth ignored? If includeExtensions is omitted, are all file types included? These details are critical for the LLM to form correct calls.