A code graph engine that indexes repositories and provides code intelligence via Model Context Protocol. Supports multi-project serving, federation, and AI-optimized code context retrieval.
LeanKG has 3 tools with complete JSON schemas and descriptions present, but quality is uneven. All tools have descriptions exceeding 20 characters and input schemas with type definitions. However, descriptions are verbose and sometimes conflate multiple concerns, parameter relationships are underdocumented, error handling is not specified in tool definitions, and output schemas are not formally documented. The 'import' tool conflates 8+ distinct actions (repo, dir, docs, memory, session, ontology, read) into a single tool, violating single-responsibility principle. Schema completeness is moderate (60-70% range), but lack of output documentation and error guidance for LLM recovery prevents higher scores.
Import content into LeanKG: index a repository or directory of repositories (action=repo|dir, path), or curate agent memory (action=memory, command=create|str_replace|insert|delete|rename|add|replace|remove). Legacy tool name 'set' is superseded by this tool. Use action=dir with path="." to import the current directory as a scoped index target (FR-P2).
Query LeanKG. Empty action routes down the ladder: L1 exact identifier → L2 fuzzy keyword → L3 semantic (vectors). Every answer carries retrieval{rung,reason} + freshness. action=memory searches agent memory; action=exact|fuzzy|semantic pins a rung; graph verbs (impact/path/callers/callees/context/explain) and session/ontology reads are also available. Legacy tool name 'get' is superseded.
LeanKG health: inventory, freshness (fresh|possibly_stale|cold), watermark, backend, embeddings state (stamped models, vectors), last embed run.
Tool 'import' violates single-responsibility principle: conflates repo indexing, directory indexing, document curation, memory management, session handling, ontology operations, and file reading into one tool. This forces LLMs to reason about 8+ distinct action modes and creates ambiguous parameter relationships (e.g., 'command' is only valid when action=memory, 'file' only when action='add'|'replace'|'remove'). Should split into: import_repository, import_directory, curate_memory, manage_session, manage_ontology, read_compressed_file.
Parameter relationships are undocumented. In 'import' tool: 'command' param is only valid with action='memory' or action='session', but this constraint is buried in the description string. 'file' param is only for add/replace/remove but accepted at schema level for all actions. 'args' is an arbitrary object accepting mode/lines/fresh/turns/scope/cwd/bank depending on action/command combination. LLMs cannot infer these dependencies from the schema alone.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 43 | 2026-07-28+ | v2 |
Output schemas are not documented for any tool. Tool descriptions do not include 'Returns:' sections specifying the structure of responses. LLMs cannot plan downstream operations or extract required fields (e.g., what fields does import return? Does query return a list with pagination info?). This violates pattern:tool and pattern:paginated-result.
Error handling is not specified in tool definitions. No guidance on retryable vs. fatal errors, no recovery suggestions, no invalid-value feedback. If query returns no results, is that an error or empty success? If import fails on partial content, what does LLM see? Tools must define error cases and recovery paths per pattern:recovery-guide.
Descriptions contain internal details and examples that distract from LLM-friendly intent. E.g., 'query' description mentions 'L1 exact identifier → L2 fuzzy keyword → L3 semantic (vectors)' and 'rung/reason/freshness', these are implementation details. LLM-facing descriptions should state WHAT (search a code graph), WHEN (when you need to find a symbol), and HOW (pass a query term or identifier), not HOW it works internally.
The 'action' enum in both tools is overloaded and under-described. In 'import', action values range from 'repo' (index a repository) to 'memory' (curate memory) to 'session' (manage session) to 'read' (compress file read). These are semantically different operations. In 'query', action values include search verbs (exact, fuzzy, semantic), graph operations (impact, path, callers), and reading operations (session, ontology, portfolio). No parameter description explains the semantic boundary between these groups.
The 'args' parameter in both tools is a catch-all object with no schema. Description lists possible keys (depth, to, command, scope, etc.) but does not enforce them at schema level. Valid args keys depend on the action/command combination, making this parameter essentially untyped and forcing LLMs to guess valid combinations.
Parameter naming inconsistencies. 'import' uses 'path' for both repository/directory path AND memory file path (e.g., 'MEMORY.md, topics/x.md'). 'query' uses 'query' as the search term. No distinction between term-based search, identifier lookup, or graph traversal parameters. Should have separate, clearly named parameters: repository_path, directory_path, memory_file, search_query, graph_identifier, etc.
Descriptions do not guide LLM on when to call which tool or how to chain them. E.g., when should you call 'import' vs. skip to 'query'? If you need to update memory, which 'command' value applies? The 'status' tool's purpose is clear (health check), but 'import' and 'query' descriptions mix implementation details with use-case guidance, making it hard for LLMs to choose the right tool at the right time.