The server defines 4 tools with clear names, reasonable descriptions, and structured input schemas. However, there are significant gaps in output schema documentation, error handling guidance, and parameter validation rules. Tool names are well-chosen (verb-noun pattern), and descriptions are in the acceptable range (150-250 chars), but lack specificity about when to use each tool vs. alternatives. Parameter descriptions are present but generic. Output schemas are inferred from code but not explicitly documented for the LLM. Error handling is minimal, no recovery guidance, no retryability classification, no actionable error messages.
Tools (4)
analyze_codebaseread onlysource verified76/100
Analyze a codebase and return a full profile.
explain_decisionread onlysource verified73/100
Explain an architectural decision based on codebase evidence.
This is a heuristic-based explanation (no LLM). It searches git history,
file structure, and dependencies for evidence related to the question.
Output schemas not documented. Tools return dicts and strings, but LLMs cannot see the structure of returned fields (e.g., what fields does analyze_codebase.to_dict() contain?). Forces LLMs to guess at response shape and breaks composition (cannot chain outputs to downstream tools).
No error handling or recovery guidance. Tools validate the repo path but provide only generic ValueError messages. No guidance for LLM on retryability, user-fixable vs. fatal errors, or alternative actions (e.g., 'not a git repo, check the path or try a different directory').
Parameter descriptions lack actionable constraints. 'max_files' has no range guidance (LLM could pass 1 or 1000000). 'style' enum is mentioned in description but not declared as a constrained enum type in the schema. No mention of valid enum values in the schema.
Recommendations
Document output schemas explicitly for all tools. For 'analyze_codebase', specify: 'Returns CodebaseProfile with fields: name (str), code_structure (dict with primary_language, src_layout, has_tests, etc.), dependencies (dict with framework, dependencies list), git_history (dict with total_commits, contributors, hot_files), patterns (dict with architecture_patterns, naming_conventions, test_patterns).' This enables LLMs to extract and chain fields.
Add error classification to all tools. Wrap path validation errors with guidance: 'Path validation failed: <reason>. Retryable: false. User action: Check that the path is a valid git repository root (contains .git directory).'
Declare 'style' parameter as an enum in the schema with values ['standard', 'minimal']. Update description to explain: 'standard': generates full documentation with all sections; 'minimal': compact version focusing on key patterns and risks.
Add result limits to tool descriptions and enforce in code. E.g., 'Returns a CodebaseProfile analyzing up to 500 files (default). For large repos, increase max_files to 1000 or higher, but be aware token usage scales with codebase size.' Explicitly cap hot_files and conventions lists to top 20 items.
Differentiate tool purposes in descriptions. Rewrite: 'analyze_codebase': Full structural and historical analysis, use this first to get a complete profile. 'explain_decision': Targeted evidence-based explanation of a specific architectural choice, faster than analyze_codebase. 'generate_claude_md': Generates a markdown document for sharing with Claude or other tools. 'map_tribal_knowledge': Identifies knowledge silos and undocumented conventions, use this to assess onboarding risk.'
Tool descriptions do not explain when to use one vs. another. 'analyze_codebase' vs. 'explain_decision' vs. 'map_tribal_knowledge' are all data-extraction tools on the same resource, but the descriptions don't disambiguate their purposes for the LLM. This forces the LLM to reason about subtle differences.
No pagination or result limit enforcement. 'analyze_codebase' returns a full CodebaseProfile; 'map_tribal_knowledge' returns all conventions, silos, and risk areas. No documented limit on result size. Large profiles could exhaust token budgets. Rubric baseline: tools should cap results at 20-50 items and offer pagination.
Tool composition is unclear. If 'explain_decision' calls 'analyze_repo' internally, the LLM doesn't know that calling both tools sequentially will re-analyze the repo twice (wasting compute). No guidance on which tools to call in sequence and what data flows between them.
Add parameter descriptions for 'include_git_history' explaining trade-offs: 'Set to false for speed if git history is not needed (speeds up large repos). Default true provides fuller context.'
Provide pagination guidance. Modify 'map_tribal_knowledge' to accept optional 'limit' parameter (default 20) and return {'conventions': [...], 'knowledge_silos': [...], 'risk_areas': [...], 'total_conventions': int, 'total_silos': int}. Document: 'Results are limited to 20 items each by default. Set limit higher (up to 500) for complete analysis, but expect larger token usage.'
Add a tool composition note in the server instructions or tool descriptions: 'Calling analyze_codebase, explain_decision, and map_tribal_knowledge all re-analyze the codebase. For efficiency, call analyze_codebase once, then filter results locally; or call explain_decision and map_tribal_knowledge which internally cache analysis.' (This requires code changes to support caching or state, consider for next version.)