Decision provenance for AI-coded codebases: the why, and what was already tried and rejected.
Selvedge provides 8 tools with reasonable naming discipline and mostly complete schemas. However, descriptions are sparse and inconsistent. Most tool descriptions (log_change, diff, blame, changeset, search, stale_decisions) are under 100 characters, below the production baseline of 194 chars and below the minimum of 10 - 1024 chars for adequate context. Parameter descriptions exist but are minimal. The server has decent structure but lacks the depth needed for production-grade tool discoverability. No output schema documentation visible. This is solidly C+ territory, better than median community servers but missing critical polish.
Get the most recent change + context for an entity
Retrieve all events in a named feature/task group
Get change history for an entity
Filtered history across all entities
Log a code change - record a change event (incl. dual-event renames)
Prior change attempts on an entity + inferred outcome
Full-text search across all events
Tool descriptions are severely underdeveloped. Six of eight tools (diff, blame, history, search, and stale_decisions) have descriptions under 50 characters. The production baseline is 194 chars (p10=34, p90=392). Short descriptions lack context about when to call each tool, what it returns, and how it differs from similar tools.
Output schemas are not documented anywhere in the provided source. Tools like history (no input params) and changeset (returns events) have no visible schema declaration for their responses. LLMs cannot plan downstream calls or extract data without knowing what fields to expect.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 55 | 2025-06-18+ | v2 |
Decisions past their revisit_after that are still in use
Parameter descriptions are minimal or missing for some tools. 'history' has zero input parameters documented. 'stale_decisions' has zero input parameters documented. Even where parameters exist (entity_path, query), descriptions are terse and lack guidance on format, valid ranges, or dependencies.
Tool naming creates ambiguity. 'log_change', 'diff', 'blame', 'history', and 'changeset' all deal with change history but perform different operations. Descriptions do not clearly explain when to use each vs. the others. An LLM may confuse 'diff' (change history for one entity) with 'history' (filtered across all entities) or 'changeset' (named feature group).
Error handling guidance is not visible. No recovery hints documented for cases like entity_path not found, search_query returning no results, or changeset_id not found. Tools should guide the LLM on what to do next.
Pagination and result limiting not documented. Tools like 'history' and 'search' may return large result sets, but no limit/offset or pagination parameters are visible. The rubric baseline requires paginated results with explicit limits to prevent context window exhaustion.