Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server exhibits significant definition quality gaps. While it attempts to expose 17 tools covering knowledge graph operations, audit chain management, and code analysis, the implementation suffers from: (1) duplicate tool definitions (record_event and record_correction appear twice with different schemas), (2) incomplete and vague descriptions that fail to guide LLM tool selection, (3) inconsistent parameter schemas with missing type information and descriptions, (4) no documented output schemas across any tool, and (5) poor error handling guidance. The tool names themselves are reasonable (verb-noun pattern), but the lack of clarity around when to use similar tools (query_fact vs multiple constraint-related tools) and missing parameter guidance severely limit agent usability. The server appears to be in early-stage development with schema definitions present but not production-ready.
Tools (17)
add_constraintwritesource verified63/100
Constraint added to the knowledge graph
add_factwritesource verified62/100
Fact added to the knowledge graph
get_audit_log_headread onlysource verified52/100
Audit chain head fetched
get_constraintsread onlysource verified57/100
Get constraints (linting rules, patterns, conventions) for a file
get_related_bugsread onlysource verified60/100
Get bugs fixed in a file and assess regression risk
ingest_pr_reviewswritesource verified62/100
Pull GitHub PR review comments and convert them into learned constraints in the knowledge graph
Duplicate tool definitions with inconsistent schemas: record_event appears twice (once at mcp_tool_dictionary.py with 'session_id', 'entities', 'reasoning' params; once at server.py with minimal params), and record_correction appears twice with different required fields. This ambiguity prevents correct LLM invocation.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract return values without knowing what fields to expect. This violates the pattern:response-shaper requirement that tools return structured, typed output.
Consolidate duplicate tool definitions. Choose one canonical schema for record_event (include session_id, entities, reasoning) and record_correction (include original_fact_id, correction_text, reasoning). Remove duplicates from mcp_tool_dictionary.py and server.py.
Document output schemas for all tools. For query_fact, document that it returns {facts: [{id, text, entity, evidence_path}], total_count, next_cursor}. For add_fact, return {fact_id, created_at, verified}. Use JSON Schema format visible in tool registration.
Expand tool descriptions to 100-200 characters with explicit WHAT/WHEN guidance. Example: 'query_fact: Search the knowledge graph for facts about code entities (APIs, functions, classes). Use this BEFORE modifying facts to check what is already known. Returns matching facts with relevance scores. Use entity_type filter (api|function|class) to narrow results.'
Add descriptions to every parameter. For query_fact.context, write: 'Optional domain context to improve relevance (e.g. "authentication layer", "payment processing"). Influences ranking of returned facts.' For record_event.entities, write: 'List of code entity names or file paths involved (e.g. ["UserService.login", "src/auth.ts"]). Used to build the dependency graph.'
Introduce a clear usage taxonomy in tool descriptions to disambiguate overlapping tools. E.g., 'Use query_fact to READ. Use add_fact to WRITE unverified observations. Use add_constraint to WRITE rules and patterns. Use record_event to LOG agent activities. Use record_correction to AUDIT prior facts.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 0 points across a rubric change (v1 → v2)
51/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
51
<=2025-11-25
v2
2026-03-09
D
51
-
v1
source verified
65/100
Query the knowledge graph for facts about entities (APIs, functions, classes, etc.)
record_correctionwritesource verified48/100
Record a user correction to Claude's output (high-priority learning signal)
record_correctionwritesource verified48/100
Correction to a prior fact recorded on the signed chain
record_decisionwritesource verified45/100
Record a decision trace: what the agent proposed and how the human responded
record_eventwritesource verified52/100
Record a development event (file edit, test run, etc.)
record_eventwritesource verified52/100
Event recorded on the signed chain
resolve_contradictionwritesource verified53/100
Contradiction between facts resolved
seed_projectwritesource verified57/100
Scan the project codebase and populate the knowledge graph with entities and relationships from existing code
validate_changeread onlysource verified62/100
Validate a proposed code change against known constraints
Vague and incomplete descriptions across most tools. Examples: 'Constraint added to the knowledge graph' (add_constraint) tells the LLM nothing about WHEN to use this vs add_fact or record_event. 'Record a development event' (record_event) lacks context about what qualifies as an event or what consequences recording has. Descriptions average 35-50 chars (below the rubric baseline of 194 chars for A+ tools) and lack WHAT/WHEN/WHY guidance.
Parameter descriptions missing or trivial. The 'context' param in query_fact says 'Additional context for the query' (vague). The 'proposed_content' param in validate_change lacks guidance on format. The 'change_description' in get_related_bugs is undescribed. Per the rubric baseline, 100% of A+ tools have param descriptions, this server has ~40% coverage.
No error handling guidance. Tools do not document what errors are retryable, user-fixable, or fatal. For example, if seed_project fails on a malformed git repo, does the LLM retry, ask the user to fix .git, or give up? No guidance is provided.
Overlapping tool responsibilities create selection confusion. query_fact, add_fact, add_constraint, record_event, record_correction, and resolve_contradiction all modify or query the knowledge graph. The LLM has no clear guidance on which tool to invoke for a given scenario. E.g., when should the LLM call add_constraint vs record_event?
Missing parameter type information in several tools. The 'decision_context', 'agent_proposal', and 'human_response' params in record_decision are typed as 'object' with no description of their internal structure. The 'evidence' param in record_event is similarly vague. This forces LLMs to guess structure.
record_decisionrecord_event
Add error handling guidance to each tool description. Example for seed_project: 'Returns success if all files indexed. On git errors, returns {status: "failed", error: "Git not initialized", next_step: "Initialize repo with git init or check WORLD_MODEL_DB_PATH"}.' Per pattern:recovery-guide.
Constrain free-form parameters with enums or patterns. For constraint_types in get_constraints, replace the open array with enum validation: ensure only valid values from ["linting", "architecture", "testing", "api_contract", "style"] are accepted. Document allowed values in description.
Remove the 'context' parameter from query_fact or replace it with structured fields (domain: string, focus_area: enum). Untyped 'object' params invite hallucination.
For record_decision, split the untyped object params into explicit fields or provide a schema example. E.g., decision_context: {scenario: string, options: string[]}, agent_proposal: {tool_name: string, params: object}, human_response: {accepted: boolean, feedback: string}.
Add pagination support to tools returning lists. query_fact should accept limit (1-100, default 20) and offset/next_cursor parameters. Document in description: 'Returns up to 20 facts per call. Use next_cursor to fetch more.'
Implement idempotency keys for write operations (add_fact, add_constraint, record_event). Accept an optional idempotency_id parameter so retried calls do not create duplicates. Document: 'If you retry with the same idempotency_id, the same fact ID is returned without creating a duplicate.'
Clarify permission boundaries. Do write tools (seed_project, ingest_pr_reviews) require special authorization? Document required scopes in descriptions: 'Requires repo:read and world-model:write permissions.' Per pattern:scope-declaration.