RuleFlow presents 31 tools with minimal visible schema documentation and sparse parameter descriptions. Source code review reveals tool definitions in mcp_rules_assistant/tools.py, but the provided code sample does not include actual tool registration, input/output schemas, or parameter type definitions. Tools are listed with names and brief descriptions only; no parameter constraints, enums, or return types are visible. This makes it impossible to verify schema completeness. The naming pattern uses dot-notation (e.g., project.detect, rules.init) which is non-standard for MCP tools and suggests possible under-the-hood routing rather than discrete tool registration. Descriptions range from 5-20 words, which is below the 50-200 char baseline for LLM-optimized descriptions. No evidence of error handling guidance, output schema documentation, or composition patterns across tools.
Tools (31)
ci.autofixwritesource verified25/100
Auto-fix CI by regenerating missing steps
ci.generatewritesource verified27/100
Generate CI workflow (GitHub Actions)
ci.validateread onlysource verified27/100
Validate generated CI workflow content
compliance.commitmentwritesource verified32/100
Return AI compliance commitment text and optionally write to project
NO VISIBLE INPUT/OUTPUT SCHEMAS for any tool. Provided source sample does not include actual tool registration code, parameter type definitions, or return type documentation. Per scoring rules, schema score MUST be 0 when schemas are not visible.
Rename tools to use verb_noun convention: 'detect_project_language', 'switch_project_context', 'validate_rules', 'ingest_rules_from_docs' instead of dot-notation. This makes intent parseable from the name alone.
Expand ALL tool descriptions to 50-200 characters. Include WHAT, WHEN to call (vs similar tools), and what output to expect. Example for project.detect: 'Analyze the project directory to identify the primary programming language, framework, and complexity level. Call this first to determine which rules profile applies.'
For EVERY parameter, add a description explaining what it controls, expected format, valid values, and constraints. Example: 'profile (string, required): Rules profile to apply. Choose from: "minimal", "standard", "strict". Defaults to "standard" if omitted.'
Document return types and field names for ALL tools. Use JSON Schema notation or structured examples. Example for memory.snapshot: 'Returns {"turns": [{"role": string, "content": string, "timestamp": ISO8601}], "summary": string, "turn_count": integer}'.
Add error handling guidance to WRITE operations. For each tool that modifies state, document: (1) what can go wrong, (2) whether the error is retryable, (3) what the LLM should do next. Example for fs.apply_patch: 'If patch fails due to file conflicts, call fs.revert_last_patch() to undo. If permissions denied, request elevated access or check file ownership.'
Clarify tool composition: map outputs to inputs. Document which tools should be called in sequence. Example: 'Call project.detect to identify language, then pass the language to rules.init. Call rules.init to populate rules, then call rules.enforce to apply them.'
Score history
Overall score trend
↑ 34 points across a rubric change (v1 → v2)
34/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
34
2026-07-28+
v2
2026-03-09
F
0
-
v1
28/100
Summarize weak/groups/near in one payload
env.diagnoseread onlysource verified28/100
Diagnose environment/tools/config presence
env.preparewritesource verified22/100
Prepare language environment
fs.apply_patchwritesource verified30/100
Guarded write with checks
git.install_hookswritesource verified25/100
Install git hooks
ide.scaffoldwritesource verified27/100
Generate per-IDE integration scaffold
memory.append_turnwritesource verified25/100
Append a conversation turn
memory.snapshotread onlysource verified27/100
Produce last 20 turns and summary
memory.toggle_autowritesource verified22/100
Enable/disable rolling memory
nl.commandwritesource verified27/100
Natural language command dispatcher
plan.setwritesource verified27/100
Set plan fields (status/current/next)
plan.suggest_nextread onlysource verified27/100
Suggest next steps from memory/plan
plan.updatewritesource verified25/100
Update project plan markdown
project.detectread onlysource verified25/100
Detect project language/framework/complexity
project.linkwritesource verified25/100
Add cross-project relationship link
project.switchwritesource verified25/100
Switch or create project context
rules.enforcewritesource verified22/100
Set enforcement level
rules.ingestwritesource verified25/100
Ingest project rules from docs
rules.initwritesource verified22/100
Select general rules by profile
rules.maximaread onlysource verified28/100
Return coverage upper-bounds (maxima) from compiled rules
rules.onboardwritesource verified28/100
Interactive onboarding to choose and apply rules profile
Tool names use dot-notation (project.detect, rules.init) instead of verb_noun convention. This is non-standard for MCP tool registration and suggests possible internal routing rather than discrete tool definitions. LLMs expect names like 'detect_project', 'switch_project', 'validate_rules' to parse intent from the name alone.
Descriptions are too brief (5-20 words, <100 chars). Baseline for LLM-optimized descriptions is 50-200 chars. Current descriptions ('Detect project language/framework/complexity', 'Switch or create project context') lack WHEN/WHY guidance and do not explain when to call vs similar tools.
No visible parameter descriptions for any tool. Tool definitions list names only; no descriptions of what each parameter controls, expected format, valid ranges, or constraints.
No error handling guidance visible. Tools marked as WRITE operations (project.switch, rules.init, fs.apply_patch, git.install_hooks, etc.) have no documented recovery steps, error categories, or what to do if the operation fails.
No visible documentation of output schemas or return types. Tools like memory.snapshot, coverage.report, rules.maxima claim to return structured data ('Produce last 20 turns and summary', 'Summarize weak/groups/near in one payload') but no field definitions are visible.
Tool composition unclear: no evidence that tool A's output fields match tool B's input parameter names. For example, project.detect vs project.switch, unclear if detect returns a 'project_id' that switch accepts, or if resolution requires manual mapping.
Overlapping tool purposes without clear disambiguation. Multiple tools operate on 'rules' (rules.init, rules.ingest, rules.validate, rules.enforce, rules.onboard, rules.maxima, rules.resolve) with no documented distinctions. LLMs will struggle to select the right one.
Reduce tool count by consolidating overlapping rules tools. Merge rules.init, rules.onboard into a single 'configure_rules' tool. Merge rules.validate, rules.maxima, rules.resolve into a 'analyze_rules' tool. This reduces LLM selection complexity.
Add enums and constraints to parameters. Example: 'enforcement_level (enum): "off", "warn", "error", "fatal". Default: "warn".' Enums are machine-parseable and prevent LLM hallucination.
For tools returning lists (coverage.report, coverage.near), add pagination support: 'limit (integer, 1-100, default 20): Max results per page. next_cursor (string, optional): Pagination cursor for next batch.' Document the total count in responses.
Add tool annotations (toolAnnotations feature) to indicate read-only, destructive, and idempotent tools. Example: 'project.detect is read-only (readOnlyHint=true). fs.apply_patch is destructive (destructiveHint=true). memory.append_turn is idempotent (idempotentHint=true).' This helps agents reason about side effects.