A proposed 16+ layer model for understanding and discussing AI system architectures. Documentation and proposal management system with workflow automation.
AILIS exposes 39 tools across 6 Python scripts. While tool names follow verb_noun conventions (check_heading_hierarchy, extract_version_from_package_json, generate_metrics_report), critical quality gaps severely limit production readiness. Most tools lack input parameter descriptions entirely, parameters are defined with type but no description text explaining what they control or accept. Output schemas are not documented anywhere in the source. Descriptions are present but often generic (e.g., 'Check if headings follow proper hierarchy') without explaining WHEN to use the tool or what the LLM should do with results. No error handling guidance, no recovery paths, and no indication of which tools are idempotent vs destructive. The tools appear to be utility functions for CI/CD workflows (markdown linting, version checking, changelog generation) rather than agent-facing tools with LLM-optimized interfaces.
Analyze version consistency across repository files.
Calculate metrics from workflow runs.
Categorize commits and PRs by type.
Check for images without alt text.
Check if headings follow proper hierarchy (no level jumps).
Check for non-descriptive link text.
Compile the final README.
Parameter descriptions missing or minimal across all 39 tools. Parameters defined with type but no explanation of what they control, valid ranges, or constraints. E.g., 'days' in fetch_workflow_runs has no description of valid range (1-365?) or default behavior.
Output schemas not documented. No indication of what fields tools return, their types, or structure. LLMs cannot plan downstream tool calls or extract required data without knowing response shape.
No error handling guidance. Tools provide no recovery hints, categorization of errors as retryable/user-fixable/fatal, or actionable error messages. E.g., if a file is not found, no suggestion to check the path or list available files.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | 2026-07-28+ | v2 |
Extract version from Cargo.toml.
Extract latest version from CHANGELOG.md.
Extract version references from markdown files.
Extract version from package.json.
Extract version from pyproject.toml.
Extract version from plain text file.
Fetch recent workflow runs from GitHub API.
Add blank lines around headings.
Add blank lines around lists.
Add language specifiers to code blocks.
Fix emphasis style according to markdownlint rules: bold=asterisk, italic=underscore.
Break long lines at logical points.
Fix all markdown issues in a file.
Ensure file ends with single newline.
Generate comprehensive metrics report for all workflows.
Generate dynamic proposal listing.
Generate changelog section for a version.
Generate workflow status badges.
Get contributing information.
Generate contributors section.
Generate documentation links.
Generate footer content.
Get commits from git history.
Get merged pull requests from GitHub API.
Get project description.
Get current project statistics.
Load existing changelog content.
Load the README template.
Parse conventional commit message.
Print human-readable metrics summary.
Scan repository for version information across multiple file formats.
Update the changelog with new version.
Destructive tools (generate_metrics_report, fix_markdown_file, update_changelog) lack idempotency hints and confirmation patterns. No indication whether repeated calls with same input are safe or cause duplicate side effects.
Tool descriptions are generic and lack WHEN/WHY context. E.g., 'Check if headings follow proper hierarchy' does not explain when an LLM should call this vs other accessibility checks, or what to do with the result.