Production-ready AI coding agent with semantic search, MCP server, REST/WS API, and export tooling
ctxai has 11 tools with mixed quality. Naming follows verb_noun convention (list_, index_, query_, etc.) and is generally clear. However, parameter descriptions are inconsistent, some tools like query_codebase and index_codebase have detailed parameter docs, while others like git_commit and git_add have minimal parameter descriptions. Most critically, NO tool definitions include output schemas in the source code, the rubric requires documented return types for A+ tools. All tool schemas use proper JSON Schema with types, which is good. The bash tool includes policy-controlled execution with sandbox support, which is sophisticated. Error handling is present but not systematically articulated in tool descriptions, agents cannot determine if errors are retryable or fatal. Word_count tool is trivial but properly defined.
Execute one policy-approved command inside the repository (no shell operators).
Gated git staging operation.
Gated branch mutation operation.
Gated git commit operation.
Show contained repository diffs.
Show recent commit history.
Show repository status.
Index a codebase for semantic search. Creates embeddings and stores them in a vector database.
No output schemas documented for any tool. Rubric baseline requires 100% of A+ tools to have documented return types. LLMs cannot plan downstream tool calls or extract required fields without knowing output structure.
Git tools (git_status, git_diff, git_log, git_commit, git_add, git_branch) have minimal descriptions (20-50 chars). Baseline for descriptions is 194 chars average for p50 tools; these fall well below 10th percentile. Descriptions do not explain WHEN to use each tool or what distinguishes them.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
List all available code indexes with their statistics.
Query an indexed codebase using natural language. Returns relevant code chunks with metadata.
Count words in a string.
Destructive operations (git_commit, git_add, git_branch) are marked as DESTRUCTIVE risk but lack confirmation/dry-run patterns. Rubric pattern:confirmation-request states irreversible operations should support a confirmation step to prevent agent mistakes.
bash tool accepts a bare 'command' string with sandbox policy. While policy enforcement is present in implementation, the tool description does not enumerate approved commands, sandbox behavior, or what happens if a command is rejected. LLMs cannot understand constraints.
Error handling guidance is absent from tool descriptions. Rubric pattern:recovery-guide requires error responses to tell the LLM what to do next. No tool description states what errors are retryable vs fatal, or offers recovery suggestions.
query_codebase parameter 'explain' is a boolean but its semantic meaning is unclear. Description says 'Include per-result retrieval explanation' but does not define what format that explanation takes or how the LLM should use it. Vague parameter semantics cause wrong invocations.
index_codebase has no documented timeout behavior or retry guidance. Tool includes timeout_seconds parameter but does not explain what happens on timeout, does indexing resume? Rollback? LLM cannot recover from timeouts without guidance.