A collection of shell script-based AI tools for code analysis, debugging, chat, code review, and more. Designed as a token-efficient alternative to MCP servers, using shell scripts that are invoked on-demand rather than loaded as persistent tools.
SQUAD is a shell-based tool suite with 9 tools registered via custom shell scripts. Critical issues prevent a higher score: (1) All tools lack explicit input schemas in the source code, they rely on shell script argument parsing, not JSON Schema. This makes schemas invisible to MCP clients expecting standard schema declarations. (2) Tool descriptions are present but generic and LLM-unfriendly. Most describe the tool's purpose but omit WHEN to use it, WHAT it returns, prerequisites, and error recovery guidance. (3) Parameters are documented in shell comments and inline help, not as structured JSON Schema with types, enums, and constraints. (4) No output schemas are documented anywhere. LLMs cannot predict what fields will be returned. (5) Error handling is ad-hoc, shell scripts echo JSON error objects, but there is no pattern for error classification, recovery guidance, or actionable error messages. (6) Tool names are descriptive but not verb_noun formatted consistently (e.g. 'apilookup' should be 'lookup_api'; 'challenge' should be 'challenge_statement'). (7) The codebase is shell-native with a custom MCP-like interface, NOT a proper MCP server, it defines tools manually and does not register them via MCP's standard tool registration mechanism.
Holistic code and architecture analysis tool. Performs a senior software analyst technical audit to help engineers understand how a codebase aligns with long-term goals, architectural soundness, scalability, and maintainability.
API/SDK documentation lookup guidance tool. Provides structured guidance for looking up API documentation. Does NOT perform web searches itself - provides guidance for web search.
Critical thinking and thoughtful disagreement tool. Prevents reflexive agreement by forcing critical thinking and reasoned analysis. Does NOT call AI - it wraps prompts for critical evaluation.
Chat tool for collaborative thinking with AI models. One-to-one equivalent supporting collaborative brainstorming, validation of ideas, and well-reasoned second opinions on technical decisions.
Multi-CLI bridge to external AI CLIs. Bridges to external AI CLI tools including Claude, Gemini, Codex, and Aider. Can query single or multiple CLIs in parallel.
No JSON Schema Input Definitions. All 9 tools lack visible JSON Schema definitions. Parameters are documented in shell comments and inline help text, not as structured schema that MCP clients can parse and validate. This makes parameter discovery, type checking, and LLM-guided input construction impossible.
No Output Schema Documentation. None of the 9 tools document their return structure. LLMs cannot predict what fields to extract or what downstream tool calls might follow. This breaks tool chaining and forces LLMs to guess structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 22 | 2026-07-28+ | v2 |
Systematic code review with expert analysis. Combines deep architectural knowledge with precision analysis to deliver actionable feedback covering architecture, maintainability, performance, and correctness.
Multi-model consensus gathering tool. Consults multiple AI models independently and synthesizes their perspectives to identify points of agreement and disagreement.
Systematic root cause analysis and debugging assistance. Provides systematic root cause analysis based on investigation presented, with ranked hypotheses, evidence, and minimal fix recommendations.
Demonstration script that showcases SQUAD tools. Checks available providers, runs a demo chat query, and displays available tools.
Generic Descriptions Lack LLM Guidance. Descriptions present but generic. Most say WHAT the tool does but omit WHEN to use it, WHAT it returns, prerequisites, and error recovery. Example: 'API/SDK documentation lookup guidance tool' tells LLM the name, not when to call it instead of a web search. Descriptions should be 50 - 200 chars and answer: what, when, why, what-returns.
Non-Standard Tool Names. Names do not follow verb_noun convention. 'apilookup' should be 'lookup_api'; 'challenge' should be 'challenge_statement' or 'evaluate_statement'; 'clink' is a brand name, not an action verb. LLMs parse intent from the name first, non-standard names reduce clarity and increase confusion when many tools are available.
No Parameter Enums or Type Constraints. Parameters like 'type' (in analyze, codereview) and 'provider' (in chat, clink) enumerate allowed values in descriptions, not as JSON Schema enums. LLMs cannot validate input against constraints and may hallucinate invalid values. Example: codereview 'type' says enum values are present, but they are not enforced as schema constraints.
Weak Error Handling & No Recovery Guidance. Error responses are generic JSON objects ('status: error, content: message'). No error classification (retryable vs fatal), no guidance on next steps, no actionable error messages. Example: 'ANTHROPIC_API_KEY not set' is clear, but most errors lack this specificity. Pattern: recovery-guide.
No Tool Annotations (readOnlyHint, destructiveHint). Tools like 'analyze', 'codereview', 'debug' are read-only but lack tool annotations. Tools like 'demo' do not declare intent. MCP spec (2026-07-28) supports tool annotations, missing them reduces LLM confidence and prevents automated safety checks.
Parameter Descriptions Lack Specificity. Descriptions exist but are vague. Example: 'Model to use (default: claude-sonnet-4-20250514)' does not explain what happens if an invalid model is passed, what constraints apply, or why a user might pick a different model. Should include: format constraints, valid range, and rationale for the default.
Required Parameters Without Defaults Risk Friction. Tools like 'analyze' and 'codereview' require 'files' parameter. No guidance on whether to accept directories, globs, or single files. No defaults. Weak descriptions force LLMs to ask clarifying questions instead of providing sensible defaults.
Missing Pagination & Result Limits. No tools declare a result limit or pagination strategy. LLM context window can be exhausted if, e.g., 'chat' returns unbounded history or 'consensus' returns unlimited model responses. Tools should cap results (e.g. 20 - 50 items) and offer pagination.