A FastAPI-based agent builder with LangChain integration that provides tools for AI agents including weather, git operations, interviews, vector search, and MCP client support
Agent Builder API exposes 17 tools with significant quality gaps across naming, descriptions, and schemas. While tools are registered with basic descriptions, they lack comprehensive parameter documentation, output schema definitions, and actionable error guidance. Several tools have vague or incomplete descriptions (e.g., 'calculates sum', 'Lists keys from JSON specification'). Parameter schemas are partially present but inconsistently documented. No tool annotations (readOnlyHint/destructiveHint) are visible despite clear risk classifications. The python_interpreter tool is DESTRUCTIVE but lacks confirmation/dry-run patterns. Most tools lack guidance on chaining, prerequisites, or recovery paths, critical for agent planning.
Model can provide direct answers
provides the staged code diff just before commit
Provides the code diff for a pull request
Response for the user when they greet.
Returns a list of relevant document snippets for a textual query retrieved from the internet.
Retrieve relevant info from a vectorstore that contains description about a job
Gets a value from JSON specification
Missing output schemas for 6+ tools (git_diff_tool, vectorstore_search, job_description_tool, resume_search_tool, json_spec_list_keys, json_spec_get_value). LLMs cannot plan downstream calls or extract required fields without documented return types.
Vague or generic descriptions on multiple tools. 'calculates sum' (sum_operation), 'Lists keys from JSON specification' (json_spec_list_keys), 'Response for the user when they greet' (greeting_tool) are under 50 chars and lack context for when/why to use them. Cannot answer: What does it do? When instead of alternatives? What does it return?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 38 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Lists keys from JSON specification
Executes python code and returns the result. The code runs in astatic sandbox without interactive mode, so print output or save output to a file.
Retrieve relevant info from a vectorstore that contains information on a candidate resume.
Saves the correct answer along with rating and explanation of the user answer for each question
Saves the interview programming skills
calculates sum
Returns if weather is hot or cold based on input
Provides current temperature for a given city in celsius
Retrieve relevant info from a vectorstore that contains information from Paul Graham about how to write good essays.
Returns clothing for the given temperature input
Tools with no input schema visible (git_diff_tool, vectorstore_search, job_description_tool, resume_search_tool, json_spec_list_keys, json_spec_get_value). Cannot validate inputs, cannot generate accurate schema for schema-aware clients. Parameter definitions are mandatory for production tools.
python_interpreter tool is DESTRUCTIVE but has no confirmation, dry-run, or sandbox validation patterns. Agents can execute arbitrary code and modify/delete data without explicit approval. Requires integration with pattern:confirmation-request or pattern:capability-gating.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) declared in tool definitions, despite clear risk classifications (READ_ONLY, WRITE, DESTRUCTIVE). Clients cannot gate or warn on dangerous operations.
Missing parameter descriptions. Tools like 'vectorstore_search', 'job_description_tool', 'resume_search_tool' lack any visible input parameter guidance. LLMs cannot infer what inputs are required or how to construct valid calls.
No chaining IDs in responses (inferred from source). Tools like 'temperature_tool' do not document what fields downstream tools (e.g., weather_clothing_tool) expect, forcing agents to guess or perform redundant lookups.
No error recovery guidance. Tools lack descriptions of what to do on failure (retryable? user-fixable? fatal?). E.g., 'git_pull_request_diff_tool' does not specify: What if the URL is invalid? What if the PR is private?
'directly_answer' tool name and description are vague. Does it just repeat user input? Does it reason? Should agents call it before other tools? Naming should follow verb_noun pattern (e.g., 'respond_to_query').