Two tools with complete input schemas and descriptions, but both have significant gaps in documentation quality and parameter descriptions. Tool naming follows verb_noun convention (create_, get_), which is strong. However, descriptions are functional but lack depth regarding when to use each tool, error recovery guidance, and parameter format expectations. Input schemas are well-structured (Pydantic models with types), but parameter descriptions within the schema are inconsistent in quality. The create_llamator_run tool has a complex nested schema but several parameters lack actionable constraints (e.g., 'Optional API key' doesn't explain format or security implications). The get_llamator_run tool is simpler but has minimal documentation on expected output structure. Neither tool includes documented output schemas, which violates pattern:tool requirements. Error handling is present in implementation (_await_job_completion timeout, _build_error_notice) but not surfaced in the tool description to guide LLM recovery.
Tools (2)
create_llamator_runwriteauthsource verified68/100
Create a LLAMATOR job and return the aggregated result after completion.
Tool descriptions lack actionable recovery guidance. 'Create a LLAMATOR job and return the aggregated result after completion' does not explain timeout behavior, retry eligibility, or what to do if the job fails.
Output schemas are not documented in tool descriptions. LLMs cannot infer the structure of aggregated results, error objects (LlamatorJobInfo.error), or whether pagination is involved. The response type LlamatorRunToolResponse exists in code but is not visible in the tool definitions provided.
Parameter 'api_key' in tested_model is marked optional with description 'Optional API key' but does not clarify security implications (server-side injection, why it's optional, format/length constraints). Credentials as parameters violates pattern:secret-injection, this should use environment-based configuration, not tool parameters.
create_llamator_run
Recommendations
Add explicit output schema documentation to both tools. Example: 'Returns LlamatorRunToolResponse with fields: {status: str, job_id: str, result: {aggregated: dict[str, dict[str, int]]}, error: {error_type: str, message: str} | null}'
Move api_key out of tool parameters. Instead, configure OpenAI/compatible endpoints via server environment (LLAMATOR_MODEL_API_KEY, etc.) and reference by endpoint name in tool parameters (e.g., 'model_config_name': 'openai-test'). Document in tool description: 'Model credentials are configured server-side; pass the model config name, not the API key.'
Upgrade 'preset_name' from description hint to formal enum schema constraint. List valid presets: ['all', 'rus', 'owasp:llm01', ...]. Include in description: 'Built-in preset name from the list of enum values.'
Add timeout and retry guidance to create_llamator_run description: 'Polls job status for up to [timeout_seconds] seconds (default 300). Timeout errors are retryable; call again with the same job_id to resume waiting. Job completion status: SUCCEEDED or FAILED.'
Add error recovery examples to both tools: 'If job not found, call with a different job_id or create a new run. If timeout occurs, call get_llamator_run(job_id) to check status without re-submitting. If results are incomplete, check run_config.enable_reports and artifacts_path settings.'
Document 'artifacts_path' behavior: 'Relative path where LLAMATOR stores artifacts (reports, logs). Directory created automatically if missing. Path must be relative (no absolute paths, no /. or ../ traversal). Artifacts retrievable via [artifacts_download_url if enabled].'
Nested parameter descriptions are sparse. 'debug_level' is described as 'LLAMATOR log verbosity (0=WARNING, 1=INFO, 2=DEBUG)', good, but 'artifacts_path' only says 'Relative path inside server artifacts root' without clarifying what happens if the path doesn't exist, whether it's created, or what paths are invalid.
'preset_name' accepts string values with enum hint 'e.g. all, rus, owasp:llm01' in description. This should be a formal enum constraint in the schema, not an example. LLMs may hallucinate preset names like 'all_extended' or 'owasp:llm02' if not restricted by schema.
Error classification not documented. Tools do not explain whether timeout errors are retryable, what constitutes user-fixable vs fatal errors, or how to distinguish 'job failed' from 'job not found'.
No documentation of result limits. If aggregated results contain hundreds of test cases, LLMs will struggle with context bloat. Rubric baseline: cap results at 20-50 items and document the limit in tool description.
'model_description' parameter lacks clarity on purpose. Is this a user-facing label, used for internal logging, or passed to the model? Ambiguity forces LLMs to guess.
create_llamator_run
Clarify 'model_description' purpose in parameter description: 'Optional human-readable label for the tested model (e.g., "GPT-4 production v1.2.3"). Used in reports and logs for identification. Not sent to the model itself.'
Add pagination/limit guidance to get_llamator_run if results are large: 'If aggregated results exceed N items, only the top N by severity are returned. Call with limit parameter to adjust, or download full artifacts archive for complete results.'
Add to both tool descriptions: 'Requires Redis and ARQ worker running. Verify worker availability before calling; returns error if job queue unavailable.'
Document system_prompts parameter format: 'List of strings, each a system prompt to be tested. Empty list or null means use default prompts. Max N prompts per run (server limit enforced).'
Add security note to create_llamator_run: 'This tool initiates arbitrary LLM calls to the tested model. Ensure tested_model configuration uses trusted, authorized endpoints only. API keys and credentials are NOT logged by this tool, but request parameters may be audited.'