The server demonstrates solid foundation with 14 well-named tools following verb_noun conventions (authentik_apps_list, authentik_apps_create, etc.). All tools have descriptions and input schemas are visible for most. However, there are critical gaps: several tools lack complete parameter documentation, output schemas are entirely undocumented, and error handling lacks recovery guidance. The code shows thoughtful implementation (access tier gating, annotation support) but falls short of production-grade quality in LLM-specific areas like response shaping and error classification. Average tool score: 72/100.
Check whether a specific user has access to an application.
Create a new application with name, slug, and optional provider, group, and metadata.
Delete an application by its slug. This action is irreversible.
Get a single application by its slug.
List applications with optional filters for name, slug, group, search, and more.
Set an application's icon to an external URL (sets the meta_icon field), or clear the current icon. Provide either icon_url to set the icon, or clear: true to remove it.
Two tools (authentik_apps_check_access, authentik_property_mappings_by_type_get) have NO input schema visible in the provided source code.
No output schema documentation anywhere in the codebase. The rubric requires documented return types for all tools (100% of A+ tools have them). Agents cannot plan downstream tool calls without knowing what fields are returned. For example, authentik_apps_list returns paginated results, but the response structure is invisible to LLMs.
Error handling is present (sanitizeError function in src/core/tools.ts) but does not include recovery guidance. Errors are wrapped in a generic 'Error: <message>' format with no categorization of retryable vs fatal vs user-fixable. The rubric requires actionable error responses like 'Try search_users() with a partial name' instead of raw error codes.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Update an existing application. Only provided fields are modified (partial update).
Get a property mapping by type and UUID.
List property mappings of a specific type. Valid types: notification, provider_google_workspace, provider_microsoft_entra, provider_rac, provider_radius, provider_saml, provider_scim, provider_scope, source_kerberos, source_ldap, source_oauth, source_plex, source_saml, source_scim.
Delete a property mapping by its UUID. This action is irreversible.
Get a single property mapping by its UUID (cross-type).
List all property mappings across all types.
Test a property mapping by UUID. Optionally provide user, context, and format_result.
List all available property mapping types that can be created.
Destructive tools (authentik_apps_delete, authentik_property_mappings_delete) have no confirmation or dry-run mechanism. The rubric requires irreversible operations to support a confirmation step to prevent catastrophic agent mistakes. Current implementation allows immediate execution.
Tool descriptions lack dependency hints. For example, authentik_apps_check_access and authentik_property_mappings_test should explain when to call them relative to other tools. No guidance on multi-step sequences like 'Get app list first to find the slug, then call check_access.'
Parameter descriptions are sparse for some tools. authentik_apps_check_access is completely missing parameter documentation in the provided source. Others lack format guidance: 'ordering' parameter in list tools says 'Field to order by' but doesn't explain the minus-prefix convention. Parameter descriptions should be 50-100 chars for clarity.
No response field naming consistency check visible. If authentik_apps_list returns 'slug' field, tools accepting applications must also use 'slug' not 'app_slug' or 'slugId'. Mismatched naming forces LLM to reason about field mappings (pattern: response-field-naming).
List tools (authentik_apps_list, authentik_property_mappings_list) accept page/page_size but do not document the total count or next_cursor in the description. The rubric requires pagination responses to include metadata for context planning.