Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server has critical deficiencies in definition quality. Both tools have descriptions present but severely lack parameter documentation and structured output schemas. The tool naming follows some conventions (_get_pull_request, _archive_to_database) but underscore prefixes suggest internal implementation details rather than user-facing tool names. Most critically, there are NO visible input parameter schemas beyond raw Python type hints in the code. The output schemas are completely undocumented, callers have no way to know what structure to expect from these tools. Error handling exists in the implementation (try/catch blocks with logging) but produces generic responses ('{}' on error) that provide no guidance to agents. Per-tool analysis: _get_pull_request returns a dict but no documented schema; _archive_to_database takes a generic Dict[str, Any] parameter with no type constraints or validation hints.
No documented output schemas. Both tools return dict/str but no schema definition visible. Agents cannot know what fields to expect (e.g., does _get_pull_request return 'pr_id', 'pr_url', 'files'?). This violates the fundamental contract between tool and caller.
Parameter 'pr_data' in _archive_to_database is typed as Dict[str, Any] with no constraints or description of required fields. An LLM cannot know which fields are mandatory, what their types are, or what structure to provide. This invites malformed calls.
Tool names start with underscore (_get_pull_request, _archive_to_database), signaling private/internal methods rather than public API. Tool names should be verb-noun pairs without underscore prefix (get_pull_request, archive_pull_request_review).
_get_pull_request_archive_to_database
Recommendations
Document the complete return schema for both tools. For _get_pull_request, explicitly list all fields returned (pr_title, pr_description, pr_author, created_timestamp, updated_timestamp, pr_state, files_changed_count, file_diffs with sub-fields: file_path, file_status, lines_added, lines_removed, total_modifications, diff_patch, source_url, api_contents_url). For _archive_to_database, document what the success response contains (e.g., { 'status': 'archived', 'document_id': '...' }).
Remove underscores from tool names. Use 'get_pull_request' and 'archive_pull_request_review' (or 'store_pull_request_analysis' for clarity about what 'review' means).
Add structured error responses. Instead of returning '{}' on error, return { 'error': '<category>', 'message': '<details>', 'recoverable': <bool>, 'next_step': '<hint>' }. Categories: 'not_found', 'rate_limited', 'unauthorized', 'invalid_input', 'server_error'. Example: { 'error': 'not_found', 'message': 'Pull request #999 not found in owner/repo', 'recoverable': false, 'next_step': 'Verify the repository owner, name, and PR number are correct.' }
Add limit and offset parameters to _get_pull_request. Accept optional 'file_limit' (default 50, max 500) to cap the file_diffs array. Return 'files_total' and 'files_returned' counts. If files_total > files_returned, return a 'next_offset' for pagination.
Expand parameter descriptions with format/type info: 'owner: GitHub organization or username (string, 1-39 chars, alphanumeric and hyphens)'; 'repository: Repository name (string, 1-255 chars)'; 'number: Pull request number (positive integer, typically 1-10000)'; 'pr_data: Object containing PR metadata. Required fields: { files_changed_count: integer, file_diffs: array of {file_path, file_status, ...}, ... }. Unknown fields are preserved.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Error responses are non-actionable. Both tools return empty dict '{}' or generic error string on failure. LLMs receive no guidance on what went wrong, whether to retry, or what to do next. Example: retrieve_pull_request_data returns None on error, then _get_pull_request returns '{}', the agent cannot distinguish 'PR not found' from 'rate limited' from 'auth failed'.
GitHub token and MongoDB URI are loaded from environment variables in the code, but no documentation of how to configure them or what happens if they are missing. The github_fetch.py raises ValueError if GITHUB_TOKEN is missing, but this happens at module import time (not tool call time), making it hard for agents to recover.
No pagination support. If a PR has many files (e.g., 500+), the entire file_diffs array is returned inline. No limit parameter, no cursor/offset, no total_count. This violates the paginated-result pattern and risks exhausting context windows.
Tool descriptions are too generic. '_get_pull_request' says 'Obtain details from a GitHub pull request' but does not explain WHEN to use it, WHAT details are returned, or what the return schema is. '_archive_to_database' says 'Store pull request review in the database' but does not say what 'review' means or what format pr_data should be.
No idempotency guarantees. _archive_to_database performs insert_one without checking for duplicates or deduplication. If an agent retries after a network timeout, the same PR review gets inserted twice. No upsert logic, no idempotent markers (idempotentHint).
Parameter descriptions are minimal. 'owner' is described as 'Repository owner', is this a GitHub username, organization name, or numeric ID? 'number' is 'Pull request number', is it a string or int? These ambiguities force LLMs to guess or try multiple formats.
_get_pull_request
Implement idempotency for _archive_to_database. Either use upsert logic (replace if pr_title + timestamp match) or include a deduplication_key parameter that agents can use. Return the operation type in the response ('inserted' vs 'updated').
Add permission/scope declarations. Document that _get_pull_request requires 'read:public_repo' or 'read:private_repo' (depending on repo visibility) and _archive_to_database requires 'write:mongodb' (custom scope). This clarifies what credentials are needed.
Add dependency hints in tool descriptions. E.g., 'Use this tool to fetch a PR's files and metadata. If you need to search for a PR number first, use search_pull_requests (if available).' This prevents agents from making wrong assumptions.
Document timeout behavior. If GitHub API or MongoDB becomes slow, what happens? Currently, there is no explicit timeout, the tool could hang. Add 'Calls time out after 30 seconds. If GitHub or MongoDB is unreachable, an error is returned with the service name.'
Strip verbose API responses. The file_diffs array includes 'source_url' and 'api_contents_url' which are rarely needed by agents. Keep only 'file_path', 'file_status', 'lines_added', 'lines_removed', 'total_modifications', 'diff_patch' in the default response. Offer an optional 'include_urls' param if agents need the full data.