A local MCP safety layer for developer workflow hygiene with Git hygiene monitoring, security scanning, semantic history search, intent validation, and session management
FlowCheck has 9 tools with clear, domain-specific names (get_flow_state, set_rules, search_history, verify_intent, sanitize_content, detect_injections, start_session, end_session) that follow verb_noun convention. However, the server exhibits significant quality gaps: (1) Input schemas are not visible in the provided source code, only parameter descriptions appear in the tool docstrings, making it impossible to verify JSON Schema compliance; (2) Descriptions are present but vary in quality, some tools (e.g., get_flow_state) include rich context (~200 chars) while others (e.g., start_session) are terse (~100 chars); (3) No output schemas are documented, forcing LLMs to infer result structure; (4) Error handling is minimal, most tools catch exceptions and return generic error dicts rather than actionable recovery guidance; (5) Security tools (sanitize_content, detect_injections) lack concrete examples of what redaction/flagging looks like, leaving LLMs uncertain about the output format. Per-tool evaluation shows most tools are in the 45 - 55 range due to missing schemas and underdocumented outputs.
Detect prompt injection patterns in content. Scans for common prompt injection attack vectors including: - Instruction overrides - Role hijacking - Context manipulation - Delimiter attacks
End a FlowCheck session and generate audit summary. Closes a session and returns a summary of actions performed during the session.
Get current flow state metrics with security scanning. Returns metrics about repository health including: - minutes_since_last_commit: Time elapsed since last commit - uncommitted_lines: Total lines changed - uncommitted_files: Number of modified files - branch_name: Current Git branch - status: Health indicator (ok, warning, danger) - security_flags: Any detected security issues
Get actionable recommendations with security awareness. Analyzes repository and returns suggestions based on: - Commit frequency and change size thresholds - Branch age and main-branch synchronization - Security scan results
Sanitize content by redacting secrets and PII. Use this before including file contents in prompts or outputs. Replaces sensitive data with [REDACTED_TYPE] tokens.
Input and output schemas not visible in source code. No JSON Schema definitions found in tool registrations. Cannot verify parameter type constraints, required vs optional fields, or result structure. LLMs must infer all schema details from descriptions alone, risking misuse.
Output schemas completely undocumented. No descriptions of return types or field names. E.g., get_flow_state returns a dict with 'security_flags' field, but the structure, type (list/dict/string?), and possible values are not documented. Agents cannot reliably chain outputs to downstream tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Search commit history semantically. Find commits by meaning rather than exact keyword matching. Example: "authentication changes" finds commits about OAuth, login, tokens.
Update FlowCheck configuration thresholds. Supported parameters: - max_minutes_without_commit: Minutes before suggesting checkpoint (default: 60) - max_lines_uncommitted: Lines before suggesting split (default: 500)
Start a new FlowCheck session for audit correlation. Creates a new session ID for tracking related tool calls and actions. Returned session_id should be included in subsequent requests for audit purposes.
Validate current work against ticket requirements. Checks if code changes align with the stated ticket/task. Flags scope creep using GitHub Issues integration.
Error handling lacks actionable recovery guidance. Exception handlers return generic {'error': str(e), 'status': 'error'} responses. No indication of whether errors are retryable, user-fixable, or fatal. No suggestions for next steps (e.g., 'Try search_users() first'). Agents have no way to self-correct.
Security tools (sanitize_content, detect_injections) lack concrete output specifications. What does a redacted response look like? What fields are in a detection result? Are injection patterns returned as a list, dict, or boolean? Incomplete specs force LLMs to guess, risking misuse.
Parameter descriptions lack format/constraint specifications. E.g., set_rules accepts a 'config' object but does not document valid keys, value types, ranges, or required/optional fields. No enum constraints for status values (ok, warning, danger). LLMs may pass invalid config and get cryptic errors.
Search result limits not documented. search_history accepts 'top_k' with default 5, but no maximum is specified. No pagination guidance for large result sets. If search returns >100 commits, token cost explodes and reasoning degrades.
Session management tools (start_session, end_session) do not document session ID format, lifetime, or storage. Description says 'correlate tool calls for auditing' but does not explain where session_id is stored, how to retrieve it, or whether it persists across restarts.
Dependency hints missing. E.g., verify_intent requires a ticket_id but does not explain how to find one. No suggestion to call search_history or list_issues first. Tools should guide multi-step planning.