Per-sentence AI-text detection with pattern diagnostics — CLI and MCP server.
The server defines 5 tools with complete input schemas and descriptive docstrings. Tool naming follows verb_noun convention (detect_*, get_*, list_*) and is clear. Descriptions are well-crafted and include context about model size, first-run behavior, and output structure. However, there are gaps in output schema documentation, error handling guidance, and some parameter descriptions lack format constraints. The server is well-intentioned and production-adjacent, but falls short of A-grade rigor on schema completeness and error recovery patterns.
Analyse two texts and report both scores. Args: text_a: First text to analyse. text_b: Second text to analyse. model: 'desklib' (default) or 'light'. device: 'auto', 'cpu', or 'cuda'.
Detect AI-generated text in a local file (the server reads it, so large drafts never enter the client's context window). Args: path: Absolute path to a UTF-8 text file. model: 'desklib' (default) or 'light'. device: 'auto', 'cpu', or 'cuda'. include_all_sentences: Also return every scored sentence, not just AI-flagged ones. First-run note: see detect_ai_text — call get_model_status first if the model may still need downloading.
Detect AI-generated text, sentence by sentence. Args: text: The text to analyse. model: 'desklib' (default, most accurate, ~1.7GB) or 'light' (small ONNX, ~126MB). device: 'auto' (GPU if available, else CPU), 'cpu', or 'cuda'. include_all_sentences: Also return every scored sentence, not just AI-flagged ones. First-run note: if the model isn't downloaded yet this call waits for the background fetch from Hugging Face (the default is ~1.7 GB). Call get_model_status first so you can warn the user about the wait. Returns an overall AI score, verdict, per-sentence AI-flagged findings, style pattern totals, and sentence-length variation (SDSL) metrics.
Check model download status and size. Args: model: Model to query (default: 'desklib'). Returns a one-line status dict without any blocking I/O: cached / downloading / absent. Safe to call before detect tools to warn the user about long waits. Model startup (downloading, importing torch) runs in a background thread — this tool returns instantly.
Output schemas are not formally documented in tool definitions. The docstrings describe what is returned (e.g. 'overall AI score, verdict, per-sentence findings, style patterns'), but input/output schema objects are not explicitly declared in the FastMCP tool registration. LLMs cannot plan downstream reasoning without knowing the exact field names and types of returned objects.
Model parameter accepts 'desklib' or 'light' but is defined as a free-form string with no enum constraint. This invites hallucinated values. Should use JSONSchema enum: ['desklib', 'light'] to be self-documenting.
Device parameter lacks enum constraint. Description says 'auto', 'cpu', or 'cuda' but schema is a free string. Should be enum: ['auto', 'cpu', 'cuda'] to prevent LLM from passing invalid values like 'mps' or 'gpu'.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 60 | 2026-07-28+ | v2 |
List available detection models with their download status. Returns a list of all available models, their descriptions, sizes, and whether they are cached locally.
No error handling guidance in tool descriptions. If a model download times out, or a file is unreadable, or the path is invalid, the error response should guide the LLM on what to do next (retry with timeout override, check path exists, etc.). Currently no recovery_guide patterns present.
Tool description for detect_ai_file mentions 'large drafts never enter the client's context window' but does not specify a maximum file size. LLMs need to know when this tool is appropriate vs when it may fail due to resource limits.
No idempotency or confirmation pattern documented for detection tools. While these are read-only operations (low risk), the server does not declare that repeated calls are idempotent, an explicit statement would help agents make confident retry decisions.