An MCP server that analyzes text for stylometric tells of AI writing, character artifacts, citation consistency, originality, and baseline writing style comparison. Provides tools for AI writing detection, academic integrity checking, and citation verification.
SignsOfAI demonstrates strong foundational quality in tool definitions. All 10 tools have clear, detailed descriptions (194-400+ characters, well above the 34-char p10 baseline). Tool names use action verbs consistently (inspect_, compare_, search_, check_, extract_, write_, measure_, analyze_). Input schemas are fully defined with typed parameters and per-parameter descriptions. However, output schemas are not explicitly documented in the visible code, return types are inferred from record definitions rather than formally declared in schema metadata. Three tools (measure_predictability, check_paraphrase) send data to external servers, which introduces a security/privacy consideration that should be more prominently flagged in descriptions. Error handling guidance is largely absent, tools do not document recovery paths or what LLMs should do on failure. The composition is excellent: each tool has a single, clear responsibility, and most accept human-friendly inputs (text, language codes) rather than opaque IDs.
Analyzes text for the stylometric tells of AI writing (English & Spanish): overused vocabulary, rhetorical crutches, syntactic tells, and low burstiness (uniform sentence rhythm). Returns an overall 0-100 "reads like AI" score, a plain-language verdict, per-category counts, document statistics, and a list of findings — each with the exact offending text, why it reads as AI, and an actionable fix. Runs fully offline; the text never leaves the machine. This is a signal, not proof of AI authorship.
Compares a document against its own reference list and reports where the two disagree: a source cited in the text that appears nowhere in the bibliography, a number cited beyond the end of a numbered list, one DOI on two different works, a malformed DOI, a publication year that has not happened yet, a duplicated entry. Works for English and Spanish, numbered (IEEE/Vancouver) and author-year (APA/MLA) styles, and returns the line of every problem. Runs FULLY OFFLINE and looks nothing up: it cannot tell you whether a well-formed reference is a real paper, only whether the document contradicts itself. That is often enough, because an invented bibliography tends to fail against itself first. Nothing is sent anywhere. A missing reference is usually a slip rather than dishonesty, and it is always the writer's to explain — the correct response to a finding is to ask them for the source.
Compares two or more documents AGAINST EACH OTHER to find copied passages — a cohort of student submissions, a draft against its sources. For each document pair it returns the overlap percentage (case- and accent-insensitive) and the actual shared passages as evidence. This is NOT a whole-internet index like Turnitin; it only compares the documents you provide, fully offline. It surfaces evidence and lets a human judge — it never accuses.
Output schemas not formally documented in tool registration. Return types are inferred from C# record definitions (CharacterReport, BaselineComparison, etc.), but MCP schema metadata is not visible in the registration code. LLMs cannot introspect the exact structure of responses without reverse-engineering from descriptions.
measure_predictability and check_paraphrase send text to external server (SignsOfAI API endpoint) but this is buried in the description's NOTE clause rather than prominently flagged. Security-sensitive operations (data transmission) should be explicitly declared upfront, not as an afterthought. Description reads: 'NOTE: unlike the offline tools, this SENDS THE TEXT to the SignsOfAI server...', a client reading only the summary might miss this.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 70 | 2026-07-28+ | v2 |
Finds REWORDED copies between two texts — same meaning, different words, including across languages (e.g. English vs Spanish) — that a literal copy check can't see. It embeds each sentence and compares cosine similarity. NOTE: this SENDS BOTH TEXTS to the SignsOfAI server to embed them (endpoint from SIGNSOFAI_API_ENDPOINT). Requires the embedding feature to be enabled on the server.
Compares one piece of writing against several earlier pieces by the SAME person, using function-word frequencies (Burrows's Delta). Returns how far the questioned text sits from that writer's centre, alongside how far each of the writer's own pieces sits from it — measured identically, so the scale is the writer's own variation rather than a threshold invented by this tool. Also returns which function words differ most, with rates per 1,000 words, and how many words are used at a rate the writer has never used them at. Runs fully offline; nothing is sent anywhere. WHAT THIS CANNOT DO: it cannot tell you who wrote something. There is no "different author" result and there must not be one in your summary either. Style moves with the assignment, the genre, the deadline, a co-author, an editor, and with a person simply getting better. A text outside the range is a reason to ask what changed; it is NEVER a conclusion, an accusation, or evidence of misconduct. The most valuable outcome is the reassuring one: a text INSIDE the range settles a suspicion, and saying so plainly is usually the most useful thing you can do with this tool. It refuses to answer on thin evidence and returns "Undetermined" instead of a number — do not work around that by rerunning with less text or by estimating one yourself.
Extracts the most DISTINCTIVE phrases from a document — long, specific, proper-noun- or number-bearing wording most worth checking on the web — and returns each with ready-made exact-phrase search links (Google, Bing, DuckDuckGo). It does NOT search the web itself; it hands you the searches to run. Offline.
Reports characters present in a text that typing does not produce: invisible/zero-width characters, letters borrowed from another alphabet to impersonate Latin ones (a Cyrillic "а" for an "a"), text direction controls, and hidden tag characters. Tools that rewrite text to defeat AI detectors insert these deliberately. Returns the exact codepoint, line and column of every occurrence, plus whether they are clustered (which ordinary copy-paste from a web page or a PDF produces) or spread through the whole document (which is what a rewriting tool leaves behind). Language-independent and fully offline. This is a checkable fact about a file, NOT proof of dishonesty and NOT a claim about who wrote the text: legitimate documents pick these up from PDFs, web pages and multilingual writing. The correct response to a finding is to ask the writer how the document was produced.
Measures how PREDICTABLE (generic) a language model finds the phrasing — its perplexity. Predictable, generic wording is common in AI writing, but formulaic human text scores predictable too and stylized AI can score varied: it is a signal, not proof. NOTE: unlike the offline tools, this SENDS THE TEXT to the SignsOfAI server to run the model (endpoint from SIGNSOFAI_API_ENDPOINT; defaults to the hosted API).
Searches the catalog of AI-writing "signs" the analyzer looks for (English & Spanish) — each with why it reads as AI and how to fix it. Useful as a reference / study aid, or to explain a finding in depth. Filter by keyword, language ("en"/"es"), and/or category (Lexical, Rhetorical, Syntactic, Statistical). Offline.
Produces the full analysis as a Markdown document a person can keep, forward to a writer, or take to an academic-integrity committee — the finished artefact rather than a summary to paraphrase. It contains the score, the signals that counted and the ones found at a rate people write at, the characters found in the file with their line and column, and the places where the document's citations disagree with its own bibliography. Checkable facts are named at the top and kept apart from the score, which is an opinion about prose. Every report prints how often this build is wrong, measured for the language actually analysed against texts written before 2022, and names the rules known to fire on human writing so the reader can weigh evidence that leans on one. Below the threshold that measurement supports, no verdict is given at all. Runs FULLY OFFLINE. The result contains material from the document, so treat it as you would the coursework itself: hand it to the person who asked, do not post it anywhere. Prefer this over paraphrasing the other tools' output when the user wants something to send, save, print or attach.
No error handling guidance. None of the 10 tools document what happens on failure or what the LLM should do next. For example, compare_to_baseline states 'It refuses to answer on thin evidence and returns Undetermined instead of a number', but the description does not explain how the LLM should interpret or act on that result. No recovery hints like 'If you get Undetermined, try again with more text' are provided.
Parameter 'earlierWork' in compare_to_baseline is documented as 'Earlier pieces by the same writer. At least ~1,400 words in total across them.' but the structure of BaselineSample[] is not visible in the provided code. What fields does BaselineSample have? Is it {title?: string, text: string}? LLMs cannot construct correct input without seeing the schema.
check_originality accepts 'documents: array' with 'optional title and its text' but the array element structure is not defined. Does it expect {title?: string, text: string} or something else? Without visible schema, LLMs must guess.