MCP server for querying and analyzing corpus data from MQuery, providing tools for concordances, frequency analysis, collocations, and text type analysis
Strong tool naming and schema definitions across all 8 tools. All tools have clear, action-oriented names (corpus_info, term_frequency, freqs, etc.) that start with verbs and convey intent clearly. Descriptions are present and substantive (avg ~150-200 chars), meeting the 10-1024 char baseline well. All parameters have explicit type definitions and descriptions. However, output schemas are not documented, responses are returned as plain text via httpRequest without structured schema definition, limiting downstream tool chaining. Error handling is present but minimal, soft/hard error categorization exists but recovery guidance is absent. Tool annotations (readOnlyHint, idempotentHint, destructiveHint, openWorldHint) are correctly applied to all tools, demonstrating spec alignment. Parameter descriptions are consistently detailed and include CQL domain context. No security issues detected (no secrets in parameters, read-only operations only). Composition is clean, each tool has single responsibility, though output schema documentation would strengthen composition patterns.
Calculate collocations of matching words with words that appear within a specified context window.
Retrieve concordance (key word in context) for a given CQL query
Get information about a corpus, including important information about its structure required for proper CQL queries
Calculate a frequency distribution of the first word of matching KWICs.
Retrieve frequency, instances per million (IPM), and Average Reduced Frequency (ARF) of a searched term within a corpus. The result is for all the matching entries given the query, regardless of the number of concrete matching words (n-grams).
Calculates frequencies of all the values of a requested structural attribute found in structures matching required query (e.g. all the authors via doc.author)
Output schemas not documented. Tools return plain text via httpRequest() without explicit structured response schemas. LLMs cannot determine what fields are in responses, breaking downstream tool chaining and forcing unstructured parsing of text results.
Error handling lacks recovery guidance. Tools distinguish soft vs hard errors via httpRequest() error object, but error responses return bare error messages without actionable next steps. E.g., 'corpus not found' should suggest 'Try corpus_info() to list available corpora.'
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | <=2025-11-25 | v2 |
Get available values for a text type (structural attribute) constraint
Shows the text types (= values of predefined structural attributes) of a searched term. This tool provides a similar result to the `text_types` called multiple times on a fixed set of attributes (typically: publication years, authors, text types, media
Parameter descriptions lack constraint hints. Parameters like 'q' (CQL query string) do not document valid syntax, length limits, or common mistakes. 'flimit' and 'max_items' lack min/max bounds. 'window' for collocations lacks units and range.
Some parameter defaults not justified. 'max_items' defaults to 20, 'flimit' to 1, 'attr' to 'lemma' in freqs/collocations. Descriptions do not explain why these are sensible defaults or under what conditions users should override them.