MCP server for ontology-first RAG pipelines, exposing SPARQL query tools and ontology navigation capabilities backed by RDF data
OntoRAG MCP provides 6 SPARQL-focused tools with consistent verb-noun naming (sparql_select, sparql_construct, describe, list_by_class, outgoing, incoming). All tools have descriptions present (10-60 chars each). Input schemas are clearly defined with parameter types (string, integer) and descriptions. However, several critical gaps limit the score: (1) descriptions are quite terse (most under 60 chars), lacking the 50-200 char LLM-optimized range; (2) output schemas are not documented in code, return types are inferred from implementation (Dict[str, Any]); (3) no enum constraints on the 'accept' parameter (which should list valid MIME types); (4) limited error guidance, IRI sanitization raises ValueError but no recovery hints; (5) tool composition is reasonable but tools lack pagination and result limits are hardcoded; (6) parameter descriptions are minimal, e.g., 'RDF serialization format (MIME type)' doesn't explain valid values or defaults well enough for LLMs. The code shows deliberate input validation (_sanitize_iri) which is good, but this is not documented in tool descriptions.
DESCRIBE a resource by IRI.
Incoming edges to a resource.
List instances of a class.
Outgoing edges from a resource.
Run a SPARQL CONSTRUCT/DESCRIBE and return RDF as text.
Run a SPARQL SELECT/ASK query and return SPARQL Results JSON.
Output schemas not documented. Tools return Dict[str, Any] with no schema documentation, making it impossible for LLMs to know what fields to extract or pass to downstream tools.
Tool descriptions are too terse (30-60 chars). Best practice is 50-200 chars. Missing context on WHEN to use each tool, WHAT it returns, and HOW to interpret results. Descriptions do not explain the semantic difference between sparql_select vs list_by_class, or when to call describe vs outgoing.
The 'accept' parameter (in sparql_construct and describe) lacks enum constraint. Valid values should be explicitly listed as enum: ['text/turtle', 'application/rdf+xml', 'application/ld+json', 'application/n-triples']. Currently LLMs must guess or remember valid MIME types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 65 | 2026-07-28+ | v2 |
Input validation (IRI sanitization) is implemented but not documented. The regex pattern and error conditions should be described in the 'iri' parameter documentation so LLMs understand what IRIs are valid.
Hardcoded result limits (limit=50 for list_by_class, limit=100 for outgoing/incoming) are not exposed as tool parameters. LLMs cannot request more results if needed. Also, no offset/page/cursor parameter for pagination.
Error handling is minimal. ValueError exceptions on invalid IRIs are raised but the error messages ('Invalid IRI scheme', 'IRI contains invalid characters') are not documented in tool descriptions. No recovery guidance, should suggest 'Use describe() first to check available IRIs'.
Parameter descriptions lack format and constraint details. E.g., 'Maximum number of instances to return' (for limit) does not state the valid range (1-50? 1-1000?). 'IRI of the resource to describe' does not explain what an IRI is or show the expected format.