Demonstration of concepts for MCP security with mTLS and Biscuit cryptographic authorization tokens
MCP-Biscuit-PoC has 8 tools with varying quality. Most tools have descriptions and input schemas present, but descriptions are often generic or lack LLM-optimized guidance. Schema quality is moderate, most parameters have types and descriptions, but several lack critical constraint documentation (enums, ranges, formats). Error handling and security considerations are largely absent from tool definitions. The server demonstrates a proof-of-concept approach rather than production-grade tooling. No tool annotations (readOnlyHint, destructiveHint) are visible despite the presence of both read-only and write operations.
Get schema information about a database table including column definitions
Execute a SQL query against the connected database
Generate data visualization recommendations and code for query results
Get the status of database connections
Get detailed information about a specific HIPAA regulation
List HIPAA regulations from eCFR with filtering and pagination
Convert natural language description to SQL query
Missing tool annotations (readOnlyHint, destructiveHint, idempotentHint). The execute_query tool marked as WRITE risk but has no destructiveHint annotation to warn LLMs of potential data modification. This increases risk of unintended destructive operations.
Insufficient error handling guidance. Tool descriptions do not explain failure modes, recovery paths, or when to retry. For example, execute_query lacks guidance on SQL syntax errors, connection failures, or timeout handling.
execute_query parameter 'query' has no input validation or constraint documentation. LLMs cannot infer that SQL injection is a risk or that certain query patterns are forbidden. Description lacks any mention of parameterized queries or safe patterns.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Search HIPAA regulations by keyword with full-text search
Output schemas are not documented in tool definitions. Agents cannot plan downstream operations because they don't know what fields to expect. For example, list_regulations returns structured data, but the response format is not declared.
Descriptions for execute_query and natural_language_to_sql are too generic (50-55 chars). They lack WHEN to use, prerequisites, or expected outcomes. LLMs cannot reliably distinguish when to call these vs. similar tools.
Pagination parameters (limit, offset) in list_regulations lack documented constraints. No mention of maximum limits, default pagination behavior, or total count in responses. LLMs may request excessively large result sets.
natural_language_to_sql 'context' parameter is vague (object type, optional, generic description). No specification of what keys/fields are expected or how schema metadata should be structured. LLMs will struggle to construct valid context.
generate_visualization 'visualization_type' enum is present and well-constrained, but the tool lacks guidance on when each type is appropriate or how query result structure determines suitability.
No security-related parameter descriptions. execute_query accepts raw SQL but does not document input sanitization, parameterized query requirements, or whether user-supplied SQL is validated server-side. Risk of SQL injection is not communicated.
Tool composition issue: search_regulations and list_regulations are nearly identical in purpose (both return regulation sets with filtering), but distinction is unclear. LLMs may conflate them or use incorrectly.