PromptLayer's MCP server exposes a curated, well-designed set of 49 domain-specific tools for prompt management, logging, evaluation, and workflow orchestration. Tool definitions are precise with typed parameters, clear intent-focused descriptions, and logical grouping around core use cases (prompts, datasets, evaluations, workflows, logging). No generic 'request_body' patterns detected. Descriptions consistently explain WHEN and WHY to use each tool. Minor issues: some descriptions truncated in schema display, a few tools lack usage context detail, and workflow/report management introduces API-internal naming (workflows vs agents, reports vs evaluations). Tool count is at the threshold of curation (49 tools, just within the 40-50 sweet spot), suggesting intentional design rather than auto-generation.
Add a single column to an existing evaluation pipeline. Use this to extend a pipeline incrementally instead of recreating the entire report. Column names must be unique within the pipeline. For column types and configuration, see https://docs.promptlayer.com/features/evaluations/column-types.
Add a request log as a row to the draft dataset version. Extracts input variables, metadata, scores, tags, prompt, and response. Requires create-draft first.
Create a dataset group. An empty draft version (version_number=-1) is created automatically. Names must be unique per workspace.
Create a dataset version by uploading base64-encoded CSV/JSON content. Processed asynchronously. Max 100MB.
Create a dataset version from request log history using filter criteria. Populated asynchronously.
Create a draft dataset version for a dataset group. Optionally copy rows from an existing version. Only one draft can exist per group.
API-internal naming inconsistencies: 'workflows' are user-facing agents, 'reports' are evaluations. Creates terminology mismatch in tool names vs. UI context.
Truncated descriptions in schema display (e.g., 'publish-prompt-template', 'get-prompt-template-raw'). Full intent documentation may be incomplete in MCP responses.
Complex nested schema for filter_group in 'search-request-logs' with inline enum definitions may be challenging for agents to construct correctly without examples.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-04-14 | C | 65 | 2025-11-25 | v1 |
Create a folder for organizing resources. Nest with parent_id. Names unique within parent.
Attach a release label to a prompt template version. Requires prompt_id (path) and body with prompt_version_number and label name.
Create an evaluation pipeline (called 'report' in the API) linked to a dataset group. The recommended approach is to add LLM assertion columns that use a language model to score each row. For all available column types, search the PromptLayer docs or visit https://docs.promptlayer.com/features/evaluations/column-types.
Create OpenTelemetry-compatible spans in bulk for distributed tracing. Each span can optionally include a log_request.
Create a new agent (called 'workflow' in the API) or a new version of an existing one. For new: use 'name'. For versioning: use workflow_id or workflow_name.
Permanently delete entities (prompts, agents, datasets, evaluations, folders). WARNING: This is destructive and cannot be undone. Use cascade=true to recursively delete all folder contents.
Delete a release label from a prompt template version.
Archive a single evaluation pipeline by ID. Prefer this over delete-reports-by-name when you have the report_id, since names can collide.
Delete a single column from an evaluation pipeline. Cannot delete DATASET columns. Surrounding columns shift left to fill the gap.
Archive all evaluation pipelines matching the given name.
Rename a folder. Name must be unique within the parent folder.
Update an existing evaluation column's type, configuration, name, or position. Use this to fix a bug in a CODE_EXECUTION script or change a column's settings without recreating the whole pipeline. Cannot edit DATASET columns.
Get paginated rows from a dataset. Each row is an array of cells with {type: 'dataset', value: ...}. Supports search via the q parameter.
Get paginated evaluation results with dataset inputs and eval outcomes. Each row has dataset cells ({type: 'dataset', value: ...}) followed by eval cells ({type: 'eval', status: 'PASSED'|'FAILED', value: ...}).
List entities (prompts, agents, datasets, evaluations, folders, etc.) in a folder. Returns root-level entities if folder_id is omitted. Use flatten=true to include all nested contents. Supports search and type filtering.
Retrieve a fully rendered prompt ready to send to an LLM. Fills in input_variables, resolves snippets, and returns provider-formatted parameters. Use label (e.g. 'prod') or version number to pin a specific version; defaults to latest. WARNING: Snippets are baked into the output — @@@snippet@@@ references are lost. Do NOT use this for editing and re-publishing prompts. Use get-prompt-template-raw instead.
Retrieve prompt template data for inspection or editing. Does not apply input variables. IMPORTANT: Set resolve_snippets=false to preserve @@@snippet_name@@@ references — this is required if you plan to edit and re-publish the prompt, otherwise snippet references will be lost. The response includes a 'snippets' array listing all referenced snippets. Set include_llm_kwargs=true to also get provider-specific parameters.
Get evaluation pipeline details including columns and configuration. Use get-report-score for the computed score.
Get the computed score for an evaluation pipeline.
Retrieve a single request's full payload by ID, returned as a prompt blueprint. Includes the prompt template content, model configuration, provider, token counts, cost, timing data, and trace_id (if the request was part of a trace). Useful for debugging, replaying requests, or extracting data for evaluations.
Get autocomplete suggestions for request log search fields. Returns possible values for a given field, optionally filtered by a prefix string. Useful for building search UIs or discovering available values (e.g. which engines, tags, or metadata keys exist in your logs). FIELDS: engine, provider_type, prompt_id, prompt, tags, metadata_keys, status, tool_names, output_keys, input_variable_keys, metadata_values, output_values, input_variable_values. For metadata_values/output_values/input_variable_values, also provide metadata_key to specify which key. Rate limited to 10 req/min.
Find all prompts that reference a given snippet. Returns prompt names, versions, and labels that use it.
Retrieve all spans for a given trace ID. Each span includes metadata and, if it generated a request log, the associated request_log_id. Useful for inspecting execution flow across multiple LLM calls in a traced operation.
Get a single agent (called 'workflow' in the API) by ID or name. Returns the agent details including full node configuration, edges, and version info. Optionally filter by version number or release label.
List all release labels for an agent (workflow). Returns each label with its name, ID, and the version it points to.
Poll for agent execution results by execution ID. Returns results when complete, or indicates still running.
List datasets with pagination. Filter by name, status, dataset_group_id, prompt_id, report_id, etc.
List evaluation pipelines (called 'reports' in the API) with pagination. Filter by name, status. Set include_runs=true to include batch runs nested under each evaluation.
List all release labels assigned to a prompt template.
List all prompt templates in the workspace with pagination. Filter by name, release label, or status.
List all agents (called 'workflows' in the API) in the workspace with pagination.
Log an LLM request/response pair to PromptLayer. Input and output must be in Prompt Blueprint format: {type:'chat', messages:[{role, content:[{type:'text', text}]}]}. Supports structured outputs, tool calls, extended thinking, and error tracking.
Move entities (prompts, agents, datasets, evaluations, folders) into a target folder. Omit folder_id to move to workspace root.
Move a release label to a different prompt version. Provide prompt_version_number to reassign the label.
Partially update an agent. Merges node changes into a new version. Set a node value to null to remove it.
Create a new version of a prompt template. Body has two required objects: prompt_template (with prompt_name, tags, folder_id) and prompt_version (with prompt_template content in chat/completion format, commit_message, metadata). Optionally assign release_labels. IMPORTANT: If the prompt uses snippets, preserve @@@snippet_name@@@ markers in the content. Do not inline snippet text — this breaks snippet references.
Rename or retag an evaluation pipeline. Provide name, tags, or both. Use this instead of recreating a misnamed pipeline.
Resolve a folder path (e.g. 'foo/bar') to a folder ID.
Execute an evaluation pipeline. Runs all columns against the dataset and produces scores. Name is required.
Execute an agent by name. Returns results synchronously, or returns immediately with an execution ID if callback_url is set for async webhook delivery.
Publish a draft dataset version by assigning it a real version number. Processed asynchronously.
Search and filter request logs using structured filters, free-text search, and sorting. Rate limited to 10 req/min, max 25 results/page. FILTER SYNTAX: Use filter_group to combine filters with AND/OR logic. Each filter is {field, operator, value, nested_key?}. Operators by field type: - String fields (engine, provider_type): is, is_not, in, not_in - Text fields (input_text, output_text): contains, not_contains, starts_with, ends_with - Numeric fields (cost, latency_ms, input_tokens, output_tokens): eq, neq, gt, gte, lt, lte, between (value=[min,max]), is_null, is_not_null - Datetime fields (request_start_time, request_end_time): is, before, after, between (value=[start,end] as ISO 8601) - Boolean fields (is_json, is_tool_call, is_plain_text): is_true, is_false - Array fields (tags, metadata_keys, tool_names, output_keys, input_variable_keys): contains, not_contains, in, not_in, is_empty, is_not_empty - Nested fields (metadata, output, input_variables): key_equals, key_not_equals, key_contains, in, not_in, is_empty, is_not_empty — requires nested_key EXAMPLES: Find GPT-4o requests: {filter_group: {logic:"AND", filters: [{field:"engine", operator:"is", value:"gpt-4o"}]}} Expensive requests: {filter_group: {logic:"AND", filters: [{field:"cost", operator:"gte", value:0.10}]}} By metadata: {filter_group: {logic:"AND", filters: [{field:"metadata", operator:"key_equals", value:"customer_123", nested_key:"user_id"}]}} Free-text search: {q: "refund policy"} Complex AND/OR: {filter_group: {logic:"OR", filters: [{field:"tags", operator:"contains", value:"prod"}, {logic:"AND", filters: [...]}]}}
Configure custom scoring for an evaluation pipeline. Specify which column_names contribute to the score, with optional custom code.
Tools like 'create-dataset-version-from-filter-params' and 'create-report' use 'additionalProperties: {}' for configuration objects, reducing schema discoverability and validation.
Optional parameters with complex interdependencies (e.g., 'create-dataset-version-from-file' vs 'create-dataset-version-from-filter-params') require agent understanding of mutual exclusivity patterns.