Automated generation of Model Context Protocol servers from Windows binaries and other source artifacts. Includes discovery pipeline, REST API, web UI, and MCP server generation.
MCP Factory presents a critical architecture issue: it is fundamentally an HTTP API microservice with async job-based processing, not a stateless request-response MCP server. The tool definitions exist in code but lack proper MCP schema registration. Most tool descriptions are present but generic or incomplete. Parameter schemas are partially visible but lack rigor. Error handling is minimal. The 'chat' tool conflates multiple concerns (semantic retrieval + tool invocation + conversation history) into a single tool. No parameter type validation or constraint documentation. Schemas are inferred from API endpoint patterns rather than explicitly registered in MCP format. This design pattern (async jobs, polling, stateful conversation tracking) violates MCP's stateless request-response model.
Upload a binary, start async discovery. Section 2-3: Analyze a binary file and discover invocable features.
Section 2.b: Analyze a file path already on the server's filesystem. Body: {path, hints?}. Returns {job_id, status_url}; poll GET /api/jobs/{id}.
Stream chat responses with semantic tool retrieval using OpenAI. Maintains conversation history and invokes selected tools.
Download generated artifacts (mcp.json, server code) for a completed generation job.
Execute an invocable tool by name with given arguments. Supports DLL import, CLI, and GUI execution methods.
Section 4: Generate an MCP server from selected invocables. Takes job_id and optional component_name, returns mcp_schema, mcp.json config, and server code.
Async job-based architecture violates MCP stateless request-response model. Tools return job_id/status_url and require polling via get_job_status, forcing multi-step sequences that break tool composition.
'chat' tool conflates semantic tool retrieval, LLM invocation, conversation history, and tool execution into a single tool. This violates single-responsibility principle and makes the tool unmodular and hard to test.
Parameter schemas lack complete type definitions and constraints. 'args' parameter in execute_tool is a bare object with no nested schema. 'messages' in chat lacks structure definition. 'selected' array in generate has no item schema.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Poll job status. Returns current progress, status, and result of an analysis or generation job.
Health check endpoint.
Missing parameter descriptions for critical fields. 'args' in execute_tool has a generic description. 'hints' in both analyze tools lack detail on what hints are expected. 'component_name' in generate has no guidance on naming conventions.
No error handling guidance. Tool descriptions do not indicate failure modes, retryable vs permanent errors, or recovery paths. Users cannot distinguish between transient failures (retry) and user-fixable errors (e.g., invalid hints).
Output schemas not documented for any tool. Descriptions do not specify return types or fields. LLMs cannot infer what analyze returns beyond 'job_id, status_url' (inferred from description text). generate description mentions 'mcp_schema, mcp.json config, and server code' but return types are not formally declared.
No idempotency guarantees. analyze and generate are destructive (create jobs, generate artifacts) but descriptions do not indicate whether repeated calls with same params produce the same result. Agents will not safely retry on failure.
Pagination and result limits not documented. get_job_status polling pattern is not ideal for MCP. No guidance on how many results download returns or whether job listings are paginated.
health tool has minimal documentation ('Health check endpoint' is 23 chars, barely above threshold) and no utility for LLM agents. Health checks belong in infrastructure, not as user-facing tools.