Distributed P2P scientific computing and agent orchestration system with formal verification, Python tool sandboxing, and neuromorphic/cryptographic cores
P2PCLAW MCP server has 4 tools with significant quality gaps. Tool names lack action verbs (runPythonTool, extractCodeBlocks, generateExecutionHash, storeExecutionHash), making intent unclear. Descriptions are present but minimal (58-93 chars). Input schemas are partially visible but lack comprehensive parameter documentation. The server executes arbitrary Python code with domain-based whitelisting, a high-risk operation, but error handling, permission gates, and recovery guidance are not evident in the provided source. Tool composition is reasonable (separate extraction, hashing, storage), but schema completeness and parameter descriptions fall short of production baselines.
Extract executable code blocks (python, lean4, sympy, sage) from paper markdown content
Generate SHA-256 execution hash of code and output for verification
Execute sandboxed Python code with domain-specific import whitelisting and timeout/memory constraints
Store execution hash with metadata in memory and Gun.js distributed database
Tool names lack clear action verbs; 'runPythonTool' and 'storeExecutionHash' are unclear. Per Arcade pattern:tool, names should start with verbs that reflect the action (execute_python, persist_execution_record). Generic names like 'run' and 'store' force LLMs to reason about intent rather than inferring it from the name alone.
No output schemas documented for any tool. Critical gap per pattern:tool and baseline (100% of A+ tools have documented return types). LLMs cannot plan downstream calls or extract required data (e.g., does generateExecutionHash return {hash: string} or {hash: string, timestamp: number}?). Tool composition breaks without known output structure.
Arbitrary Python code execution (runPythonTool) is a high-risk operation, but no permission gate, audit logging, or confirmation step is evident. Per pattern:permission-gate and pattern:confirmation-request, destructive/sensitive tools must verify authority and optionally require confirmation. No dry-run mode visible.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Error handling is not visible in the provided source. Per pattern:recovery-guide, error responses must tell the LLM what to do next (e.g., 'Domain not whitelisted for physics, try domain: mathematics'). No indication of retryable vs. fatal errors, invalid input constraints, or corrective actions.
Parameter descriptions lack format/range constraints. E.g., timeout is 'Timeout in milliseconds (default 60000)' but does not specify bounds (min/max). Per pattern:constrained-input, numeric parameters should include min - max ranges (e.g., 'timeout in milliseconds (1000 - 300000, default 60000)'). Domain enum is clear, but others are vague.
storeExecutionHash metadata parameter type is under-constrained. Listed as 'object' with properties listed but no 'required' array. LLMs cannot determine which fields are optional vs. mandatory (is 'success' required or optional?). Should be strict JSON Schema with required array and additionalProperties: false to prevent silent data loss.
Tool composition lacks chaining ID guidance. Per mxe:include-chaining-ids, if runPythonTool execution result needs to be stored via storeExecutionHash and hashed via generateExecutionHash, the runPythonTool response should include an 'execution_id' or 'request_id' that ties the steps together. Without it, the agent has to reason about sequencing.