A toolkit repository for scientific workflows, containing multiple skills and tools for bioinformatics analysis including cell-type annotation, peak atlas processing, and Shiny multiome visualization.
SciAgent-toolkit exhibits severe definition quality gaps across all 10 tools. While tool names are generally action-oriented (mllmct, apply_primary_rescue_filter, validate_*), the overwhelming majority of tools lack visible input schemas in the provided source code. Of the 10 tools, only 2 have reasonably detailed input parameter documentation visible (apply_primary_rescue_filter, calculate_adaptive_thresholds in R source; prepare_for_shinymultiome). The remaining 8 tools have either minimal schema documentation or exist only as references to external scripts/CLIs without explicit MCP registration. Descriptions vary widely, some are quite detailed (mllmct: 164 chars describing version-locking and reproducibility; apply_primary_rescue_filter: 269 chars), while others are minimal (check_ports: 141 chars, but vague on return format). Critical issue: tools 1, 4, 5, 6, 8, 9, 10 appear to be inferred from file references (pyproject.toml, .sh, .R, .py, install.sh) rather than explicitly registered MCP tool definitions with JSON schemas. This violates the schema rule: 'If a tool has NO input schema at all, its schema score MUST be 0.' Per-tool analysis reveals consistent structural deficits: no enum constraints, no min/max bounds on numeric parameters, no documented output schemas, no error recovery guidance.
Primary + Rescue peak filter. Adaptive per-cluster cell-count threshold: clamp(rate * n, floor, cap). PRIMARY: per cluster adaptive threshold max(floor, min(rate*n, cap)). Keep a peak if it clears the bar in ANY ONE cluster. RESCUE: keep peaks with n_strategies == 4 AND >= rescue_global_min cells globally.
Adaptive per-cluster cell-count threshold: clamp(rate * n, floor, cap).
Decision Pause S3 collision probe (shinymultiome flavour). Reads a space-separated list of TCP ports as args (or defaults to 8090). Prints FREE/TAKEN per port. Always exits 0; the agent reads stdout to decide whether to fire Decision Pause S3.
Install a local scio release tarball on the machine. Implements content-addressed installation under a prefix, with receipt tracking, atomic extraction, and reversible uninstall. Validates checksums before extraction.
Version-locked CLI for mLLMCelltype multi-LLM consensus cell-type / cell-state annotation with token capture, forced determinism, Python-recomputed consensus metrics, and a full reproducibility trace.
Phase A object preflight. Reads a Signac multiome Seurat .rds, normalises it for ShinyMultiome.UiO hosting (assay names, Annotation seqlevels, Fragment paths, group.by columns), optionally computes LinkPeaks and motif footprints, and writes a deploy-ready .rds.
No visible JSON Schema input schemas for 8 of 10 tools. Tools are inferred from external files (pyproject.toml, shell scripts, R scripts) rather than explicitly registered with MCP-compliant JSON schemas.
Tool definitions appear inferred rather than explicitly registered. Only R function signatures and CLI argument lists visible; no explicit MCP tool registration with capabilities and schemas.
No documented output schemas. Tools return unstructured data or results not documented in structured format. Callers cannot plan downstream tool use without explicit schema documentation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Substitute .env values into shinymultiome docker-compose and nginx templates. Renders docker-compose.shinymultiome.yml and nginx-shinymultiome.conf from templates by replacing <<KEY>> placeholders with .env values.
Project bindings, CRAFT rendering, and guardrail checks. Main CLI entrypoint with subcommands: craft (render SCIO:CRAFT block), link (link catalog trees and ensure hooks), lint (run project guardrail checks).
Phase B post-up smoke test. Runs after docker compose up -d. Confirms the container is healthy at multiple layers: image build, R package availability, R startup of the app, nginx reverse proxy, websocket upgrade, and basic-auth.
Pre-deploy schema check for Signac multiome objects. Runs before docker compose up to fail fast if Phase A was skipped or incomplete. Validates assay names, ChromatinAssay class, Annotation presence and UCSC seqlevels, Fragment accessibility, group.by columns, optional Links, and optional motifs + footprint enrichment.
Numeric parameters lack bounds. E.g., apply_primary_rescue_filter 'floor', 'rate', 'cap', 'rescue_global_min' have no min/max constraints documented. R function signatures show defaults (floor=15, rate=0.02, cap=200) but no validation rules or bounds for LLM-provided values.
No enum constraints for categorical parameters. Tools like 'scio' (subcommands: craft, link, lint), 'render_compose' (templates override), and 'install' (uninstall, list flags) accept specific values but do not enumerate them in schemas. LLMs must guess valid options.
No error handling or recovery guidance. Tools have no documented error cases, categories (retryable vs. fatal), or recovery steps. E.g., validate_signac_rds may fail if RDS invalid or fragments missing, but no error message or guidance provided.
Descriptions are generic or overly technical without LLM-optimized phrasing. E.g., 'check_ports' description (141 chars) mentions 'Decision Pause S3 collision probe' (domain jargon) but does not clearly state WHAT the tool does or WHEN to call it. 'render_compose' (111 chars) is vague: 'Substitute .env values into templates', why? When?
No documented parameter order, interdependencies, or side effects. E.g., 'prepare_for_shinymultiome' has 10 parameters; interdependencies (e.g., 'if compute_links=TRUE, LinkPeaks() must succeed first') not documented. 'install' has --uninstall and --list flags that likely conflict but no mutual exclusion documented.
Inconsistent naming conventions. Most tools use verb_noun pattern (validate_*, apply_*, calculate_*, check_*, prepare_*, render_*), which is good. However, 'mllmct' (acronym, no verb), 'scio' (noun-only), and 'install' lack clarity. 'install' could mean 'install_scio' to parallel the rest.
Parameter descriptions sometimes reference domain-specific jargon without expansion. E.g., 'validate_signac_rds' references 'ChromatinAssay', 'UCSC seqlevels', 'LinkPeaks()', 'motifs', 'footprints' without explaining what these mean for an LLM unfamiliar with Signac. Descriptions assume bioinformatics domain knowledge.