MCP server for local AI music generation via ACE-Step 1.5, with optional vocal/instrumental stem splitting
SongForge is a well-structured music generation MCP server with strong domain-specific documentation. Most tools (11/12) have detailed descriptions (100-400+ chars) that explain what they do, when to use them, and key prerequisites. Input schemas are present and mostly complete with type declarations. However, there are notable gaps: several tools lack comprehensive parameter descriptions (e.g., advanced_settings in generate_vocal_track is an object with no nested schema), some parameters are under-described, output schemas are documented only in prose rather than formal JSON Schema, and several tools mix read/write operations (e.g., generate_vocal_track_takes is mentioned but not defined in source). The server demonstrates strong domain expertise in music generation and includes helpful context about ACE-Step quirks, reference audio risks, and idempotency guarantees, this is above-average for a specialized domain server. Naming is generally clear (verb_noun pattern, e.g. analyze_reference_audio, transcribe_instrumental_to_midi) with one exception: generate_vocal_track_takes is referenced but source is incomplete, cap at 50.
Starts measuring real BPM and musical key from an audio file (tempo tracking + key-profile correlation) — not a guess. Returns {"job_id": str} immediately; poll check_vocal_track_status(job_id) exactly as for generate_vocal_track (same tool, same registry). Usually finishes in a few seconds, but duration scales with the input file's length with no fixed cap, so this still goes through the job/poll pattern rather than blocking. Feed the completed result into generate_vocal_track's advanced_settings, e.g. {"BPM (Beats Per Minute)": result["bpm"], "Key": f"{result['key']} {result['mode'].title()}"}. Does NOT detect chord progression — that needs a real chord chart (identify the song, look it up via web search against Chordify/Ultimate-Guitar/etc., then put it directly in the caption, e.g. "chord progression: Gb - Db - Ab - Bbm") since ACE-Step has no chord-sequence input. Useful as a key cross-check, or when no chord chart exists (e.g. private/unreleased audio).
Moves a file this server previously produced (a full-mix render or a stem) out of the output folder and into a `.trash` subfolder — NOT a permanent delete. Deliberately reversible: the calling model (or a manipulated/injected instruction) being wrong about what's safe to remove should never be able to permanently destroy a generation. To actually reclaim disk space, empty `.trash` manually via File Explorer — there is no tool for that, on purpose.
Starts trimming and/or fading a file this server previously produced — writes a NEW file, never overwrites the original. Returns {"job_id": str} immediately; poll check_vocal_track_status(job_id) exactly as for generate_vocal_track (same tool, same registry).
generate_vocal_track's advanced_settings parameter is typed as 'object' with no nested schema or enum constraints for valid ACE-Step setting names. LLMs cannot know valid keys (e.g. 'BPM (Beats Per Minute)' vs 'BPM' vs 'beats_per_minute'). Description provides one example but does not enumerate all valid overrides.
No formal output schemas in JSON Schema format. Descriptions include prose summaries of return types (e.g. 'Returns {"job_id": str}' or 'on completion it returns vocals_path/instrumental_path') but these are not machine-parseable and lack field descriptions. LLMs must infer structure from text.
Several parameters lack descriptions or have minimal descriptions: remix_source_path, remix_melody_retention, remix_no_fsq in generate_vocal_track are described only briefly; extra_stems in split_vocal_stems links to Demucs model filenames without explaining when/why to use them (e.g. 'htdemucs_6s.yaml' vs 'htdemucs_ft.yaml', what is the difference?).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 61 | 2026-07-28+ | v2 |
Start generating a full EDM/vocal track (music + vocals together) via ACE-Step 1.5. Returns {"job_id": str} immediately — does not wait for generation to finish. Poll check_vocal_track_status(job_id) to follow it through; see that tool's own docs for how sparse updates should be. Renders a complete song, not an isolated vocal — the default, complete deliverable. Leave split_stems at its default False unless the user's current request explicitly asks for stems — do not set it True on your own judgment, and do not carry it over from an earlier request in the same conversation. A plain "generate a track" request is complete once the track exists; the finished mix auto-plays via the OS's default player right after this job completes, which splitting stems first would delay. Also don't call split_vocal_stems unless the user explicitly asks for the stems.
Returns the actual note data (pitch, start, end, velocity) from a MIDI file this server produced — real ground truth, not a summary. Use this before recreating, describing, or importing a transcribed MIDI file's content anywhere: neither transcribe_instrumental_to_midi's own result nor Reaper's own MIDI tools expose real note content from an external file, so without calling this there is no actual data to work from, only the file path and a note count — attempting to "recreate" the notes without this produces fabricated content, not the real transcription. Paginated (offset/max_results) since a real transcription can have hundreds of notes — call again with a higher offset for more.
Lists finished tracks sitting in this server's output folder, newest first — for when a past generation's job_id has been lost (jobs are in-memory only and don't survive a server restart) or the user asks what's already been made. Does not include split-out stems (see split_vocal_stems) — only full-mix renders.
Lists recent background jobs (from generate_vocal_track, generate_vocal_track_takes, split_vocal_stems, analyze_reference_audio, or transcribe_instrumental_to_midi), newest first — for when a job_id has been lost mid-conversation. Jobs are in-memory only: finished/errored jobs older than 1 hour are evicted, and none survive a server restart.
List audio-separator's available vocal-separation models, sorted by vocal SDR score descending. Use this to pick or suggest a specific model for split_vocal_stems's `model` param when the default ("vocals_mel_band_roformer.ckpt") or its usual alternative ("model_bs_roformer_ep_368_sdr_12.9628.ckpt") aren't separating cleanly — pass any "filename" value from this list's results directly to split_vocal_stems's `model` param as a literal string. Fast, synchronous, no job polling needed (typically ~1s, no GPU/audio processing involved).
Plays a file this server previously produced, on demand. Every completed `generate_vocal_track` already auto-plays its result — this tool is the fallback for when that didn't work or wasn't visible (e.g. only flashed in the taskbar), or when the user asks to hear an earlier result again later in the conversation.
Checks whether reference vocal clips already exist for a named voice, or prepares a new one from a YouTube link or a local audio file already on disk. Call with ONLY voice_name first (no youtube_url/local_audio_path) whenever the user wants to use a particular voice. Two possible results: - {"status": "found", ...} — includes total_duration_seconds and meets_recommended_minimum. One song typically yields only 1-2 minutes of actual vocal once instrumental sections are excluded, so a single clip essentially never meets the recommended minimum (currently 600s) — if meets_recommended_minimum is False, tell the user how much they have and that more clips (ideally from different songs) are needed, don't treat one clip as done. - {"status": "not_found", "message": ...} — none exist yet. Ask the user for either a YouTube link or a local audio file path for a song featuring this voice (recommend an acoustic version if using YouTube, since sparser instrumentation separates into a cleaner vocal). Set the expectation up front that multiple clips will likely be needed, not just one. Do not search for or guess a link/path yourself. Once the user gives a youtube_url OR a local_audio_path (exactly one, not both), call this again with voice_name plus that one source to actually prepare it — this returns {"job_id": str} immediately (downloads if from YouTube, separates the vocal either way, and saves it into this voice's own library folder, clearly labeled with both the voice name and source). Poll with check_vocal_track_status(job_id) exactly as for generate_vocal_track (same tool, same registry); on completion it returns the same found-style status (clips, total_duration_seconds, meets_recommended_minimum) so you can tell the user whether to keep going or stop.
Only call this when the user's current request explicitly asks for the track to be split into stems — never proactively, and never just because it was used earlier in the conversation. A plain "generate a track" request is complete once generate_vocal_track finishes; do not follow it with this tool on your own judgment, and do not assume which stems are wanted (vocals/instrumental only, or also drums/bass/guitar/piano via extra_stems below) — ask, if the user's request didn't already say. Start splitting a full mix into a vocals-only stem and an instrumental-only stem via audio-separator. Returns {"job_id": str} immediately — poll check_vocal_track_status(job_id) exactly as for generate_vocal_track (same tool, same registry); on completion it returns vocals_path/instrumental_path instead of audio_path. Idempotent per (audio_path, model) pair — splitting the same file with the same model again returns the existing stems instantly instead of re-running; a different model always re-runs.
Starts converting an instrumental audio file into MIDI via basic-pitch, then splitting it into separate bass/melody/chords tracks by pitch register and note-overlap density — NOT real per-instrument-class separation (can't tell a pad from a saw-stack), but genuinely separate, DAW-assignable parts instead of one flat blob of every note layered together. No drums/percussion captured, pitched content only. Import the three split tracks as separate Reaper tracks, not the flat one — that's the actual usable starting point for a human to re-orchestrate. Returns {"job_id": str} immediately; poll check_vocal_track_status(job_id) exactly as for generate_vocal_track (same tool, same registry). Usually finishes in a few seconds, but duration scales with the input file's length with no fixed cap, so this still goes through the job/poll pattern rather than blocking. Idempotent — calling it again on a file already transcribed returns the existing MIDI files instantly instead of re-running the model, so if you've lost track of an earlier result it is always cheap and safe to call this again rather than assuming nothing exists yet.
Mutual exclusivity constraints are documented in prose (e.g. 'Mutually exclusive with reference_audio_path') but not enforced via structured parameter dependencies or validation. LLMs may pass both parameters despite the warning.
Error handling is described in prose (tool docstrings mention 'SongForgeMCPError' and error codes from songforge_mcp_shared.error_codes) but no formal error response schema is documented. LLMs cannot know what fields are in an error response or how to classify errors as retryable vs fatal.
Tool generate_vocal_track_takes is mentioned in list_recent_jobs output docs ('from generate_vocal_track, generate_vocal_track_takes, ...') but the tool definition is not visible in provided source. If it exists, it may be in a separate file not included; if inferred, cap score at 50.