An MCP server for analyzing audio files using librosa.
This server exhibits significant definition quality gaps across most dimensions. While tool naming generally follows verb-noun patterns, descriptions are often vague or missing critical context. Parameter descriptions are present but frequently lack constraints, ranges, and actionable guidance. Input schemas exist but lack rigor: many parameters accept untyped float/int without bounds, and there are no enums for constrained values. Output schemas are minimally documented, most tools return file paths or floats with no explanation of what downstream tools should do with those values. Error handling is absent; there is no guidance on retryability, user-fixable vs fatal errors, or recovery steps. Security is a major concern: the server downloads files from user-supplied URLs and YouTube links without validation, sanitization, or sandboxing; there is no rate limiting, permission gating, or audit logging. The overall composition is reasonable (8 focused tools), but critical production-readiness patterns are missing throughout.
Computes the beat track of the given audio time series using librosa. The beat track is a representation of the audio signal in terms of its rhythmic content, which is useful for music analysis. The beat track is computed using the following parameters: - hop_length: The number of samples between frames. - start_bpm: The initial estimate of the tempo (in BPM). - tightness: The tightness of the beat tracking (default is 100). - units: The units of the beat track (default is "frames"). It can be frames, samples, time.
Computes the chroma CQT of the given audio time series using librosa. The chroma CQT is a representation of the audio signal in terms of its chromatic content, which is useful for music analysis. The chroma CQT is computed using the following parameters: - path_audio_time_series_y: The path to the audio time series (CSV file). It's sometimes better to take harmonics only - hop_length: The number of samples between frames. - fmin: The minimum frequency of the chroma feature. - n_chroma: The number of chroma bins (default is 12). - n_octaves: The number of octaves to include in the chroma feature. The chroma CQT is saved to a CSV file with the following columns: - note: The note name (C, C#, D, etc.). - time: The time position of the note in seconds. - amplitude: The amplitude of the note at that time. The path to the CSV file is returned.
Downloads a file from a given URL and returns the path to the downloaded file. Be careful, you will never know the name of the song.
Downloads a file from a given youtube URL and returns the path to the downloaded file. Be careful, you will never know the name of the song.
Parameter descriptions lack actionable constraints and bounds. Most numeric parameters (hop_length, start_bpm, n_octaves, fmin, ac_size, max_tempo, tightness) have NO min/max ranges, allowing LLMs to pass absurd or invalid values (e.g., hop_length=-512, n_octaves=10000, start_bpm=0.001). This violates the 'constrained-input' pattern and invites silent failures.
Free-form string parameter 'units' in beat_track should be declared as an enum, not a bare string. Description says 'Can be frames, samples, time' but schema does not enforce it, LLM can pass 'frame', 'Frame', 'time_ms' or any other value, causing silent failures or cryptic API errors.
Output schemas are undocumented. Most tools return file paths (tempo, chroma_cqt, mfcc, beat_track, download_from_url, download_from_youtube) or simple scalars, but there is NO documentation of what the LLM should do with these outputs. For file-returning tools, the file format (CSV columns, structure, data types) is not specified, making it impossible for downstream tools or the LLM to parse results correctly.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 51 | - | v1 |
Returns the total duration (in seconds) of the given audio time series.
Loads an audio file and returns the path to the audio time series Offset and duration are optional, in seconds. Be careful, you will never know the name of the song.
Computes the MFCC of the given audio time series using librosa. The MFCC is a representation of the audio signal in terms of its spectral content, which is useful for music analysis. The MFCC is computed using the following parameters: - path_audio_time_series_y: The path to the audio time series (CSV file). It's sometimes better to take harmonics only
Estimates the tempo (in BPM) of the given audio time series using librosa. Offset and duration are optional, in seconds.
No error handling or recovery guidance. Tools can fail (file not found, corrupt audio, network timeout, permission error) but descriptions provide NO actionable error guidance. An LLM that gets a failure has no idea if it's retryable, user-fixable, or fatal. This violates the 'recovery-guide' pattern.
Critical security gaps in download tools. download_from_url and download_from_youtube accept user-supplied URLs with NO validation, sanitization, rate limiting, or sandboxing. An attacker could supply a malicious URL, causing the server to download arbitrary files (path traversal), leading to DoS, data exfiltration, or execution. The server should validate URL schemes (https only), domain whitelists, file size limits, and timeouts.
Docstring errors and inconsistencies. tempo description says 'Offset and duration are optional, in seconds' but these parameters do NOT exist in the input schema, this is a copy-paste error from load(). Such inconsistencies confuse LLMs and agents.
Vague and repetitive descriptions. Multiple tools (chroma_cqt, mfcc) include the unhelpful note 'It's sometimes better to take harmonics only' without a corresponding tool to extract harmonics or guidance on when to do this. This wastes tokens and confuses the LLM.
Missing tool composition support. The load() tool performs STFT and harmonic-percussive source separation (HPSS) internally but does NOT return the harmonic/percussive components as separate outputs (the code is commented out). If LLMs need to analyze harmonics separately from percussion, they cannot, they must reload the file via a different tool. This violates the 'tool-chain' pattern and forces inefficient multi-step workflows.
No audit logging or traceability. The server provides no logging of tool calls (who called what, when, with what parameters, what happened). This violates the 'audit-trail' pattern and makes debugging, compliance, and security incident response impossible.