MCP Server for Baidu Digital Human (XiLing) - provides tools for generating digital human videos, managing voices, figures, and related media processing
This server exhibits significant quality gaps across naming, descriptions, and parameter documentation. While all 13 tools are explicitly registered with FastMCP and have schemas visible in the source code, the descriptions are verbose but lack clarity on WHEN to use each tool. Many parameter descriptions are minimal or generic. The server conflates concerns (e.g., generateDh123Video and generateDhVideo overlap functionally), lacks error guidance, and provides no validation hints. Output schemas are not documented, callers cannot infer what fields to expect. The verbose Chinese-language examples in descriptions do not help LLM decision-making and waste tokens. Overall, the definitions are below production baseline (median ~45-55 for community servers); this server lands in the lower half.
#工具说明:通过上传的音频文件,克隆音色。
#工具说明:简单便捷的生成数字人视频,根据真人录制的视频及选定音色,对视频分辨率等没有要求,无需人像生成,直接生产对应的数字人视频。 # 样例1: 用户输入:用fileid为xxx的视频文件,发音人ID为yyy的音色,视频的内容是"大家好,我是数字人播报的内容",生成一个数字人视频。 思考过程: 1.用户想要用视频文件来直接生成一个视频,用户只提供了视频文件ID,发音人ID,以及内容,是一个简单的视频合成需求,需要使用"generateDh123Video"工具。 2.工具需要templateVideoId,driveType,text,person,inputAudioUrl这几个参数。 3.templateVideoId是需要使用的视频文件的ID,所以值为xxx。给的播报内容是文本,所以driveType是文本驱动,text为"大家好,我是数字人播报的内容"。发音人已经提供了ID,所以person的值是yyy # 样例2: 用户输入:视频的地址是https://open-api-test.bj.bcebos.com/ae870923-2a3b-4d5e-b6a2-e44b4025647220250417_163529_trim.mp4,用发音人ID为yyy的音色,视频的内容是"大家好,我是数字人播报的内容",生成一个数字人视频。 思考过程: 1.用户想要用视频地址的文件来直接生成一个视频,用户只提供了视频文件链接URL,发音人ID,以及内容,是一个简单的视频合成需求用户没有提到,需要使用"generateDh123Video"工具。 2.工具需要templateVideoId,driveType,text,person,inputAudioUrl这几个参数。 3.templateVideoId是需要使用的视频文件的ID,所以值为xxx。给的播报内容是文本,所以driveType是文本驱动,text为"大家好,我是数字人播报的内容"。发音人已经提供了ID,所以person的值是yyy
#工具说明:根据所选数字人像ID及发音人ID,生成数字人视频。 # 样例1: 用户输入:用数字人像ID为xxx,发音人ID为yyy的音色,视频的内容是"大家好,我是数字人播报的内容",使用横屏全身的机位,视频背景用"https://digital-human-material.bj.bcebos.com/-%5BLjava.lang.String%3B%4046f6cc1e.png",开启自动添加动作,开启字幕,生成一个1080P的数字人视频。 思考过程: 1.用户想要用人像ID生成一个数字人视频,对声音,背景,字幕,分辨率等有要求,不是一个简单的数字人视频,需要使用"generateDhVideo"工具。 2.工具需要FigureId,driveType,text,person,inputAudioUrl,width,hight,cameraID,enable,backgroundimageUrl,autoAnimoji这些参数。 3.FigureId是需要使用的人像ID,所以值为xxx。给的播报内容是文本,所以driveType是文本驱动,text为"大家好,我是数字人播报的内容"。发音人已经提供了ID,所以person的值是yyy,开启自动动作,所以autoAnimoji的值为true,开启字幕,所以enabled的值为true,分辨率为1080P,拆分为width的值为1920,hight的值为1080,backgroundimageUrl的值是"https://digital-human-material.bj.bcebos.com/-%5BLjava.lang.String%3B%4046f6cc1e.png"
Output schemas not documented. No tool returns a documented schema; callers cannot infer response fields, types, or chaining IDs. For example, generateDh123Video returns MCPVideoGenerateResponse but the structure is opaque to LLMs. This forces guesswork and prevents proper response parsing.
Verbose, examples-heavy descriptions do not clarify WHEN to use each tool. generateDh123Video and generateDhVideo descriptions include multi-line Chinese examples and internal reasoning that wastes tokens without helping LLM decision-making. Description is 1000+ chars for generateDhVideo, 5x the recommended 200 chars, and buries the key distinction (123 = simple template-based, regular = full customization).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 29 | - | v1 |
#工具说明:通过上传的视频和音频文件,生成2D小样本数字人。
#工具说明:根据输入的文本和音色ID,生成音频。
#工具说明:查询123数字人视频合成进度。 # 样例1: 用户输入:查一下taskid为xxx的123数字人视频好了没有 思考过程: 1.用户想要查询taskid为xxx的123数字人视频,需要使用"getDh123VideoStatus"工具。 2.工具需要task ID这些参数。 3.task ID的值为xxx
#工具说明:查询基础数字人视频合成进度。 # 样例1: 用户输入:查一下taskid为xxx的数字人视频好了没有 思考过程: 1.用户想要查询taskid为xxx的数字人视频,需要使用"getDhVideoStatus"工具。 2.工具需要task ID这些参数。 3.task ID的值为xxx
#工具说明:查询可用的人像ID
#工具说明:查询2D小样本数字人生成状态。
#工具说明:查询文本转语音任务状态。
#工具说明:查询音色克隆任务状态。
#工具说明:查询可用的发音人ID。 # 样例1: 用户输入:我之前克隆过哪些声音? 思考过程: 1.用户想要查询可用的发音人ID,需要使用"getVoices"工具。 2.工具需要参数,isSystem,一个参数。 3.从"克隆过的"可以推测希望查询克隆发音人ID,因此参数的值为"false" # 样例2: 用户输入:我想用一个二十岁左右温柔小姐姐的声音。 思考过程: 1.用户想要查询可用的发音人ID,需要使用"getVoices"工具。 2.工具需要参数,isSystem,一个参数。 3.用户未明确指出发音人ID的来源,因此不传任何值。 4.从接口返回的内容中寻找describe中"二十岁"左右,gender中为"female"的音色,优先推荐给用户
#工具说明:根据业务类型上传所需要的文件。 # 样例: 用户输入:上传test.mp3这个文件用于声音克隆,文件在C:/Users/username/Desktop/test.mp3。 思考过程: 1.用户想要上传文件,需要使用"uploadFiles"工具。 2.工具需要参数,file,providerType,sourceFileName三个参数。 3.file:在C:/Users/username/Desktop/test.mp3路径下,名称为test.mp3的文件;providerType:声音克隆对应的值OPEN_TTS_CLONE_LITE;sourceFileName:test.mp3
Duplicate/overlapping tool pairs lack clear differentiation: generateDh123Video vs generateDhVideo, getVoices vs cloneVoice, generateLite2d creation vs getLite2dStatus query. The LLM cannot easily infer which to call for a given user intent. For example, 'generate a video with my cloned voice' could mean either tool.
Parameter descriptions are minimal or missing context. Example: 'isSystem' in getVoices has a description but no hint about what to do if the user doesn't specify. 'videoUrl' in generateDh123Video accepts both file ID and URL, the description does not clarify when to use each or how the system disambiguates. 'driveType' in generateDhVideo is enum-constrained but the description omits when TEXT vs VOICE is appropriate.
No error guidance. All tools catch exceptions and return generic 'error' field with string message (e.g., MCPVideoGenerateResponse(error=str(e))). LLMs receive stack traces or API errors with no actionable recovery guidance. For example, if a video generation fails due to invalid figureId, the LLM learns nothing about how to correct it or which tool to call instead.
No validation constraints on numeric parameters. generateText2Audio accepts speed, volume, pitch (range 1-9) but no min/max constraints in schema. LLMs can pass 0, 100, or -5 with no guidance. Similarly, generateDhVideo resolution parameters (width, height) have no bounds, LLM could request 1px or 1000000px.
Parameter naming inconsistencies reduce discoverability. getVoices uses 'isSys', generateDhVideo uses 'figureId', but other tools use 'voiceId'. Some tools use 'inputText', others use 'text'. This forces LLMs to reason about parameter names when they should reuse knowledge from earlier calls.
Status-check tools (getDh123VideoStatus, getDhVideoStatus, getText2AudioStatus, getLite2dStatus, getVoiceCloneStatus) have minimal descriptions (<30 chars). They do not explain what fields the response contains, when polling is complete, or how to interpret status values.
No idempotency guarantees documented. Upload, generate, and clone operations are likely idempotent (retry with same ID = no new record), but this is not stated. Agents do not know whether retrying is safe.