MCP server for SmartOnlineJudge platform providing tools for user management, question/problem management, and administrative operations
The server provides 16 tools with varying quality. Strengths: tool names follow verb_noun patterns and are mostly clear; descriptions are present for all tools and explain purpose. Weaknesses: descriptions are written in Chinese, which limits LLM comprehension outside Chinese-language models; input schemas are visible but lack type constraints and enums; output schemas are NOT documented (tools return untyped strings or raw data structures); error handling is minimal, most errors are simple string messages without recovery guidance or categorization; no pagination support despite query tools; no batch operations; critical missing pieces include documented return types, parameter constraints (enums for difficulty/risk levels), and actionable error messages.
为指定题目指定编程语言创建一个判题模板。 在调用这个工具时你需要输入三个参数: - question_id: 题目的id - language_id: 评测模板对应的编程语言的id - code: 评测模板的代码
为指定题目指定编程语言创建一个内存时间限制。 在调用这个工具时你需要输入四个参数: - question_id: 题目的id - memory_limit: 题目的内存限制(单位 MB) - time_limit: 题目的时间限制(单位 ms) - language_id: 题目的编程语言的id
创建一道新的题目。 在调用这个工具时你需要输入三个参数: - title: 新题目的名称 - description: 新题目的描述 - difficulty: 新题目的难度(只能是这三个值:easy, medium, hard) - tags: 新题目的标签列表,列表元素是标签ID
为指定题目指定编程语言创建一个解题框架。 在调用这个工具时你需要输入三个参数: - question_id: 题目的id - language_id: 解题框架对应的编程语言的id - code_framework: 解题框架的代码
为指定题目创建一个测试用例。 在调用这个工具时你需要输入两个参数: - question_id: 题目的id - input_output: 测试用例的输入输出信息。
获取当前用户的信息。 在调用这个工具的时候你不需要输入任何参数。 你可以通过这个工具来获取当前用户的详细信息。
Output schemas are not documented. Tools return untyped strings (e.g., 'query_all_programming_languages() -> str') or raw dicts without describing expected fields. LLMs cannot reliably extract data or plan downstream calls when return types are opaque.
Descriptions are written in Chinese. While clear for Chinese-speaking LLMs, this excludes models trained primarily on English text and reduces interoperability. MCP servers should provide English descriptions as primary.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
查询所有编程语言信息。 在调用这个工具的时候你不需要输入任何参数。 一个编程语言通常包括以下字段: - id: 该编程语言在数据库中的唯一标识 - name: 编程语言的名称 - version: 编程语言的版本 - is_deleted: 该编程语言是否已经不使用了
查询所有标签信息。 在调用这个工具的时候你不需要输入任何参数。 一个标签通常包括以下字段: - id: 该标签在数据库中的唯一标识 - name: 标签的名称 - score: 标签的分数。分数越高,代表该标签对应的题目的难度系数越高
查询一个题目的所有判题模板信息。 在调用这个工具时你需要输入一个参数: - question_id: 题目的id 一个判题模板通常包括以下字段: - id: 该评测模板在数据库中的唯一标识 - code: 判题模板的代码 - language: 评测模板对应的编程语言(对象类型) - id: 该编程语言在数据库中的唯一标识 - name: 编程语言的名称
查询一个题目的内存时间限制。 在调用这个工具时你需要输入一个参数: - question_id: 题目的id 一个内存时间限制通常包括以下字段: - id: 该内存时间限制在数据库中的唯一标识 - memory_limit: 题目的内存限制(单位 MB) - time_limit: 题目的时间限制(单位 ms) - language: 评测模板对应的编程语言(对象类型) - id: 该编程语言在数据库中的唯一标识 - name: 编程语言的名称
查询一个题目的详细信息。 在调用这个工具时你需要输入一个参数: - question_id: 题目的id 你可以获取到该题目的部分字段对应的信息: - id: 该题目在数据库中的唯一标识 - title: 题目的标题 - description: 题目的描述 - difficulty: 题目的难度 - tags: 题目的标签 - name: 标签的名称
查询一个题目的解题框架信息。 在调用这个工具时你需要输入一个参数: - question_id: 题目的id 一个解题框架通常包括以下字段: - id: 该解题框架在数据库中的唯一标识 - code_framework: 解题框架的代码 - language: 评测模板对应的编程语言(对象类型) - id: 该编程语言在数据库中的唯一标识 - name: 编程语言的名称
查询一个题目的所有测试用例信息。 在调用这个工具时你需要输入一个参数: - question_id: 题目的id 一个测试用例通常包括以下字段: - id: 该测试用例在数据库中的唯一标识 - input_output: 测试用例的原始输入输出信息
更新一个题目的判题模板。 在调用这个工具时你需要输入两个参数: - question_id: 这个判题模板对应的题目ID - judge_template_id: 判题模板的id - code: 判题模板的代码
更新一个题目的内存时间限制。 在调用这个工具时你需要输入三个参数: - question_id: 这个内存时间限制对应的题目ID - memory_time_limit_id: 内存时间限制的id - memory_limit: 题目的内存限制(单位 MB) - time_limit: 题目的时间限制(单位 ms)
更新一个题目的解题框架。 在调用这个工具时你需要输入两个参数: - question_id: 这个解题框架对应的题目ID - solving_framework_id: 解题框架的id - code_framework: 解题框架的代码
No input validation or enum constraints. 'create_question' accepts 'difficulty' as a free-form string despite the description stating it must be 'easy, medium, or hard'. LLMs will hallucinate invalid values. Missing enums for all enum-like parameters.
Error handling is minimal. Most tools return opaque string messages like '获取编程语言失败' (get programming languages failed) with no guidance on cause or recovery. No categorization (retryable vs user-fixable vs fatal). No error codes or structured error responses.
No pagination support. Query tools (query_all_programming_languages, query_all_tags) return all records with no limit, offset, or page_size parameters. Large result sets will blow the context window. No documented result limits.
Tool composition gaps. 'query_question_info' returns tags as an array of names only, losing the tag_id needed to relate to other question objects. Tools like 'create_question' accept a 'tags' array of IDs, but there is no way to discover valid tag IDs from the queries, requires an extra lookup or hardcoding.
Parameter descriptions lack detail. E.g., 'question_id: 题目的id' (question_id: the question's id) is circular and adds no value. Descriptions should clarify format, valid range, or when the parameter is required vs optional. No indication of whether IDs are numeric system IDs or human-readable slugs.
Destructive operations (create_*, update_*) lack confirmation or dry-run support. 'create_question' and 'update_judge_template_for_question' are irreversible. No guidance on whether these tools are safe to retry. No idempotency guarantees.
Response stripping is inconsistent. 'query_all_tags' removes 'is_deleted' and 'created_at' fields (good), but other tools return raw API responses with potentially irrelevant metadata. No documented guidance on which fields are safe to expose.