A collection of MCP servers for ASR transcription, Elasticsearch indexing, speaker identification with metadata analysis, and Wasabi S3-compatible storage management
This tool, add_document, is designed to add a new document to a specified index within Elasticsearch. Parameters: index_name (string, required): The name of the Elasticsearch index where the document will be added. filename (string, required): The name of the file associated with the content being indexed. This will be stored as a field within the document. file_content (string, required): The actual content of the file that needs to be indexed. This will be stored as a field within the document. Functionality: The add_document tool constructs a document with filename and file_content fields. It then attempts to index this document into the specified index_name. doc_id value taken from a filename unique ID for the document. If doc_id exist then it will return doc_id already indexed. Returns: The tool returns a dictionary with the following keys: success (boolean): True if the document was successfully indexed, False otherwise. message (string, optional): A success message if the document was indexed. error (string, optional): An error message if the document indexing failed.
This tool, create_index, is designed to create a new index within Elasticsearch. Parameters: index_name (string, required): The name you want to give to the new index. Functionality: The tool first checks if an index with the provided index_name already exists. If it does exist, the tool will return a message indicating that the index already exists and will not attempt to create it. If it does not exist, the tool will proceed to create a new index with the specified name. Optionally, a mapping (not explicitly shown in the provided snippet but referenced) can be applied during creation to define the structure and data types of documents within the index. Returns: The tool returns a dictionary with the following keys: success (boolean): True if the index was created successfully, False otherwise. message (string, optional): A success message if the index was created. error (string, optional): An error message if the index creation failed or if the index already existed.
Deletes a document from the specified Elasticsearch index using the doc_id. Parameters: index_name (string): The name of the index to delete the document from. doc_id (string): The unique ID of the document to be deleted. Returns: A dictionary indicating success or failure.
Deletes a specified file (object) from an S3-compatible storage bucket. This tool interacts directly with the S3 client to remove an object identified by its bucket name and key (path within the bucket). Args: bucket_name (str): The name of the S3 bucket from which the file is to be deleted. (e.g., "my-documents-bucket"). key (str): The unique identifier (key) of the object within the bucket. This is essentially the full path to the file within the bucket (e.g., "reports/annual_report_2024.pdf", "images/profile.jpg"). Returns: dict: A dictionary indicating the outcome of the deletion operation. - On success: `{"message": "Successfully Deleted", "status": "Success"}` - On failure: `{"message": "Internal Server Error", "status": "Failure"}` (along with error details printed to the console).
Reads and returns the content of a file from a specified S3 bucket, with specialized handling for different file types. This tool determines the file type based on its extension and uses appropriate helper functions to extract content from DOCX, PDF, and TXT files. For other file types, it retrieves the raw binary content directly from S3. Args: bucket_name (str): The name of the S3 bucket where the file is located (e.g., "my-document-repo"). file_key (str): The full key (path and name) of the file within the bucket (e.g., "reports/latest/minutes.docx", "forms/application.pdf"). Returns: str or bytes: The extracted content of the file. - For DOCX, PDF, and TXT files, it attempts to return the textual content as a string. - For other file types, it returns the raw binary content of the file as bytes. - In case of an error, it returns a string representation of the exception.
Retrieves metadata and processes comments from a given social media URL (YouTube, Facebook, or TikTok). This tool fetches video or post metadata, including comments, from the specified URL. It then analyzes the comments to identify potential 'speakers' (which could be key entities or people mentioned). Finally, it cleans and formats the top 10 comments and includes the identified speakers in the output data. Args: url (str): The URL of the social media post or video (must contain 'youtube.com', 'youtu.be', 'facebook.com', or 'tiktok.com'). Returns: dict: A dictionary containing the processing status, a message, and the 'data' payload. The 'data' includes the original metadata, a 'comments' list (top 10 cleaned comments), and a 'speakers' list (identified by the AI model). Example Success Return: {"status": "Success", "data": {... metadata ..., "comments": [...], "speakers": [...]}, "message": "Successfully processed."} Example Failure Return: {"status": "Success", "message": "Not able to fetch metadata."}
Lists all file keys (names) present in a specified S3 bucket. This client-side function interacts with an S3-compatible service to retrieve a comprehensive list of all objects (files) stored within a given bucket. It uses a paginator to efficiently handle buckets containing a large number of objects, ensuring all keys are retrieved. Args: bucket_name (str): The name of the S3 bucket from which to list files (e.g., "my-archive-data", "company-documents"). Returns: list[str]: A list of strings, where each string is the full key (path and name) of a file within the specified bucket (e.g., "folder/subfolder/document.pdf", "image.png"). Returns an empty list if the bucket is empty or if no 'Contents' are found.
Lists all document IDs (doc_id) in the specified Elasticsearch index. Parameters: index_name (string): The name of the Elasticsearch index. Returns: A dictionary containing the list of document IDs or an error message.
Generates a pre-signed URL for an S3 object, allowing temporary, time-limited access to it. This tool creates a URL that grants temporary access permissions to a specific object in an S3-compatible bucket without requiring AWS credentials. The generated URL is valid for a default duration of 1 hour (3600 seconds) but can be customized. This is useful for securely sharing private S3 objects with users who don't have direct S3 access. Args: bucket_name (str): The name of the S3 bucket where the object is located (e.g., "my-data-archive"). key (str): The unique identifier (key) of the object within the bucket for which the pre-signed URL is to be generated (e.g., "documents/report.pdf"). expires_in (int, optional): The number of seconds for which the pre-signed URL will be valid. Defaults to 3600 seconds (1 hour). Returns: dict: A dictionary containing the result of the operation. - On success: `{"message": "Successfully Generated", "status": "Success", "presigned_url": "..."}` where "..." is the generated URL. - On failure: `{"message": "Internal Server Error", "status": "Failure"}` (along with error details printed to the console).
This tool, search_by_keyword, is designed to search a specified Elasticsearch index for documents containing a particular keyword within their file_content field. Parameters: index (string, required): The name of the Elasticsearch index to search within. keyword (string, required): The keyword or phrase to search for within the file_content field of the documents. Functionality: The search_by_keyword tool constructs a match query targeting the file_content field. This query is then executed against the specified Elasticsearch index. The tool retrieves all documents that match the provided keyword in their file_content. Returns The tool returns a list of dictionaries. Each dictionary in the list represents a document that matched the search query, containing the original source fields of the document (e.g., filename, file_content).
This tool, search_content_by_filename, is designed to retrieve the content of a document from a specified Elasticsearch index by its filename. Parameters: index (string, required): The name of the Elasticsearch index to search within. filename (string, required): The exact filename of the document to retrieve. Functionality: The search_content_by_filename tool constructs a query that attempts to find documents where the filename field matches the provided filename. It uses a bool query with should clauses to search both for an exact term match on filename.keyword (for precise matching of the whole filename) and a more general match on the filename field (to handle potential tokenization or slight variations). The tool then executes this query, limiting the results to the top 1 hit, as it expects to find a unique document for a given filename. Returns: The tool returns a dictionary containing the filename and file_content of the found document. If no document is found with the specified filename, it returns a dictionary with an "error" key.
Transcribe from URL.
Uploads a given file to a specified Wasabi bucket. This function creates a temporary directory named "temp" if it doesn't already exist. It then writes the provided file_content to a file with the given filename inside this temporary directory. Finally, it uploads this local file to the specified Wasabi bucket. Args: filename (str): The name of the file to be uploaded (e.g., "my_document.txt"). file_content (str): The content of the file as a string. bucket_name (str): The name of the Wasabi bucket where the file will be uploaded. Returns: str: A confirmation message indicating the file was uploaded successfully, or an error message if an exception occurred during the process.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 0 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 12 | - | v1 |