by kurokobo ยท v0.0.1
Speech tools including diarized speech-to-text powered by OpenAI or Azure OpenAI.
This community listing does not yet include every recommended support, privacy, pricing, and permission disclosure. Review the available package permissions before installing.
Available inside your emploidai workspace after installation.
Available inside your emploidai workspace after installation.
Tools for transcribing audio/video files using OpenAI or Azure OpenAI.
These tools are designed to provide a one-stop solution for transcribing even large files, including video files. In addition to standard transcription, it also supports the use of speaker diarization models. You can review the utterances for each speaker, replace the speaker names as needed, and output the results as text or files in various formats.
See the โถ๏ธ Demo Apps (DSL) section below for example apps using these tools.
There are two types of tools: all-in-one tools that auto-split and merge results, and step-by-step tools that run each stage separately. Note that the all-in-one tools can be slow and may hit timeouts for large/long files; if that happens, use the step-by-step tools.
Self-hosted users: If you encounter file-size errors or node/app execution timeouts at runtime, consider tuning the environment variables in the โ๏ธ Self-Hosted Tuning (Environment Variables) section below.
โ All-in-One Diarize (all-in-one)
output_format is specified.โ All-in-One Transcribe (all-in-one)
โ Split Audio (step-by-step)
โ Diarize Audio (step-by-step)
โ Transcribe Audio (step-by-step)
โ Concat Segments (step-by-step)
โ Review Speakers (step-by-step)
replace_speaker_name, and outputs text, Markdown (list/collapsible), or JSON (text or file outputs).โ Replace Speaker Name (step-by-step)
โ Format Segments (step-by-step)
๐พ Step-by-Step Transcription
๐พ All-in-One Diarization with Adjusting Speaker Names
๐พ Step-by-Step Diarization with Adjusting Speaker Names
After installing the plugin, navigate to the Tools or Plugins page and then click on the OpenAI Audio Toolkit plugin to configure it.
By clicking on the API Key Authorization Configuration button, you can set following fields to use this plugin in your app.
Service
OpenAI or Azure OpenAI.API Key
API Base URL (Optional for OpenAI, Required for Azure OpenAI)
https://your-resource.openai.azure.com).Model name or deployment name
gpt-4o-transcribe-diarize for diarization.gpt-4o-transcribe-diarize.You can add multiple authorizations for different services or accounts, and can select one of them when using the tools in your app.
If you run a self-hosted Dify instance and see file-size related errors, or node/app execution timeouts, review and tune the following environment variables.
Edit your .env file and restart the instance with docker compose down and docker compose up -d to apply changes.
UPLOAD_VIDEO_FILE_SIZE_LIMITUPLOAD_AUDIO_FILE_SIZE_LIMITNGINX_CLIENT_MAX_BODY_SIZEAPP_MAX_EXECUTION_TIMEWORKFLOW_MAX_EXECUTION_TIMEPLUGIN_MAX_EXECUTION_TIMEOUTPLUGIN_DAEMON_TIMEOUTTranscribes one or more audio/video files with speaker diarization enabled and outputs formatted text or files when output_format is specified.
Supports automatic splitting for large or long files (>25MB or >1500 seconds).
If inputs are split, each chunk is transcribed and the results are automatically merged with corrected time offsets and speaker IDs.
โ ๏ธ Performance Notice โ ๏ธ
input_files
auto_split (Optional, default: enabled)
split_audio outputs).use_silence_detection (Optional, default: disabled)
output_format is set, returns formatted text or a formatted file.plain_text, markdown_text, vtt_text, srt_text, json_text and their *_file variants.Transcribes one or more audio/video files and outputs plain text only. Supports automatic splitting for large or long files (>25MB or >1500 seconds). If inputs are split, each chunk is transcribed and the results are automatically merged.
input_files
auto_split (Optional, default: enabled)
use_silence_detection (Optional, default: disabled)
output_format (Optional, default: plain_text)
plain_text or plain_file.Returns a text message or a text file containing the concatenated transcript.
Splits audio/video files based on file size and duration limits. MP4 files with supported audio codecs may be extracted as audio.
input_files
use_silence_detection (Optional, default: disabled)
Returns one or more audio files (blobs). Files in API-native formats within limits pass through; others are transcoded and/or split.
Transcribes one or more audio files with speaker diarization enabled and outputs concatenated diarized segments as text and JSON.
Usually, use the outputs of split_audio as inputs.
input_files
split_audio outputs).Returns text and JSON messages containing:
segments: Array of diarized segments exactly as provided by the API, with the following structure:
id: Unique segment identifierstart: Segment start time in secondsend: Segment end time in secondstext: Transcribed textspeaker: Speaker identifiermetadata: Overall processing metadata
total_duration_sec: Total duration in seconds across processed filesWhen processing multiple files:
1-A, 2-B)file_1/seg_0, file_2/seg_0)Transcribes one or more audio files and outputs plain text only.
Usually, use the outputs of split_audio as inputs.
input_files
split_audio outputs).Returns a text message containing the concatenated transcript.
Concatenates an array of diarize-like outputs and normalizes segment ids and time offsets.
Usually, use the outputs of multiple diarize_audio calls as inputs.
items_json_string
segments and optional metadata.diarize_audio output shape.items_array (experimental)
segments and optional metadata.Returns text and JSON messages containing:
segments: Concatenated segments with updated id, start, and endmetadata:
total_duration_sec: Total duration across all itemsitem_count: Number of items in inputsegment_count: Total number of segmentsTo review speaker-wise utterances before replacing auto-assigned sequential speaker names with the replace_speaker_name tool, groups diarized segments by speaker and outputs text, Markdown (list/collapsible), or JSON.
Usually, use the output of diarize_audio, concat_segments, or format_segments as input.
segments_json_string
diarize_audio / concat_segments / format_segments JSON output).segments_json_file
format_segments JSON file output).output_format
plain_text, plain_file, markdown_list_text, markdown_list_file,
markdown_collapsible_text, markdown_collapsible_file, json_text, or json_file.preview_limit
0 for unlimited.Replaces auto-assigned sequential speaker names using user-provided rules and outputs text and JSON for downstream formatting.
Usually, use the output of diarize_audio, concat_segments, or format_segments as input.
segments_json_string
diarize_audio / concat_segments / format_segments JSON output).segments_json_file
format_segments JSON file output).replace_rules
from:to format (colon is not allowed in names).Speaker1:John Doe1-A:AliceReturns text and JSON messages containing the replaced segments.
Formats diarization segments into text, Markdown, VTT, or SRT.
Usually, use the output of diarize_audio, concat_segments, or replace_speaker_name as input.
segments_json_string
diarize_audio / concat_segments / replace_speaker_name text output).output_format
plain_text, markdown_text, vtt_text, srt_text, json_text or their *_file variants.See PRIVACY.md for details on data handling.
If you have any questions, suggestions, or issues regarding this plugin, please feel free to reach out to us through the following channels: