A DBSCAN-Based Automation for Speech Onset Detection and Voice Activity Detection
Download Praditor | English · 中文 | Our Paper
Praditor is a speech onset detector that automatically finds boundaries between silence and sound. It supports two detection modes and generates output in .TextGrid format — ready for use in Praat.
Tip
Need automatic time & content annotation? Check out Praasper — a VAD-Enhanced Automatic Speech Annotation pipeline for Psycholinguistic Research.
- Onset/Offset Detection (Default Mode) — Detects the start and end of sound events using DBSCAN clustering and first-derivative thresholding. Outputs
PointTierlayers for onsets (blue) and offsets (green). - Voice Activity Detection (VAD Mode) — Detects speech segments and outputs an
IntervalTierwith "sound" intervals. VAD mode uses fixed kernel parameters and processes audio in 15-second segments with intelligent boundary detection. - Batch Processing —
Run Allprocesses every audio file in the current folder, generating.TextGridand CSV summary files. - Parameter Persistence — Three save modes (Default / Folder / File) with priority-based loading. VAD mode parameters are stored separately with
_vadsuffix. - Parameter History — Navigate through up to 10 previous parameter sets with
Backward/Forwardbuttons. - CSV Export — Automatic per-folder CSV summary with timestamps for all processed audio files.
- Real-time Preview —
Testbutton estimates onset/offset count without modifying.TextGridoutput.
Try test_audio.wav and test_audio_mp3.mp3 on Praditor.
test_audio_mp3.mp3is from an online resource, whiletest_audio.wavandtest_large_audio.wavare from our own experiments.
| Button | Action |
|---|---|
File |
Import audio files (.mp3, .wav, .ogg, .aac, .flac, .amr, .wma, .aiff) |
? (Help) |
Open documentation |
Run |
Run detection on the current file |
Run All |
Batch process all audio files in the current folder |
Test |
Preview how many onsets/offsets would be detected with current parameters (does not modify .TextGrid) |
Stop |
Cancel running detection |
Trash |
Clear displayed annotations (does not modify .TextGrid) |
Read |
Reload annotations from .TextGrid file |
Onset |
Toggle onset annotations (blue) |
Offset |
Toggle offset annotations (green) |
◀ / ▶ |
Navigate to previous / next audio file |
Displays waveform with onset/offset markers and time labels. Supports zoom and scroll via mouse, keyboard, and touchpad.
Wheel ↑/Wheel ↓— Zoom amplitudeCtrl+Wheel ↑/Wheel ↓— Zoom timelineShift+Wheel ↓/Wheel ↑— Scroll timelineF5— Play / pause audio; any key to stop
- Vertical scroll — Zoom amplitude
- Horizontal scroll — Zoom timeline
- Horizontal swipe — Scroll timeline
Timeline zoom may not work on macOS touchpad. Use
Command + I/Command + Oinstead.
Nine parameters for onset (blue) and nine for offset (green), independently adjustable:
| Parameter | Description |
|---|---|
Threshold |
Coefficient for actual threshold (baseline × coefficient) |
NetActive |
Accumulated net count of above-threshold frames required |
Penalty |
Penalty for below-threshold frames |
RefLen |
Length of reference segment for baseline calculation |
KernelFrm% |
Percentage of frames retained in the kernel |
KernelSize |
Kernel size in frames |
EPS% |
Neighborhood radius in DBSCAN clustering |
LowPass |
Lower cutoff frequency of bandpass filter |
HighPass |
Higher cutoff frequency of bandpass filter |
In VAD mode, offset sliders are hidden and Penalty, RefLen, KernelFrm%, KernelSize are fixed to VAD-optimized defaults.
| Button | Action |
|---|---|
Default |
Use application-level default parameters |
Folder |
Use folder-level parameters (params.txt / params_vad.txt in current folder) |
File |
Use file-specific parameters (same name as audio file, .txt extension) |
Save |
Save current parameters to selected mode(s) |
Reset |
Reload parameters from selected mode's saved file |
Backward / Forward |
Navigate parameter history (up to 10 sets per mode) |
VAD |
Toggle VAD mode |
1/10 |
Current parameter index / total |
Priority: File > Folder > Default. When multiple modes are selected, Save writes to all selected locations; Reset loads from the highest-priority location that exists.
Basic understanding is enough. Understanding the algorithm is better.
- Basic knowledge: Go to the first section of Quick Fix.
- Advanced knowledge: Go to the second section of Quick Fix (Detailed Introduction).
- Expert knowledge: Go to Parameters.
If you would like to download the datasets that were used in developing Praditor, please refer to our OSF storage.
If you use Praditor in your research, please cite the following paper:
Liu, Z., Yu, X., Hu, W.C. et al. Praditor: A DBSCAN-based automation for speech onset detection. Behav Res 57, 247 (2025). https://doi.org/10.3758/s13428-025-02776-2
Or download .ris from the paper's About this article page.
Shout out to these remarkable contributors!
- Thank YU Xinqi, Dr. MA Yunxiao, ZHANG Sifan for their work in validating the effectiveness of Praditor's algorithm.
- Thank HU Wing Chung for her work in packaging Praditor for macOS (arm64 and universal2).
- Thank Prof. ZHANG Haoyun (University of Macau) and Prof. WANG Ruiming (South China Normal University) for their guidance and support.
Also, the funding:
- This project was funded by the National Natural Science Foundation of China (32200845), the Science and Technology Development Fund, Macao S.A.R (FDCT, 0153/2022/A), and the Multi-Year Research Grant (MYRG2022-00148-ICI) from the University of Macau to Haoyun Zhang.
Praditor is written and maintained by Tony, Liu Zhengyuan from the Centre for Cognitive and Brain Sciences, University of Macau.
If you have questions about using Praditor, its algorithm details, or need custom scripts (audio export, Excel tables, etc.), feel free to contact me at zhengyuan.liu@connect.um.edu.mo or paradeluxe3726@gmail.com.
Praditor is dual-licensed under AGPL v3 + a commercial license (LICENSE):
- AGPL v3 (default): free, open source. Academic / personal / non-profit / small orgs can use it directly. Only requirement: if you offer Praditor as a network service, you must make the source available.
- Commercial License: if you cannot accept AGPL copyleft obligations (e.g. commercial products, SaaS, large org internal use), purchase a commercial license to waive AGPL terms.
See COMMERCIAL-LICENSE.md for details.
