enginex-mr_series-sherpa-onnx

EngineX-Iluvatar/enginex-mr_series-sherpa-onnx

Archived

This repository has been archived on 2025-08-26. You can view files and clone it, but cannot push or open issues or pull requests.

Go to file

Fangjun Kuang f1b311ee4f Handle audio files less than 10s long for speaker diarization. (#1412 )

If the input audio file is less than 10 seconds long, there is only 
one chunk, and there is no need to compute embeddings or 
do clustering.

We can use the segmentation result from the speaker segmentation 
model directly.

2024-10-11 10:27:16 +08:00

.github

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

android

Fix releasing npm package and fix building Android VAD+ASR example (#1288 )

2024-08-26 10:18:48 +08:00

c-api-examples

C API for speaker diarization (#1402 )

2024-10-09 17:10:03 +08:00

cmake

C++ API for speaker diarization (#1396 )

2024-10-09 12:01:20 +08:00

dart-api-examples

Release v1.10.27 (#1359 )

2024-09-19 10:49:29 +08:00

dotnet-examples

C# API for speaker diarization (#1407 )

2024-10-10 14:29:05 +08:00

ffmpeg-examples

Fix ffmpeg c api example (#1185 )

2024-07-29 14:27:55 +08:00

flutter

Fix CI errors introduced by supporting loading keywords from buffers (#1366 )

2024-09-20 19:04:21 +08:00

flutter-examples

Update online_model.dart (#1375 )

2024-09-25 06:40:22 +08:00

go-api-examples

Go API for speaker diarization (#1403 )

2024-10-09 20:10:44 +08:00

ios-swift

Revert to onnxruntime 1.17.1 (#1131 )

2024-07-15 14:24:08 +08:00

ios-swiftui

Add MeloTTS example for ios (#1223 )

2024-08-06 14:48:54 +08:00

java-api-examples

Add APIs about max speech duration in VAD for various programming languages (#1349 )

2024-09-14 12:30:13 +08:00

kotlin-api-examples

Add Java and Kotlin API for sense voice (#1164 )

2024-07-22 14:08:40 +08:00

lazarus-examples

Add APIs about max speech duration in VAD for various programming languages (#1349 )

2024-09-14 12:30:13 +08:00

mfc-examples

fix building MFC examples (#1178 )

2024-07-28 14:07:25 +08:00

nodejs-addon-examples

JavaScript API (node-addon) for speaker diarization (#1408 )

2024-10-10 15:51:31 +08:00

nodejs-examples

Add APIs about max speech duration in VAD for various programming languages (#1349 )

2024-09-14 12:30:13 +08:00

pascal-api-examples

Text to speech API for Object Pascal. (#1273 )

2024-08-20 20:52:16 +08:00

python-api-examples

Python API for speaker diarization. (#1400 )

2024-10-09 14:13:26 +08:00

rust-api-examples

Update README to include Rust. (#1212 )

2024-08-04 12:20:05 +08:00

scripts

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

sherpa-onnx

Handle audio files less than 10s long for speaker diarization. (#1412 )

2024-10-11 10:27:16 +08:00

swift-api-examples

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

toolchains

Support RISC-V (#609 )

2024-02-26 06:57:18 +08:00

wasm

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

.clang-format

add java wrapper suppport (#117 )

2023-04-15 22:17:28 +08:00

.clang-tidy

Support clang-tidy (#1034 )

2024-06-19 20:51:57 +08:00

.flake8

add offline websocket server/client (#98 )

2023-03-29 21:48:45 +08:00

.gitignore

C API for speaker diarization (#1402 )

2024-10-09 17:10:03 +08:00

build-aarch64-linux-gnu.sh

Build websocket related binaries for embedded systems. (#1327 )

2024-09-08 17:16:58 +08:00

build-android-arm64-v8a.sh

Add Python API for clustering (#1385 )

2024-09-30 11:33:15 +08:00

build-android-armv7-eabi.sh

Add Python API for clustering (#1385 )

2024-09-30 11:33:15 +08:00

build-android-x86-64.sh

Add Python API for clustering (#1385 )

2024-09-30 11:33:15 +08:00

build-android-x86.sh

Add Python API for clustering (#1385 )

2024-09-30 11:33:15 +08:00

build-arm-linux-gnueabihf.sh

Build websocket related binaries for embedded systems. (#1327 )

2024-09-08 17:16:58 +08:00

build-ios-no-tts.sh

Add blank penalty for various language bindings. (#1234 )

2024-08-08 10:43:31 +08:00

build-ios-shared.sh

Revert to onnxruntime 1.17.1 (#1131 )

2024-07-15 14:24:08 +08:00

build-ios.sh

Add MeloTTS example for ios (#1223 )

2024-08-06 14:48:54 +08:00

build-riscv64-linux-gnu.sh

Build websocket related binaries for embedded systems. (#1327 )

2024-09-08 17:16:58 +08:00

build-swift-macos.sh

Fix CI errors. (#993 )

2024-06-12 11:42:19 +08:00

build-wasm-simd-asr.sh

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

build-wasm-simd-kws.sh

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

build-wasm-simd-nodejs.sh

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

build-wasm-simd-speaker-diarization.sh

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

build-wasm-simd-tts.sh

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

build-wasm-simd-vad-asr.sh

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

build-wasm-simd-vad.sh

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

CHANGELOG.md

Release v1.10.27 (#1359 )

2024-09-19 10:49:29 +08:00

CMakeLists.txt

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

CPPLINT.cfg

Use static libraries for MFC examples (#210 )

2023-07-13 14:52:43 +08:00

LICENSE

Use standard apache 2.0 license (#53 )

2023-02-22 11:30:46 +08:00

MANIFEST.in

Fix building wheels from source. (#632 )

2024-03-04 16:39:51 +08:00

new-release.sh

Release v1.10.27 (#1359 )

2024-09-19 10:49:29 +08:00

README.md

WebAssembly exmaple for speaker diarization (#1411 )

2024-10-10 22:14:45 +08:00

release.sh

Publish pre-compiled libs for Android. (#217 )

2023-07-15 12:25:18 +08:00

setup.py

Provide pre-built wheels with CUDA support. (#1143 )

2024-07-17 22:59:13 +08:00

README.md

Supported functions

Speech recognition	Speech synthesis
✔️	✔️

Speaker identification	Speaker diarization	Speaker identification
✔️	✔️	✔️

Spoken Language identification	Audio tagging	Voice activity detection
✔️	✔️	✔️

Keyword spotting	Add punctuation
✔️	✔️

Supported platforms

Architecture	Android	iOS	Windows	macOS	linux
x64	✔️		✔️	✔️	✔️
x86	✔️		✔️
arm64	✔️	✔️	✔️	✔️	✔️
arm32	✔️				✔️
riscv64					✔️

Supported programming languages

1. C++	2. C	3. Python	4. JavaScript
✔️	✔️	✔️	✔️

5. Java	6. C#	7. Kotlin	8. Swift
✔️	✔️	✔️	✔️

9. Go	10. Dart	11. Rust	12. Pascal
✔️	✔️	✔️	✔️

For Rust support, please see sherpa-rs

It also supports WebAssembly.

Introduction

This repository supports running the following functions locally

Speech-to-text (i.e., ASR); both streaming and non-streaming are supported
Text-to-speech (i.e., TTS)
Speaker diarization
Speaker identification
Speaker verification
Spoken language identification
Audio tagging
VAD (e.g., silero-vad)
Keyword spotting

on the following platforms and operating systems:

x86, x86_64, 32-bit ARM, 64-bit ARM (arm64, aarch64), RISC-V (riscv64)
Linux, macOS, Windows, openKylin
Android, WearOS
iOS
NodeJS
WebAssembly
Raspberry Pi
RV1126
LicheePi4A
VisionFive 2
旭日X3派
爱芯派
etc

with the following APIs

C++, C, Python, Go, C#
Java, Kotlin, JavaScript
Swift, Rust
Dart, Object Pascal

Links for Huggingface Spaces

You can visit the following Huggingface spaces to try sherpa-onnx without installing anything. All you need is a browser.

Description	URL
Speech recognition	Click me
Speech recognition with Whisper	Click me
Speech synthesis	Click me
Generate subtitles	Click me
Audio tagging	Click me
Spoken language identification with Whisper	Click me

We also have spaces built using WebAssembly. They are listed below:

Description	Huggingface space	ModelScope space
Voice activity detection with silero-vad	Click me	地址
Real-time speech recognition (Chinese + English) with Zipformer	Click me	地址
Real-time speech recognition (Chinese + English) with Paraformer	Click me	地址
Real-time speech recognition (Chinese + English + Cantonese) with Paraformer-large	Click me	地址
Real-time speech recognition (English)	Click me	地址
VAD + speech recognition (Chinese + English + Korean + Japanese + Cantonese) with SenseVoice	Click me	地址
VAD + speech recognition (English) with Whisper tiny.en	Click me	地址
VAD + speech recognition (English) with Zipformer trained with GigaSpeech	Click me	地址
VAD + speech recognition (Chinese) with Zipformer trained with WenetSpeech	Click me	地址
VAD + speech recognition (Japanese) with Zipformer trained with ReazonSpeech	Click me	地址
VAD + speech recognition (Thai) with Zipformer trained with GigaSpeech2	Click me	地址
VAD + speech recognition (Chinese 多种方言) with a TeleSpeech-ASR CTC model	Click me	地址
VAD + speech recognition (English + Chinese, 及多种中文方言) with Paraformer-large	Click me	地址
VAD + speech recognition (English + Chinese, 及多种中文方言) with Paraformer-small	Click me	地址
Speech synthesis (English)	Click me	地址
Speech synthesis (German)	Click me	地址
Speaker diarization	Click me	地址

Links for pre-built Android APKs

Description	URL	中国用户
Streaming speech recognition	Address	点此
Text-to-speech	Address	点此
Voice activity detection (VAD)	Address	点此
VAD + non-streaming speech recognition	Address	点此
Two-pass speech recognition	Address	点此
Audio tagging	Address	点此
Audio tagging (WearOS)	Address	点此
Speaker identification	Address	点此
Spoken language identification	Address	点此
Keyword spotting	Address	点此

Links for pre-built Flutter APPs

Real-time speech recognition

Description	URL	中国用户
Streaming speech recognition	Address	点此

Text-to-speech

Description	URL	中国用户
Android (arm64-v8a, armeabi-v7a, x86_64)	Address	点此
Linux (x64)	Address	点此
macOS (x64)	Address	点此
macOS (arm64)	Address	点此
Windows (x64)	Address	点此

Note: You need to build from source for iOS.

Links for pre-built Lazarus APPs

Generating subtitles

Description	URL	中国用户
Generate subtitles (生成字幕)	Address	点此

Links for pre-trained models

Description	URL
Speech recognition (speech to text, ASR)	Address
Text-to-speech (TTS)	Address
VAD	Address
Keyword spotting	Address
Audio tagging	Address
Speaker identification (Speaker ID)	Address
Spoken language identification (Language ID)	See multi-lingual Whisper ASR models from Speech recognition
Punctuation	Address
Speaker segmentation	Address

Useful links

Documentation: https://k2-fsa.github.io/sherpa/onnx/
Bilibili 演示视频: https://search.bilibili.com/all?keyword=%E6%96%B0%E4%B8%80%E4%BB%A3Kaldi

How to reach us

Please see https://k2-fsa.github.io/sherpa/social-groups.html for 新一代 Kaldi 微信交流群 and QQ 交流群.

Projects using sherpa-onnx

voiceapi

Streaming ASR and TTS based on FastAPI

It shows how to use the ASR and TTS Python APIs with FastAPI.

腾讯会议摸鱼工具 TMSpeech

Uses streaming ASR in C# with graphical user interface.

Video demo in Chinese: 【开源】Windows实时字幕软件（网课/开会必备）

lol互动助手

It uses the JavaScript API of sherpa-onnx along with Electron

Video demo in Chinese: 爆了！炫神教你开打字挂！真正影响胜率的英雄联盟工具！英雄联盟的最后一块拼图！和游戏中的每个人无障碍沟通！

Languages

C++ 38.3%

Python 16.3%

Shell 7.6%

Kotlin 5.1%

JavaScript 5.1%

Other 27.4%

README.md Unescape Escape

Supported functions

Supported platforms

Supported programming languages

Introduction

Links for Huggingface Spaces

Links for pre-built Android APKs

Links for pre-built Flutter APPs

Real-time speech recognition

Text-to-speech

Links for pre-built Lazarus APPs

Generating subtitles

Links for pre-trained models

Useful links

How to reach us

Projects using sherpa-onnx

voiceapi

腾讯会议摸鱼工具 TMSpeech

lol互动助手

README.md