Repository: jianchang512/pyvideotrans
Stars: 16905
README.md
Sponsors: Recall.ai - Meeting Transcription API
> If youβre looking for a transcription API for meetings, consider checking out Recall.ai , an API that works with Zoom, Google Meet, Microsoft Teams, and more
pyVideoTrans
<div align="center">
A Powerful Open Source Video Translation / Audio Transcription / AI Dubbing / Subtitle Translation Tool
δΈζ | Documentation | Online Q&A
  ![Platform]()
</div>
pyVideoTrans is dedicated to seamlessly converting videos from one language to another, offering a complete workflow that includes speech recognition, subtitle translation, multi-role dubbing, and audio-video synchronization. It supports both local offline deployment and a wide variety of mainstream online APIs.
<img width="1658" height="935" alt="image" src="https://github.com/user-attachments/assets/c5959e59-6014-480c-9a7d-44c2b1729d36" />
---
β¨ Core Features
- π₯ Fully Automatic Video Translation: One-click workflow: Speech Recognition (ASR) -> Subtitle Translation -> Speech Synthesis (TTS) -> Video Synthesis.
- ποΈ Audio Transcription / Subtitle Generation: Batch convert audio/video to SRT subtitles, supporting Speaker Diarization to distinguish between different roles.
- π£οΈ Multi-Role AI Dubbing: Assign different AI dubbing voices to different speakers.
- 𧬠Voice Cloning: Integrates models like F5-TTS, CosyVoice, GPT-SoVITS for zero-shot voice cloning.
- π§ Powerful Model Support:
- ASR: Faster-Whisper (Local), OpenAI Whisper, Alibaba Qwen, ByteDance Volcano, Azure, Google, etc.
- LLM Translation: DeepSeek, ChatGPT, Claude, Gemini, MiniMax, Ollama (Local), Alibaba Bailian, etc.
- TTS: Edge-TTS (Free), OpenAI, Azure, Minimaxi, ChatTTS, ChatterBox, etc.
- π₯οΈ Interactive Editing: Supports pausing and manual proofreading at each stage (recognition, translation, dubbing) to ensure accuracy.
- π οΈ Utility Toolkit: Includes auxiliary tools such as vocal separation, video/subtitle merging, audio-video alignment, and transcript matching.
- π» Command Line Interface (CLI): Supports headless operation, convenient for server deployment or batch processing.
<img width="2752" height="1536" alt="unnamed" src="https://github.com/user-attachments/assets/960e9e34-84a4-425d-b582-f726623475a8" />
---
π Quick Start (Windows Users)
We provide a pre-packaged .exe version for Windows 10/11 users, requiring no Python environment configuration.
1. Download: Click to download the latest pre-packaged version
2. Unzip: Extract the compressed file to a path (e.g., D:\pyVideoTrans).
3. Run: Double-click sp.exe inside the folder to launch.
Note:
* Do not run directly from within the compressed archive.
* To use GPU acceleration, ensure CUDA 12.8 and cuDNN 9.11 are installed.
---
π οΈ Source Deployment (macOS / Linux / Windows Developers)
We recommend using uv for package management for faster speed and better environment isolation.
1. Prerequisites
* Python: Recommended version 3.10 --> 3.12
* FFmpeg: Must be installed and configured in the environment variables.
* macOS: brew install ffmpeg libsndfile git
* Linux (Ubuntu/Debian): sudo apt-get install ffmpeg libsndfile1-dev
* Windows: Download FFmpeg and configure Path, or place ffmpeg.exe and ffprobe.exe directly in the project directory.
2. Install uv (If not installed)
macOS/Linux
curl -LsSf https://astral.sh/uv/install.sh | shWindows (PowerShell)
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"3. Clone and Install
1. Clone the repository (Ensure path has no spaces/Chinese characters)
git clone https://github.com/jianchang512/pyvideotrans.git
cd pyvideotrans2. Install dependencies (uv automatically syncs environment)
uv syncIf you need local channels for qwen-tts and qwen-asr, please execute uv sync --extra qwen-tts --extra qwen-asr
4. Launch Software
Launch GUI:
uv run sp.pyUse CLI:
View documentation for detailed parameters
Video Translation Example
uv run cli.py --task vtv --name "./video.mp4" --source_language_code zh --target_language_code enAudio to Subtitle Example
uv run cli.py --task stt --name "./audio.wav" --model_name large-v35. (Optional) GPU Acceleration Configuration
If you have an NVIDIA graphics card, execute the following commands to install the CUDA-supported PyTorch version:
Uninstall CPU version
uv remove torch torchaudioInstall CUDA version (Example for CUDA 12.x)
uv add torch==2.7 torchaudio==2.7 --index-url https://download.pytorch.org/whl/cu128
uv add nvidia-cublas-cu12 nvidia-cudnn-cu12---
π§© Supported Channels & Models (Partial)
| Category | Channel/Model | Description |
| :--- | :--- | :--- |
| ASR (Speech Recognition) | Faster-Whisper (Local) | Recommended, fast speed, high accuracy |
| | WhisperX / Parakeet | Supports timestamp alignment & speaker diarization |
| | Alibaba Qwen3-ASR / ByteDance Volcano | Online API, excellent for Chinese |
| Translation (LLM/MT) | DeepSeek / ChatGPT | Supports context understanding, more natural translation |
| | MiniMax AI | MiniMax M2.7 LLM, latest flagship model, OpenAI-compatible |
| | Google / Microsoft | Traditional machine translation, fast speed |
| | Ollama / M2M100 | Fully local offline translation |
| TTS (Speech Synthesis) | Edge-TTS | Microsoft free interface, natural effect |
| | F5-TTS / CosyVoice | Supports Voice Cloning, requires local deployment |
| | GPT-SoVITS / ChatTTS | High-quality open-source TTS |
| | 302.AI / OpenAI / Azure | High-quality commercial API |
---
π Documentation & Support
* Official Documentation: https://pyvideotrans.com (Includes detailed tutorials, API configuration guides, FAQ)
* Online Q&A Community: https://bbs.pyvideotrans.com (Submit error logs for automated AI analysis and answers)
β οΈ Disclaimer
This software is an open-source, free, non-commercial project. Users are solely responsible for any legal consequences arising from the use of this software (including but not limited to calling third-party APIs or processing copyrighted video content). Please comply with local laws and regulations and the terms of use of relevant service providers.
π Acknowledgements
This project mainly relies on the following open-source projects (partial):
* FFmpeg
* PySide6
* faster-whisper
* openai-whisper
* edge-tts
* F5-TTS
* CosyVoice
---
Created by jianchang512