đ BigVGAN: A Universal Neural Vocoder with Large-Scale Training
BigVGAN is a universal neural vocoder trained on a large scale, aiming to solve the problem of high - quality audio synthesis and providing an effective solution for speech synthesis and audio generation tasks.
Authors
Sang - gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, Sungroh Yoon
[Paper] - [Code] - [[Showcase]](https://bigvgan - demo.github.io/) - [Project Page] - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan - 66959df3d97fd7d98d97dc9a) - [Demo]

⨠Features
News
- Jul 2024 (v2.3):
- General refactor and code improvements for improved readability.
- Fully fused CUDA kernel of anti - alised activation (upsampling + activation + downsampling) with inference speed benchmark.
- Jul 2024 (v2.2): The repository now includes an interactive local demo using gradio.
- Jul 2024 (v2.1): BigVGAN is now integrated with đ¤ Hugging Face Hub with easy access to inference using pretrained checkpoints. We also provide an interactive demo on Hugging Face Spaces.
- Jul 2024 (v2): We release BigVGAN - v2 along with pretrained checkpoints. Below are the highlights:
- Custom CUDA kernel for inference: we provide a fused upsampling + activation kernel written in CUDA for accelerated inference speed. Our test shows 1.5 - 3x faster speed on a single A100 GPU.
- Improved discriminator and loss: BigVGAN - v2 is trained using a multi - scale sub - band CQT discriminator and a multi - scale mel spectrogram loss.
- Larger training data: BigVGAN - v2 is trained using datasets containing diverse audio types, including speech in multiple languages, environmental sounds, and instruments.
- We provide pretrained checkpoints of BigVGAN - v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling ratio.
đĻ Installation
This repository contains pretrained BigVGAN checkpoints with easy access to inference and additional huggingface_hub
support.
If you are interested in training the model and additional functionalities, please visit the official GitHub repository for more information: https://github.com/NVIDIA/BigVGAN
git lfs install
git clone https://huggingface.co/nvidia/bigvgan_v2_24khz_100band_256x
đģ Usage Examples
Basic Usage
Below example describes how you can use BigVGAN: load the pretrained BigVGAN generator from Hugging Face Hub, compute mel spectrogram from input waveform, and generate synthesized waveform using the mel spectrogram as the model's input.
device = 'cuda'
import torch
import bigvgan
import librosa
from meldataset import get_mel_spectrogram
model = bigvgan.BigVGAN.from_pretrained('nvidia/bigvgan_v2_24khz_100band_256x', use_cuda_kernel=False)
model.remove_weight_norm()
model = model.eval().to(device)
wav_path = '/path/to/your/audio.wav'
wav, sr = librosa.load(wav_path, sr=model.h.sampling_rate, mono=True)
wav = torch.FloatTensor(wav).unsqueeze(0)
mel = get_mel_spectrogram(wav, model.h).to(device)
with torch.inference_mode():
wav_gen = model(mel)
wav_gen_float = wav_gen.squeeze(0).cpu()
wav_gen_int16 = (wav_gen_float * 32767.0).numpy().astype('int16')
Advanced Usage
You can apply the fast CUDA inference kernel by using a parameter use_cuda_kernel
when instantiating BigVGAN:
import bigvgan
model = bigvgan.BigVGAN.from_pretrained('nvidia/bigvgan_v2_24khz_100band_256x', use_cuda_kernel=True)
When applied for the first time, it builds the kernel using nvcc
and ninja
. If the build succeeds, the kernel is saved to alias_free_activation/cuda/build
and the model automatically loads the kernel. The codebase has been tested using CUDA 12.1
.
Please make sure that both are installed in your system and nvcc
installed in your system matches the version your PyTorch build is using.
For detail, see the official GitHub repository: https://github.com/NVIDIA/BigVGAN?tab=readme-ov-file#using-custom-cuda-kernel-for-synthesis
đ Documentation
Pretrained Models
We provide the pretrained models on Hugging Face Collections.
One can download the checkpoints of the generator weight (named bigvgan_generator.pt
) and its discriminator/optimizer states (named bigvgan_discriminator_optimizer.pt
) within the listed model repositories.
Property |
Details |
Model Name |
bigvgan_v2_44khz_128band_512x, bigvgan_v2_44khz_128band_256x, bigvgan_v2_24khz_100band_256x, bigvgan_v2_22khz_80band_256x, bigvgan_v2_22khz_80band_fmax8k_256x, bigvgan_24khz_100band, bigvgan_base_24khz_100band, bigvgan_22khz_80band, bigvgan_base_22khz_80band |
Sampling Rate |
44 kHz, 24 kHz, 22 kHz |
Mel band |
128, 100, 80 |
fmax |
22050, 12000, 11025, 8000 |
Upsampling Ratio |
512, 256 |
Params |
122M, 112M, 14M |
Dataset |
Large - scale Compilation, LibriTTS, LibriTTS + VCTK + LJSpeech |
Steps |
5M |
Fine - Tuned |
No |
đ License
This project is under the MIT License.