BigVGAN v2 Open-Source Audio Synthesis Model - Multi-Sampling Rate and Frequency Band Configuration, Achieve High-Quality Sound Effects for Free!

Bigvgan V2 24khz 100band 256x

Developed by nvidia

BigVGAN is a high-performance neural vocoder that achieves high-quality audio synthesis through large-scale training, supporting multiple sampling rates and frequency band configurations.

Audio Generation Open Source License:MIT #High-fidelity audio synthesis #Multi-scale mel spectrogram #CUDA-accelerated inference

Downloads 34.03k

Release Time : 7/15/2024

Model Overview

BigVGAN is a universal neural vocoder capable of converting mel spectrograms into high-quality waveform audio. Through large-scale training and advanced architectural design, it achieves excellent audio generation results.

Model Features

Large-scale training

Trained on a diverse audio dataset including multilingual speech, environmental sounds, and musical instruments to enhance the model's generalization capability.

High-performance inference

Provides custom CUDA kernels supporting fused upsampling + activation operations, improving inference speed by 1.5-3x.

Multi-configuration support

Offers pre-trained models with various sampling rates (22kHz-44kHz) and frequency band configurations to adapt to different application scenarios.

Improved discriminator

Utilizes multi-scale sub-band CQT discriminators and multi-scale mel spectrogram loss training to enhance generation quality.

Model Capabilities

Mel spectrogram to waveform conversion

High-quality audio synthesis

Multi-sampling rate support

Fast inference

Use Cases

Speech synthesis

Text-to-speech systems

Serves as the backend vocoder for TTS systems, converting mel spectrograms into natural speech waveforms.

Generates high-quality, natural speech output

Audio enhancement

Audio super-resolution

Converts low-quality audio into high-quality waveforms.

Improves audio quality and clarity

Music generation

Musical instrument sound synthesis

Generates waveform sounds for various musical instruments.

High-quality musical instrument timbre synthesis

🚀 BigVGAN: A Universal Neural Vocoder with Large-Scale Training

BigVGAN is a universal neural vocoder trained on a large scale, aiming to solve the problem of high - quality audio synthesis and providing an effective solution for speech synthesis and audio generation tasks.

Authors

Sang - gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, Sungroh Yoon

[Paper] - [Code] - [[Showcase]](https://bigvgan - demo.github.io/) - [Project Page] - [[Weights]](https://huggingface.co/collections/nvidia/bigvgan - 66959df3d97fd7d98d97dc9a) - [Demo]

✨ Features

News

Jul 2024 (v2.3):
- General refactor and code improvements for improved readability.
- Fully fused CUDA kernel of anti - alised activation (upsampling + activation + downsampling) with inference speed benchmark.
Jul 2024 (v2.2): The repository now includes an interactive local demo using gradio.
Jul 2024 (v2.1): BigVGAN is now integrated with 🤗 Hugging Face Hub with easy access to inference using pretrained checkpoints. We also provide an interactive demo on Hugging Face Spaces.
Jul 2024 (v2): We release BigVGAN - v2 along with pretrained checkpoints. Below are the highlights:
- Custom CUDA kernel for inference: we provide a fused upsampling + activation kernel written in CUDA for accelerated inference speed. Our test shows 1.5 - 3x faster speed on a single A100 GPU.
- Improved discriminator and loss: BigVGAN - v2 is trained using a multi - scale sub - band CQT discriminator and a multi - scale mel spectrogram loss.
- Larger training data: BigVGAN - v2 is trained using datasets containing diverse audio types, including speech in multiple languages, environmental sounds, and instruments.
- We provide pretrained checkpoints of BigVGAN - v2 using diverse audio configurations, supporting up to 44 kHz sampling rate and 512x upsampling ratio.

📦 Installation

This repository contains pretrained BigVGAN checkpoints with easy access to inference and additional huggingface_hub support.

If you are interested in training the model and additional functionalities, please visit the official GitHub repository for more information: https://github.com/NVIDIA/BigVGAN

git lfs install
git clone https://huggingface.co/nvidia/bigvgan_v2_24khz_100band_256x

💻 Usage Examples

Basic Usage

Below example describes how you can use BigVGAN: load the pretrained BigVGAN generator from Hugging Face Hub, compute mel spectrogram from input waveform, and generate synthesized waveform using the mel spectrogram as the model's input.

device = 'cuda'

import torch
import bigvgan
import librosa
from meldataset import get_mel_spectrogram

# instantiate the model. You can optionally set use_cuda_kernel=True for faster inference.
model = bigvgan.BigVGAN.from_pretrained('nvidia/bigvgan_v2_24khz_100band_256x', use_cuda_kernel=False)

# remove weight norm in the model and set to eval mode
model.remove_weight_norm()
model = model.eval().to(device)

# load wav file and compute mel spectrogram
wav_path = '/path/to/your/audio.wav'
wav, sr = librosa.load(wav_path, sr=model.h.sampling_rate, mono=True) # wav is np.ndarray with shape [T_time] and values in [-1, 1]
wav = torch.FloatTensor(wav).unsqueeze(0) # wav is FloatTensor with shape [B(1), T_time]

# compute mel spectrogram from the ground truth audio
mel = get_mel_spectrogram(wav, model.h).to(device) # mel is FloatTensor with shape [B(1), C_mel, T_frame]

# generate waveform from mel
with torch.inference_mode():
    wav_gen = model(mel) # wav_gen is FloatTensor with shape [B(1), 1, T_time] and values in [-1, 1]
wav_gen_float = wav_gen.squeeze(0).cpu() # wav_gen is FloatTensor with shape [1, T_time]

# you can convert the generated waveform to 16 bit linear PCM
wav_gen_int16 = (wav_gen_float * 32767.0).numpy().astype('int16') # wav_gen is now np.ndarray with shape [1, T_time] and int16 dtype

Advanced Usage

You can apply the fast CUDA inference kernel by using a parameter use_cuda_kernel when instantiating BigVGAN:

import bigvgan
model = bigvgan.BigVGAN.from_pretrained('nvidia/bigvgan_v2_24khz_100band_256x', use_cuda_kernel=True)

When applied for the first time, it builds the kernel using nvcc and ninja. If the build succeeds, the kernel is saved to alias_free_activation/cuda/build and the model automatically loads the kernel. The codebase has been tested using CUDA 12.1.

Please make sure that both are installed in your system and nvcc installed in your system matches the version your PyTorch build is using.

For detail, see the official GitHub repository: https://github.com/NVIDIA/BigVGAN?tab=readme-ov-file#using-custom-cuda-kernel-for-synthesis

📚 Documentation

Pretrained Models

We provide the pretrained models on Hugging Face Collections. One can download the checkpoints of the generator weight (named bigvgan_generator.pt) and its discriminator/optimizer states (named bigvgan_discriminator_optimizer.pt) within the listed model repositories.

Property	Details
Model Name	bigvgan_v2_44khz_128band_512x, bigvgan_v2_44khz_128band_256x, bigvgan_v2_24khz_100band_256x, bigvgan_v2_22khz_80band_256x, bigvgan_v2_22khz_80band_fmax8k_256x, bigvgan_24khz_100band, bigvgan_base_24khz_100band, bigvgan_22khz_80band, bigvgan_base_22khz_80band
Sampling Rate	44 kHz, 24 kHz, 22 kHz
Mel band	128, 100, 80
fmax	22050, 12000, 11025, 8000
Upsampling Ratio	512, 256
Params	122M, 112M, 14M
Dataset	Large - scale Compilation, LibriTTS, LibriTTS + VCTK + LJSpeech
Steps	5M
Fine - Tuned	No

📄 License

This project is under the MIT License.

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご