Official implementation of A cappella: Audio-visual Singing VoiceSeparation, from BMVC21

Juan F. Montesinos

Last update: Oct 22, 2022

Related tags

Audio audio audiovisual bss bmvc y-net

Overview

Y-Net

Official implementation of A cappella: Audio-visual Singing VoiceSeparation, British Machine Vision Conference 2021

Project page: ipcv.github.io/Acappella/
Paper: Arxiv, BMVC (not available yet)

Running a demo / Y-Net Inference

We provide simple functions to load models with pre-trained weights. Steps:

Clone the repo or download y-net>VnBSS>models (models can run as a standalone package)
Load a model:

from VnBSS import y_net_gr # or from models import y_net_gr 
model = y_net_gr(n=1)

Check a demo fully working:

Citation

@inproceedings{acappella,
    author    = {Juan F. Montesinos and
                 Venkatesh S. Kadandale and
                 Gloria Haro},
    title     = {A cappella: Audio-visual Singing VoiceSeparation},
    booktitle = {British Machine Vision Conference (BMVC)},
    year      = {2021},

}

.
.
.
.
.
.
.
.

Training / Using DEV code

###Training The most difficult part is to prepare the dataset as everything is builded upon a very specific format.
To run training:
python run.py -m model_name --workname experiment_name --arxiv_path directory_of_experiments --pretrained_from path_pret_weights
You can inspect the argparse at default.py>argparse_default.
Possible model names are: y_net_g, y_net_gr, y_net_m,y_net_r,u_net,llcp

Testing

Go to manuscript_scripts and replace checkpoint paths by yours in the testing scripts.
Run: bash manuscript_scripts/test_gr_r.sh
Replace the paths of manuscript_scripts/auto_metrics.py by your experiment_directory path.
Run: python manuscript_scripts/auto_metrics.py to visualise results.

It's a complicated framework. HELP!

The best option to run the framework is to debug! Having a runable code helps to see input shapes, dataflow and to run line by line. Download The circle of life demo with the files already processed. It will act like a dataset of 6 samples. You can download it from Google Drive 1.1 Gb.

Unzip the file
run python run.py -m y_net_gr (for example)

Everything has been configured to run by default this way.

The model

Each effective model is wrapped by a nn.Module which takes care of computing the STFT, the mask, returning the waveform etcetera... This wrapper can be found at VnBSS>models>y_net.py>YNet. To get rid of this you can simply inherit the class, take minimum layers and keep the core_forward method, which is the inference step without the miscelanea.

FAQs

How to change the optimizer's hyperparameters?
Go to config>optimizer.json
How to change clip duration, video framerate, STFT parameters or audio samplerate?
Go to config>__init__.py
How to change the batch size or the amount of epochs?
Go to config>hyptrs.json
How to dump predictions from the training and test set
Go to default.py. Modify DUMP_FILES (can be controlled at a subset level). force argument skips the iteration-wise conditions and dumps for every single network prediction.
Is tensorboard enabled?
Yes, you will find tensorboard records at your_experiment_directory/used_workname/tensorboard
Can I resume an experiment?
Yes, if you set exactly the same experiment folder and workname, the system will detect it and will resume from there.
I'm trying to resume but found AssertionError If there is an exception before running the model
How to change the amount of layers of U-Net
U-net is build dynamically given a list of layers per block as shown in models>__init__.py from outer to inner blocks.
How to modify the default network values?
The json file config>net_cfg.json overwrites any default configuration from the model.

Python library for audio and music analysis

librosa A python package for music and audio analysis. Documentation See https://librosa.org/doc/ for a complete reference manual and introductory tut

5.6k Jan 6, 2023

?️ Open Source Audio Matching and Mastering

Matching + Mastering = ❤️ Matchering 2.0 is a novel Containerized Web Application and Python Library for audio matching and mastering. It follows a si

781 Jan 5, 2023

Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications

A Python library for audio feature extraction, classification, segmentation and applications This doc contains general info. Click here for the comple

5.1k Jan 2, 2023

Manipulate audio with a simple and easy high level interface

Pydub Pydub lets you do stuff to audio in a way that isn't stupid. Stuff you might be looking for: Installing Pydub API Documentation Dependencies Pla

6.6k Jan 1, 2023

Scalable audio processing framework written in Python with a RESTful API

TimeSide : scalable audio processing framework and server written in Python TimeSide is a python framework enabling low and high level audio analysis,

340 Jan 4, 2023

Python module for handling audio metadata

Mutagen is a Python module to handle audio metadata. It supports ASF, FLAC, MP4, Monkey's Audio, MP3, Musepack, Ogg Opus, Ogg FLAC, Ogg Speex, Ogg The

1.1k Dec 31, 2022

Python I/O for STEM audio files

stempeg = stems + ffmpeg Python package to read and write STEM audio files. Technically, stems are audio containers that combine multiple audio stream

72 Dec 23, 2022

Python library for handling audio datasets.

AUDIOMATE Audiomate is a library for easy access to audio datasets. It provides the datastructures for accessing/loading different datasets in a gener

121 Nov 27, 2022

An audio digital processing toolbox based on a workflow/pipeline principle

AudioTK Audio ToolKit is a set of audio filters. It helps assembling workflows for specific audio processing workloads. The audio workflow is split in

238 Oct 18, 2022

Comments

Missing required arguments ?

Hello, thanks for this great work and dataset, When I tried to run ,I got error below, Traceback (most recent call last): File "Desktop/Acappella-YNet/run.py", line 27, in iter_param, model, model_kwargs = VnBSS.ModelConstructor( File "/Desktop/Acappella-YNet/VnBSS/models/init.py", line 134, in build return self._build_dev() File "/Desktop/Acappella-YNet/VnBSS/models/init.py", line 140, in _build_dev model = constructor(**self.common_kwargs) TypeError: init() missing 6 required keyword-only arguments: 'remix_input', 'remix_coef', 'video_enabled', 'llcp_enabled', 'skeleton_enabled', and 'activation'

opened by Enescigdem 9
Metrics

Hi,

I run run.py and it gives me some metrics like sdr/ds, etc.. Are these metrics equal to ones without ds? I search ds term in code, but couldn't find what it is exactly. For example I got sdr/ds=13.4 for my training, can I say sdr=13.4 ?

opened by EmreOzkose 2
There is no dataset config ?

Hi, In README. md , Download the code and set your dataset paths at config>dataset_paths.json is said but there is no such file in code directory or it is not mentioned about the format of that json file. Thanks in advance

opened by Enescigdem 2

UnboundLocalError: local variable 'warped_img' referenced before assignment

Hi, thanks for releasing your code and dataset. I encounter this error on executing preprocess.py


Exception Handled:  'NoneType' object is not iterable
  0%|                                                                                                                                           |
 samples processed:   0%|
Traceback (most recent call last):
  File "preprocess.py", line 314, in <module>
    mean_face)
  File "preprocess.py", line 254, in process_samples
    del warped_img, landmarks, bbox, vid, stacked_frames, stacked_landmarks
UnboundLocalError: local variable 'warped_img' referenced before assignment

Error log is saying that 'warped_img' isn't assigned. So I checked 'warped_img' and that Errorhandling code.

In 231 line, in preprocess_sample function,

with tqdm(total=num_frames) as pbar:
                for num_frame in range(num_frames):
                    try:
                        img_raw = vid.get_data(num_frame)
                    except IndexError as e:
                        print("Processing FAILED for sample: " + sample_id)
                        break
                    if img_raw.shape[1] >= MAX_IMAGE_WIDTH:
                        asp_ratio = img_raw.shape[0] / img_raw.shape[1]
                        dim = (MAX_IMAGE_WIDTH, int(MAX_IMAGE_WIDTH * asp_ratio))
                        new_img = cv2.resize(img_raw, dim, interpolation=cv2.INTER_AREA)
                        img = np.asarray(new_img)
                    else:
                        img = img_raw
                    try:
                        **_warped_img, landmarks, bbox = fp.process_image(img)_**
                        _, aligned_landmarks, _ = fp.process_image(warped_img)
                        good_frame_ids.append(num_frame)
                    except Exception as e:
                        print("Exception Handled: ", e)
                        continue

I think face detector can't find face , but I don't know why If you give advice to me, I'm very appreciate.

opened by yddr 2

Owner

Juan F. Montesinos

PhD student at Pompeu Fabra university Barcelona

GitHub https://ipcv.github.io/Acappella/

Official implementation of A cappella: Audio-visual Singing VoiceSeparation, from BMVC21

Related tags

Overview

Y-Net

Running a demo / Y-Net Inference

Citation

Training / Using DEV code

Testing

It's a complicated framework. HELP!

The model

FAQs

You might also like...

Python library for audio and music analysis

?️ Open Source Audio Matching and Mastering

Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications

Manipulate audio with a simple and easy high level interface

Scalable audio processing framework written in Python with a RESTful API

Python module for handling audio metadata

Python I/O for STEM audio files

Python library for handling audio datasets.

An audio digital processing toolbox based on a workflow/pipeline principle

Comments

Missing required arguments ?

Metrics

There is no dataset config ?

UnboundLocalError: local variable 'warped_img' referenced before assignment

Owner

Juan F. Montesinos

[Singing Log] Let your program learn to sing!

cross-library (GStreamer + Core Audio + MAD + FFmpeg) audio decoding for Python

cross-library (GStreamer + Core Audio + MAD + FFmpeg) audio decoding for Python

Audio spatialization over WebRTC and JACK Audio Connection Kit

Audio augmentations library for PyTorch for audio in the time-domain

praudio provides audio preprocessing framework for Deep Learning audio applications

convert-to-opus-cli is a Python CLI program for converting audio files to opus audio format.

Implementation of "Slow-Fast Auditory Streams for Audio Recognition, ICASSP, 2021" in PyTorch

Audio fingerprinting and recognition in Python

kapre: Keras Audio Preprocessors