/home/tony/Work/glockenspiel/s3prl/s3prl/upstream/byol_s/byol_a/common.py:20: UserWarning: torchaudio._backend.set_audio_backend has been deprecated. With dispatcher enabled, this function is no-op. You can remove the function call. torchaudio.set_audio_backend("sox_io") /home/tony/Work/glockenspiel/s3prl/s3prl/run_downstream.py:157: UserWarning: torchaudio._backend.set_audio_backend has been deprecated. With dispatcher enabled, this function is no-op. You can remove the function call. torchaudio.set_audio_backend('sox_io') 2023-10-14 22:17:49 | INFO | s3prl.upstream.hubert.hubconf | Converting a fairseq checkpoint: /home/tony/Data/MERT/hubert_d2v2_200k.pt 2023-10-14 22:17:49 | INFO | s3prl.upstream.hubert.hubconf | To: /home/tony/Data/MERT/hubert_d2v2_200k.converted.pt 2023-10-14 22:17:56 | INFO | fairseq.models.hubert.hubert | HubertModel Config: HubertConfig(_name='hubert', label_rate=50.0, extractor_mode='default', encoder_layers=12, encoder_embed_dim=768, encoder_ffn_embed_dim=3072, encoder_attention_heads=12, activation_fn='gelu', layer_type='transformer', dropout=0.1, attention_dropout=0.1, activation_dropout=0.0, encoder_layerdrop=0.05, dropout_input=0.1, dropout_features=0.1, final_dim=256, untie_final_proj=True, layer_norm_first=False, conv_feature_layers='[(512,10,5)] + [(512,3,2)] * 4 + [(512,2,2)] * 3', conv_bias=False, logit_temp=0.1, target_glu=False, feature_grad_mult=0.1, mask_length=10, mask_prob=0.8, mask_selection='static', mask_other=0.0, no_mask_overlap=False, mask_min_space=1, mask_channel_length=10, mask_channel_prob=0.0, mask_channel_selection='static', mask_channel_other=0.0, no_mask_channel_overlap=False, mask_channel_min_space=1, conv_pos=128, conv_pos_groups=16, conv_pos_batch_norm=False, latent_temp=[2.0, 0.5, 0.999995], skip_masked=False, skip_nomask=False, checkpoint_activations=False, required_seq_len_multiple=2, depthwise_conv_kernel_size=31, attn_type='', pos_enc_type='abs', fp16=False) /home/tony/anaconda3/envs/suno_env/lib/python3.10/site-packages/torch/nn/utils/weight_norm.py:30: UserWarning: torch.nn.utils.weight_norm is deprecated in favor of torch.nn.utils.parametrizations.weight_norm. warnings.warn("torch.nn.utils.weight_norm is deprecated in favor of torch.nn.utils.parametrizations.weight_norm.") 2023-10-14 22:17:58 | INFO | fairseq.models.hubert.hubert | cannot find dictionary. assume will be used for fine-tuning [Featurizer] - Take a list of 13 features and weighted sum them. [Featurizer] - The selected feature hidden_states's downsample rate is 640 [Runner] - Start a new experiment [(512, 10, 5), (512, 3, 2), (512, 3, 2), (512, 3, 2), (512, 3, 2), (512, 2, 2), (512, 2, 2), (512, 2, 2)] 640 overall: 0%| | 0/4000 [00:00 phonemes: 0%| | 0/3515 [00:00 phonemes: 12%|█▏ | 437/3515 [00:00<00:00, 4363.82it/s] word -> phonemes: 25%|██▍ | 874/3515 [00:00<00:00, 4266.14it/s] word -> phonemes: 39%|███▊ | 1361/3515 [00:00<00:00, 4535.65it/s] word -> phonemes: 53%|█████▎ | 1848/3515 [00:00<00:00, 4662.94it/s] word -> phonemes: 66%|██████▌ | 2327/3515 [00:00<00:00, 4708.44it/s] word -> phonemes: 80%|███████▉ | 2799/3515 [00:00<00:00, 4658.51it/s] word -> phonemes: 93%|█████████▎| 3266/3515 [00:00<00:00, 4495.93it/s] word -> phonemes: 100%|██████████| 3515/3515 [00:00<00:00, 4526.60it/s] train: 0%| | 0/3515 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 20%|██ | 553/2703 [00:00<00:00, 5527.49it/s] word -> phonemes: 41%|████ | 1106/2703 [00:00<00:00, 5245.15it/s] word -> phonemes: 60%|██████ | 1632/2703 [00:00<00:00, 5201.97it/s] word -> phonemes: 81%|████████ | 2193/2703 [00:00<00:00, 5359.15it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5308.00it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 13%|█▎ | 349/2703 [00:00<00:01, 1313.18it/s] word -> phonemes: 32%|███▏ | 874/2703 [00:00<00:00, 2689.32it/s] word -> phonemes: 49%|████▉ | 1322/2703 [00:00<00:00, 3284.34it/s] word -> phonemes: 70%|██████▉ | 1882/2703 [00:00<00:00, 4029.75it/s] word -> phonemes: 90%|█████████ | 2441/2703 [00:00<00:00, 4520.60it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 3776.49it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 19%|█▊ | 505/2703 [00:00<00:00, 4915.26it/s] word -> phonemes: 40%|███▉ | 1075/2703 [00:00<00:00, 5368.75it/s] word -> phonemes: 60%|█████▉ | 1616/2703 [00:00<00:00, 5383.65it/s] word -> phonemes: 82%|████████▏ | 2206/2703 [00:00<00:00, 5583.77it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5469.47it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/3515 [00:00 phonemes: 11%|█ | 386/3515 [00:00<00:00, 3859.20it/s] word -> phonemes: 23%|██▎ | 812/3515 [00:00<00:00, 4093.51it/s] word -> phonemes: 37%|███▋ | 1296/3515 [00:00<00:00, 4433.12it/s] word -> phonemes: 50%|█████ | 1770/3515 [00:00<00:00, 4553.37it/s] word -> phonemes: 64%|██████▍ | 2262/3515 [00:00<00:00, 4683.91it/s] word -> phonemes: 78%|███████▊ | 2731/3515 [00:00<00:00, 4452.46it/s] word -> phonemes: 90%|█████████ | 3179/3515 [00:00<00:00, 4458.49it/s] word -> phonemes: 100%|██████████| 3515/3515 [00:00<00:00, 4381.90it/s] train: 0%| | 0/3515 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 22%|██▏ | 582/2703 [00:00<00:00, 5812.92it/s] word -> phonemes: 43%|████▎ | 1164/2703 [00:00<00:00, 5625.46it/s] word -> phonemes: 64%|██████▍ | 1742/2703 [00:00<00:00, 5693.46it/s] word -> phonemes: 86%|████████▌ | 2312/2703 [00:00<00:00, 5694.50it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5625.61it/s] dev: 0%| | 0/2703 [00:00