/home/tony/Work/glockenspiel/s3prl/s3prl/upstream/byol_s/byol_a/common.py:20: UserWarning: torchaudio._backend.set_audio_backend has been deprecated. With dispatcher enabled, this function is no-op. You can remove the function call. torchaudio.set_audio_backend("sox_io") /home/tony/Work/glockenspiel/s3prl/s3prl/run_downstream.py:157: UserWarning: torchaudio._backend.set_audio_backend has been deprecated. With dispatcher enabled, this function is no-op. You can remove the function call. torchaudio.set_audio_backend('sox_io') 2023-10-14 21:20:54 | INFO | s3prl.util.download | Requesting URL: https://huggingface.co/s3prl/converted_ckpts/resolve/main/wavlm_base_plus.pt 2023-10-14 21:20:54 | INFO | s3prl.util.download | Using URL's local file: /home/tony/.cache/s3prl/download/72cb34edf8a3724c720467cf40b77ad20b1b714b5f694e9db57f521467f9006b.wavlm_base_plus.pt 2023-10-14 21:20:54 | INFO | s3prl.upstream.wavlm.WavLM | WavLM Config: {'extractor_mode': 'default', 'encoder_layers': 12, 'encoder_embed_dim': 768, 'encoder_ffn_embed_dim': 3072, 'encoder_attention_heads': 12, 'activation_fn': 'gelu', 'layer_norm_first': False, 'conv_feature_layers': '[(512,10,5)] + [(512,3,2)] * 4 + [(512,2,2)] * 2', 'conv_bias': False, 'feature_grad_mult': 0.1, 'normalize': False, 'dropout': 0.1, 'attention_dropout': 0.1, 'activation_dropout': 0.0, 'encoder_layerdrop': 0.05, 'dropout_input': 0.1, 'dropout_features': 0.1, 'mask_length': 10, 'mask_prob': 0.8, 'mask_selection': 'static', 'mask_other': 0.0, 'no_mask_overlap': False, 'mask_min_space': 1, 'mask_channel_length': 10, 'mask_channel_prob': 0.0, 'mask_channel_selection': 'static', 'mask_channel_other': 0.0, 'no_mask_channel_overlap': False, 'mask_channel_min_space': 1, 'conv_pos': 128, 'conv_pos_groups': 16, 'relative_position_embedding': True, 'num_buckets': 320, 'max_distance': 800, 'gru_rel_pos': True} /home/tony/anaconda3/envs/suno_env/lib/python3.10/site-packages/torch/nn/utils/weight_norm.py:30: UserWarning: torch.nn.utils.weight_norm is deprecated in favor of torch.nn.utils.parametrizations.weight_norm. warnings.warn("torch.nn.utils.weight_norm is deprecated in favor of torch.nn.utils.parametrizations.weight_norm.") /home/tony/anaconda3/envs/suno_env/lib/python3.10/site-packages/torch/nn/functional.py:5076: UserWarning: Support for mismatched key_padding_mask and attn_mask is deprecated. Use same type for both instead. warnings.warn( [Featurizer] - Take a list of 13 features and weighted sum them. [Featurizer] - The selected feature hidden_states's downsample rate is 320 [Runner] - Start a new experiment overall: 0%| | 0/4000 [00:00 phonemes: 0%| | 0/3515 [00:00 phonemes: 13%|█▎ | 458/3515 [00:00<00:00, 4578.92it/s] word -> phonemes: 27%|██▋ | 932/3515 [00:00<00:00, 4672.90it/s] word -> phonemes: 40%|████ | 1408/3515 [00:00<00:00, 4709.72it/s] word -> phonemes: 54%|█████▎ | 1884/3515 [00:00<00:00, 4727.18it/s] word -> phonemes: 67%|██████▋ | 2361/3515 [00:00<00:00, 4741.92it/s] word -> phonemes: 81%|████████ | 2837/3515 [00:00<00:00, 4746.52it/s] word -> phonemes: 94%|█████████▍| 3312/3515 [00:00<00:00, 4621.57it/s] word -> phonemes: 100%|██████████| 3515/3515 [00:00<00:00, 4659.25it/s] train: 0%| | 0/3515 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 22%|██▏ | 585/2703 [00:00<00:00, 5840.88it/s] word -> phonemes: 44%|████▎ | 1178/2703 [00:00<00:00, 5889.58it/s] word -> phonemes: 65%|██████▌ | 1767/2703 [00:00<00:00, 5875.15it/s] word -> phonemes: 87%|████████▋ | 2355/2703 [00:00<00:00, 5810.70it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5784.30it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 14%|█▍ | 389/2703 [00:00<00:02, 1032.95it/s] word -> phonemes: 32%|███▏ | 871/2703 [00:00<00:00, 2073.54it/s] word -> phonemes: 50%|█████ | 1364/2703 [00:00<00:00, 2878.50it/s] word -> phonemes: 68%|██████▊ | 1829/2703 [00:00<00:00, 3386.80it/s] word -> phonemes: 84%|████████▎ | 2259/2703 [00:00<00:00, 3651.01it/s] word -> phonemes: 100%|█████████▉| 2695/2703 [00:00<00:00, 3857.38it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 3058.87it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 20%|█▉ | 534/2703 [00:00<00:00, 5337.80it/s] word -> phonemes: 41%|████ | 1097/2703 [00:00<00:00, 5507.97it/s] word -> phonemes: 61%|██████ | 1652/2703 [00:00<00:00, 5523.58it/s] word -> phonemes: 82%|████████▏ | 2221/2703 [00:00<00:00, 5586.13it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5511.78it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/3515 [00:00 phonemes: 12%|█▏ | 408/3515 [00:00<00:00, 4075.63it/s] word -> phonemes: 25%|██▍ | 865/3515 [00:00<00:00, 4365.69it/s] word -> phonemes: 38%|███▊ | 1341/3515 [00:00<00:00, 4542.16it/s] word -> phonemes: 51%|█████▏ | 1803/3515 [00:00<00:00, 4570.46it/s] word -> phonemes: 65%|██████▍ | 2277/3515 [00:00<00:00, 4629.27it/s] word -> phonemes: 78%|███████▊ | 2740/3515 [00:00<00:00, 4514.62it/s] word -> phonemes: 91%|█████████ | 3192/3515 [00:00<00:00, 4469.01it/s] word -> phonemes: 100%|██████████| 3515/3515 [00:00<00:00, 4465.55it/s] train: 0%| | 0/3515 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 22%|██▏ | 592/2703 [00:00<00:00, 5912.37it/s] word -> phonemes: 44%|████▍ | 1184/2703 [00:00<00:00, 5828.27it/s] word -> phonemes: 66%|██████▌ | 1789/2703 [00:00<00:00, 5925.25it/s] word -> phonemes: 88%|████████▊ | 2382/2703 [00:00<00:00, 5887.03it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5821.92it/s] dev: 0%| | 0/2703 [00:00