/home/tony/Work/glockenspiel/s3prl/s3prl/upstream/byol_s/byol_a/common.py:20: UserWarning: torchaudio._backend.set_audio_backend has been deprecated. With dispatcher enabled, this function is no-op. You can remove the function call. torchaudio.set_audio_backend("sox_io") 2023-10-11 21:18:18 | WARNING | root | Pytorch pre-release version 2.1.0.dev20230831+cu118 - assuming intent to test it /home/tony/Work/glockenspiel/s3prl/s3prl/run_downstream.py:157: UserWarning: torchaudio._backend.set_audio_backend has been deprecated. With dispatcher enabled, this function is no-op. You can remove the function call. torchaudio.set_audio_backend('sox_io') Some weights of the model checkpoint at m-a-p/MERT-v1-330M were not used when initializing MERTModel: ['encoder.pos_conv_embed.conv.weight_g', 'encoder.pos_conv_embed.conv.weight_v'] - This IS expected if you are initializing MERTModel from the checkpoint of a model trained on another task or with another architecture (e.g. initializing a BertForSequenceClassification model from a BertForPreTraining model). - This IS NOT expected if you are initializing MERTModel from the checkpoint of a model that you expect to be exactly identical (initializing a BertForSequenceClassification model from a BertForSequenceClassification model). Some weights of MERTModel were not initialized from the model checkpoint at m-a-p/MERT-v1-330M and are newly initialized: ['encoder.pos_conv_embed.conv.parametrizations.weight.original0', 'encoder.pos_conv_embed.conv.parametrizations.weight.original1'] You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference. [Featurizer] - The selected feature hidden_states's downsample rate is 320 [Runner] - Start a new experiment mert_330M_encode_norm2 {'model_config': None, 'refresh': False} using suno's k-means clusters: mert_330M_encode_norm2 overall: 0%| | 0/4000 [00:00 phonemes: 0%| | 0/3515 [00:00 phonemes: 14%|█▎ | 482/3515 [00:00<00:00, 4819.22it/s] word -> phonemes: 28%|██▊ | 993/3515 [00:00<00:00, 4984.23it/s] word -> phonemes: 43%|████▎ | 1519/3515 [00:00<00:00, 5107.08it/s] word -> phonemes: 58%|█████▊ | 2030/3515 [00:00<00:00, 4980.85it/s] word -> phonemes: 72%|███████▏ | 2529/3515 [00:00<00:00, 4714.83it/s] word -> phonemes: 86%|████████▌ | 3014/3515 [00:00<00:00, 4755.96it/s] word -> phonemes: 99%|█████████▉| 3492/3515 [00:00<00:00, 4745.85it/s] word -> phonemes: 100%|██████████| 3515/3515 [00:00<00:00, 4821.69it/s] train: 0%| | 0/3515 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 21%|██▏ | 579/2703 [00:00<00:00, 5785.15it/s] word -> phonemes: 44%|████▎ | 1177/2703 [00:00<00:00, 5897.74it/s] word -> phonemes: 66%|██████▌ | 1782/2703 [00:00<00:00, 5965.93it/s] word -> phonemes: 88%|████████▊ | 2379/2703 [00:00<00:00, 5913.09it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5625.28it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 20%|██ | 549/2703 [00:00<00:00, 5489.40it/s] word -> phonemes: 42%|████▏ | 1130/2703 [00:00<00:00, 5675.71it/s] word -> phonemes: 63%|██████▎ | 1711/2703 [00:00<00:00, 5735.32it/s] word -> phonemes: 85%|████████▍ | 2285/2703 [00:00<00:00, 5611.05it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5604.12it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 22%|██▏ | 599/2703 [00:00<00:00, 5976.14it/s] word -> phonemes: 45%|████▍ | 1203/2703 [00:00<00:00, 6013.34it/s] word -> phonemes: 68%|██████▊ | 1829/2703 [00:00<00:00, 6122.55it/s] word -> phonemes: 90%|█████████ | 2443/2703 [00:00<00:00, 6126.38it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 6065.88it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/3515 [00:00 phonemes: 12%|█▏ | 414/3515 [00:00<00:00, 4135.54it/s] word -> phonemes: 26%|██▌ | 904/3515 [00:00<00:00, 4580.22it/s] word -> phonemes: 40%|████ | 1409/3515 [00:00<00:00, 4793.64it/s] word -> phonemes: 54%|█████▍ | 1903/3515 [00:00<00:00, 4850.44it/s] word -> phonemes: 68%|██████▊ | 2389/3515 [00:00<00:00, 4841.36it/s] word -> phonemes: 82%|████████▏ | 2881/3515 [00:00<00:00, 4866.32it/s] word -> phonemes: 96%|█████████▌| 3368/3515 [00:00<00:00, 4777.69it/s] word -> phonemes: 100%|██████████| 3515/3515 [00:00<00:00, 4770.31it/s] train: 0%| | 0/3515 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 21%|██ | 564/2703 [00:00<00:00, 5630.22it/s] word -> phonemes: 42%|████▏ | 1136/2703 [00:00<00:00, 5681.02it/s] word -> phonemes: 64%|██████▍ | 1729/2703 [00:00<00:00, 5794.02it/s] word -> phonemes: 85%|████████▌ | 2309/2703 [00:00<00:00, 5660.29it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5655.80it/s] dev: 0%| | 0/2703 [00:00