/home/tony/Work/glockenspiel/s3prl/s3prl/upstream/byol_s/byol_a/common.py:20: UserWarning: torchaudio._backend.set_audio_backend has been deprecated. With dispatcher enabled, this function is no-op. You can remove the function call. torchaudio.set_audio_backend("sox_io") 2023-10-13 22:54:05 | WARNING | root | Pytorch pre-release version 2.1.0.dev20230831+cu118 - assuming intent to test it /home/tony/Work/glockenspiel/s3prl/s3prl/run_downstream.py:157: UserWarning: torchaudio._backend.set_audio_backend has been deprecated. With dispatcher enabled, this function is no-op. You can remove the function call. torchaudio.set_audio_backend('sox_io') [Featurizer] - Take a list of 13 features and weighted sum them. [Featurizer] - The selected feature hidden_states's downsample rate is 960 [Runner] - Start a new experiment mert_test_25hz {'model_config': None, 'refresh': False} Loading model from Local mert_test_25hz load local model succeed load local feat extractor Wav2Vec2FeatureExtractor { "do_normalize": false, "feature_extractor_type": "Wav2Vec2FeatureExtractor", "feature_size": 1, "padding_side": "right", "padding_value": 0, "return_attention_mask": true, "sampling_rate": 24000 } overall: 0%| | 0/4000 [00:00 phonemes: 0%| | 0/3515 [00:00 phonemes: 12%|█▏ | 414/3515 [00:00<00:00, 4137.63it/s] word -> phonemes: 24%|██▎ | 828/3515 [00:00<00:00, 4021.55it/s] word -> phonemes: 36%|███▋ | 1276/3515 [00:00<00:00, 4225.53it/s] word -> phonemes: 48%|████▊ | 1701/3515 [00:00<00:00, 4233.03it/s] word -> phonemes: 61%|██████ | 2142/3515 [00:00<00:00, 4293.96it/s] word -> phonemes: 73%|███████▎ | 2572/3515 [00:00<00:00, 4248.66it/s] word -> phonemes: 85%|████████▌ | 2998/3515 [00:00<00:00, 4248.85it/s] word -> phonemes: 97%|█████████▋| 3424/3515 [00:00<00:00, 4246.33it/s] word -> phonemes: 100%|██████████| 3515/3515 [00:00<00:00, 4214.96it/s] train: 0%| | 0/3515 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 18%|█▊ | 483/2703 [00:00<00:00, 4823.70it/s] word -> phonemes: 36%|███▌ | 966/2703 [00:00<00:00, 4783.70it/s] word -> phonemes: 54%|█████▍ | 1471/2703 [00:00<00:00, 4900.78it/s] word -> phonemes: 73%|███████▎ | 1962/2703 [00:00<00:00, 4541.32it/s] word -> phonemes: 90%|█████████ | 2438/2703 [00:00<00:00, 4614.28it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 4654.65it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 22%|██▏ | 602/2703 [00:00<00:00, 6012.45it/s] word -> phonemes: 45%|████▍ | 1204/2703 [00:00<00:00, 5586.26it/s] word -> phonemes: 68%|██████▊ | 1840/2703 [00:00<00:00, 5923.09it/s] word -> phonemes: 90%|█████████ | 2435/2703 [00:00<00:00, 5683.56it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5717.26it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 20%|█▉ | 531/2703 [00:00<00:00, 5301.71it/s] word -> phonemes: 40%|████ | 1087/2703 [00:00<00:00, 5452.02it/s] word -> phonemes: 60%|██████ | 1633/2703 [00:00<00:00, 5408.93it/s] word -> phonemes: 82%|████████▏ | 2214/2703 [00:00<00:00, 5563.00it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5423.37it/s] dev: 0%| | 0/2703 [00:00 phonemes: 0%| | 0/3515 [00:00 phonemes: 11%|█▏ | 398/3515 [00:00<00:00, 3976.93it/s] word -> phonemes: 24%|██▎ | 831/3515 [00:00<00:00, 4182.29it/s] word -> phonemes: 36%|███▌ | 1271/3515 [00:00<00:00, 4278.24it/s] word -> phonemes: 49%|████▉ | 1725/3515 [00:00<00:00, 4380.77it/s] word -> phonemes: 62%|██████▏ | 2171/3515 [00:00<00:00, 4408.85it/s] word -> phonemes: 74%|███████▍ | 2613/3515 [00:00<00:00, 4411.88it/s] word -> phonemes: 87%|████████▋ | 3062/3515 [00:00<00:00, 4437.14it/s] word -> phonemes: 100%|██████████| 3515/3515 [00:00<00:00, 4391.50it/s] train: 0%| | 0/3515 [00:00 phonemes: 0%| | 0/2703 [00:00 phonemes: 18%|█▊ | 478/2703 [00:00<00:00, 4777.44it/s] word -> phonemes: 39%|███▉ | 1055/2703 [00:00<00:00, 5357.25it/s] word -> phonemes: 61%|██████ | 1640/2703 [00:00<00:00, 5580.56it/s] word -> phonemes: 81%|████████▏ | 2199/2703 [00:00<00:00, 5480.91it/s] word -> phonemes: 100%|██████████| 2703/2703 [00:00<00:00, 5427.88it/s] dev: 0%| | 0/2703 [00:00