google/fleurs
Viewer • Updated • 768k • 99.2k • 445
How to use deepdml/whisper-tiny-es-mix-norm with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("automatic-speech-recognition", model="deepdml/whisper-tiny-es-mix-norm") # Load model directly
from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq
processor = AutoProcessor.from_pretrained("deepdml/whisper-tiny-es-mix-norm")
model = AutoModelForSpeechSeq2Seq.from_pretrained("deepdml/whisper-tiny-es-mix-norm", device_map="auto")This model is a fine-tuned version of openai/whisper-tiny on the Common Voice 17.0 dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Wer Raw | Cer Raw | Wer | Cer |
|---|---|---|---|---|---|---|---|
| 0.3392 | 0.0385 | 1000 | 0.5449 | 28.8875 | 10.7007 | 28.7842 | 10.6805 |
| 0.2970 | 0.0769 | 2000 | 0.4834 | 25.9414 | 9.4732 | 25.9167 | 9.4688 |
| 0.2931 | 0.1154 | 3000 | 0.4402 | 24.4542 | 9.0891 | 24.4377 | 9.0862 |
| 0.3371 | 0.1538 | 4000 | 0.4287 | 23.8031 | 8.9298 | 23.7936 | 8.9282 |
| 0.4178 | 0.1923 | 5000 | 0.4111 | 23.1951 | 8.6842 | 23.1925 | 8.6838 |
| 0.2361 | 0.2308 | 6000 | 0.3823 | 20.7530 | 7.6529 | 20.7524 | 7.6528 |
| 0.3117 | 0.2692 | 7000 | 0.3746 | 21.4333 | 8.3745 | 21.4314 | 8.3741 |
| 0.3198 | 0.3077 | 8000 | 0.3682 | 20.9369 | 7.8742 | 20.9369 | 7.8742 |
| 0.2277 | 1.0035 | 9000 | 0.3481 | 19.8135 | 7.4420 | 19.8135 | 7.4420 |
| 0.2065 | 1.0420 | 10000 | 0.3423 | 19.0578 | 7.0356 | 19.0578 | 7.0356 |
| 0.1846 | 1.0804 | 11000 | 0.3353 | 18.7769 | 6.9491 | 18.7769 | 6.9491 |
| 0.1821 | 1.1189 | 12000 | 0.3343 | 18.3040 | 6.6649 | 18.3040 | 6.6649 |
| 0.2689 | 1.1573 | 13000 | 0.3397 | 19.2632 | 7.2140 | 19.2632 | 7.2140 |
| 0.2155 | 1.1958 | 14000 | 0.3329 | 18.8315 | 7.1981 | 18.8315 | 7.1981 |
| 0.1951 | 1.2343 | 15000 | 0.3257 | 18.3072 | 6.7254 | 18.3072 | 6.7254 |
| 0.3108 | 1.2727 | 16000 | 0.3247 | 18.5639 | 7.0022 | 18.5639 | 7.0022 |
| 0.3006 | 1.3112 | 17000 | 0.3245 | 17.9553 | 6.5605 | 17.9553 | 6.5605 |
| 0.1631 | 2.007 | 18000 | 0.3128 | 17.6130 | 6.6351 | 17.6130 | 6.6351 |
| 0.1956 | 2.0455 | 19000 | 0.3122 | 17.9471 | 6.7893 | 17.9471 | 6.7893 |
| 0.1764 | 2.0839 | 20000 | 0.3151 | 17.6967 | 6.5664 | 17.6967 | 6.5664 |
| 0.2162 | 2.1224 | 21000 | 0.3135 | 17.6808 | 6.6306 | 17.6808 | 6.6306 |
| 0.1736 | 2.1608 | 22000 | 0.3110 | 17.2941 | 6.3855 | 17.2941 | 6.3855 |
| 0.1488 | 2.1993 | 23000 | 0.3088 | 17.4982 | 6.4258 | 17.4982 | 6.4258 |
| 0.3267 | 2.2378 | 24000 | 0.3105 | 17.6624 | 6.6477 | 17.6624 | 6.6477 |
| 0.1671 | 2.2762 | 25000 | 0.3077 | 17.4634 | 6.5241 | 17.4634 | 6.5241 |
| 0.1411 | 2.3147 | 26000 | 0.3069 | 17.3220 | 6.4775 | 17.3220 | 6.4775 |
Please cite the model using the following BibTeX entry:
@misc{deepdml/whisper-tiny-es-mix-norm,
title={Fine-tuned Whisper tiny ASR model for speech recognition in Spanish},
author={Jimenez, David},
howpublished={\url{https://huggingface.co/deepdml/whisper-tiny-es-mix-norm}},
year={2026}
}