SpeechLLM

SpeechLLM is a multi-modal LLM trained to predict the metadata of the speaker's turn in a conversation. speechllm-2B model is based on HubertX audio encoder and TinyLlama LLM. The model predicts the following:

SpeechActivity : if the audio signal contains speech (True/False)
Transcript : ASR transcript of the audio
Gender of the speaker (Female/Male)
Age of the speaker (Young/Middle-Age/Senior)
Accent of the speaker (Africa/America/Celtic/Europe/Oceania/South-Asia/South-East-Asia)
Emotion of the speaker (Happy/Sad/Anger/Neutral/Frustrated)

Usage

# Load model directly from huggingface
from transformers import AutoModel
model = AutoModel.from_pretrained("skit-ai/speechllm-2B", trust_remote_code=True)

model.generate_meta(
    audio_path="path-to-audio.wav", #16k Hz, mono
    audio_tensor=torchaudio.load("path-to-audio.wav")[1], # [Optional] either audio_path or audio_tensor directly
    instruction="Give me the following information about the audio [SpeechActivity, Transcript, Gender, Emotion, Age, Accent]",
    max_new_tokens=500, 
    return_special_tokens=False
)

# Model Generation
'''
{
  "SpeechActivity" : "True",
  "Transcript": "Yes, I got it. I'll make the payment now.",
  "Gender": "Female",
  "Emotion": "Neutral",
  "Age": "Young",
  "Accent" : "America",
}
'''

Try the model in Google Colab Notebook. Also, check out our blog on SpeechLLM for end-to-end conversational agents(User Speech -> Response).

Model Details

Developed by: Skit AI
Authors: Shangeth Rajaa, Abhinav Tushar
Language: English
Finetuned from model: HubertX, TinyLlama
Model Size: 2.1 B
Checkpoint: 2000 k steps (bs=1)
Adapters: r=4, alpha=8
lr : 1e-4
gradient accumulation steps: 8

Checkpoint Result

Dataset	Type	Word Error Rate	Gender Acc	Age Acc	Accent Acc
librispeech-test-clean	Read Speech	6.73	0.9496
librispeech-test-other	Read Speech	9.13	0.9217
CommonVoice test	Diverse Accent, Age	25.66	0.8680	0.6041	0.6959

Cite

@misc{Rajaa_SpeechLLM_Multi-Modal_LLM,
author = {Rajaa, Shangeth and Tushar, Abhinav},
title = {{SpeechLLM: Multi-Modal LLM for Speech Understanding}},
url = {https://github.com/skit-ai/SpeechLLM}
}

Downloads last month: 96

Safetensors

Model size

2B params

Tensor type

F32

Datasets used to train skit-ai/speechllm-2B

Collection including skit-ai/speechllm-2B

SpeechLLM

Collection

Multi-Modal speech LLMs • 2 items • Updated Jun 26, 2024 • 1

Evaluation results

Test WER on LibriSpeech (clean)
test set self-reported

6.730
Test WER on LibriSpeech (other)
test set self-reported

9.130
Test WER on Common Voice 16.1
test set self-reported

25.660
Test Age Accuracy on Common Voice 16.1
test set self-reported

60.410
Test Accent Accuracy on Common Voice 16.1
test set self-reported

69.590