Performance LLMs - Base Models
Collection
22 items
•
Updated
•
7
This is a llamafied qwen-14b for compatibility with the llama software ecosystem.
I used this script to make the model and used the tokenizer of CausalLM, as suggested in the comments of the script.
https://github.com/hiyouga/LLaMA-Factory/blob/main/tests/llamafy_qwen.py
Detailed results can be found here
Metric | Value |
---|---|
Avg. | 63.09 |
AI2 Reasoning Challenge (25-Shot) | 55.20 |
HellaSwag (10-Shot) | 82.31 |
MMLU (5-Shot) | 66.11 |
TruthfulQA (0-shot) | 45.60 |
Winogrande (5-shot) | 76.56 |
GSM8k (5-shot) | 52.77 |