TheBloke
/

orca_mini_v2_7B-GPTQ

+---
+inference: false
+license: other
+---
+<!-- header start -->
+<div style="width: 100%;">
+    <img src="https://i.imgur.com/EBdldam.jpg" alt="TheBlokeAI" style="width: 100%; min-width: 400px; display: block; margin: auto;">
+</div>
+<div style="display: flex; justify-content: space-between; width: 100%;">
+    <div style="display: flex; flex-direction: column; align-items: flex-start;">
+        <p><a href="https://discord.gg/theblokeai">Chat & support: my new Discord server</a></p>
+    </div>
+    <div style="display: flex; flex-direction: column; align-items: flex-end;">
+        <p><a href="https://www.patreon.com/TheBlokeAI">Want to contribute? TheBloke's Patreon page</a></p>
+    </div>
+</div>
+<!-- header end -->
+# Pankaj Mathur's Orca Mini v2 7B GPTQ
+These files are GPTQ 4bit model files for [Pankaj Mathur's Orca Mini v2 7B](https://huggingface.co/psmathur/orca_mini_v2_7b).
+It is the result of quantising to 4bit using [GPTQ-for-LLaMa](https://github.com/qwopqwop200/GPTQ-for-LLaMa).
+## Repositories available
+* [4-bit GPTQ models for GPU inference](https://huggingface.co/TheBloke/orca_mini_v2_7B-GPTQ)
+* [2, 3, 4, 5, 6 and 8-bit GGML models for CPU+GPU inference](https://huggingface.co/TheBloke/orca_mini_v2_7B-GGML)
+* [Unquantised fp16 model in pytorch format, for GPU inference and for further conversions](https://huggingface.co/psmathur/orca_mini_v2_7b)
+## How to easily download and use this model in text-generation-webui
+Please make sure you're using the latest version of text-generation-webui
+1. Click the **Model tab**.
+2. Under **Download custom model or LoRA**, enter `TheBloke/orca_mini_v2_7B-GPTQ`.
+3. Click **Download**.
+4. The model will start downloading. Once it's finished it will say "Done"
+5. In the top left, click the refresh icon next to **Model**.
+6. In the **Model** dropdown, choose the model you just downloaded: `orca_mini_v2_7B-GPTQ`
+7. The model will automatically load, and is now ready for use!
+8. If you want any custom settings, set them and then click **Save settings for this model** followed by **Reload the Model** in the top right.
+  * Note that you do not need to and should not set manual GPTQ parameters any more. These are set automatically from the file `quantize_config.json`.
+9. Once you're ready, click the **Text Generation tab** and enter a prompt to get started!
+## How to use this GPTQ model from Python code
+First make sure you have [AutoGPTQ](https://github.com/PanQiWei/AutoGPTQ) installed:
+`pip install auto-gptq`
+Then try the following example code:
+```python
+from transformers import AutoTokenizer, pipeline, logging
+from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
+import argparse
+model_name_or_path = "TheBloke/orca_mini_v2_7B-GPTQ"
+model_basename = "orca-mini-v2_7b-GPTQ-4bit-128g.no-act.order"
+use_triton = False
+tokenizer = AutoTokenizer.from_pretrained(model_name_or_path, use_fast=True)
+model = AutoGPTQForCausalLM.from_quantized(model_name_or_path,
+        model_basename=model_basename,
+        use_safetensors=True,
+        trust_remote_code=False,
+        device="cuda:0",
+        use_triton=use_triton,
+        quantize_config=None)
+# Note: check the prompt template is correct for this model.
+prompt = "Tell me about AI"
+prompt_template=f'''USER: {prompt}
+ASSISTANT:'''
+print("\n\n*** Generate:")
+input_ids = tokenizer(prompt_template, return_tensors='pt').input_ids.cuda()
+output = model.generate(inputs=input_ids, temperature=0.7, max_new_tokens=512)
+print(tokenizer.decode(output[0]))
+# Inference can also be done using transformers' pipeline
+# Prevent printing spurious transformers error when using pipeline with AutoGPTQ
+logging.set_verbosity(logging.CRITICAL)
+print("*** Pipeline:")
+pipe = pipeline(
+    "text-generation",
+    model=model,
+    tokenizer=tokenizer,
+    max_new_tokens=512,
+    temperature=0.7,
+    top_p=0.95,
+    repetition_penalty=1.15
+)
+print(pipe(prompt_template)[0]['generated_text'])
+```
+## Provided files
+**orca-mini-v2_7b-GPTQ-4bit-128g.no-act.order.safetensors**
+This will work with AutoGPTQ, ExLlama, and CUDA versions of GPTQ-for-LLaMa. There are reports of issues with Triton mode of recent GPTQ-for-LLaMa. If you have issues, please use AutoGPTQ instead.
+It was created with group_size 128 to increase inference accuracy, but without --act-order (desc_act) to increase compatibility and improve inference speed.
+* `orca-mini-v2_7b-GPTQ-4bit-128g.no-act.order.safetensors`
+  * Works with AutoGPTQ in CUDA or Triton modes.
+  * LLaMa models also work with [ExLlama](https://github.com/turboderp/exllama}, which usually provides much higher performance, and uses less VRAM, than AutoGPTQ.
+  * Works with GPTQ-for-LLaMa in CUDA mode.  May have issues with GPTQ-for-LLaMa Triton mode.
+  * Works with text-generation-webui, including one-click-installers.
+  * Parameters: Groupsize = 128. Act Order / desc_act = False.
+<!-- footer start -->
+## Discord
+For further support, and discussions on these models and AI in general, join us at:
+[TheBloke AI's Discord server](https://discord.gg/theblokeai)
+## Thanks, and how to contribute.
+Thanks to the [chirper.ai](https://chirper.ai) team!
+I've had a lot of people ask if they can contribute. I enjoy providing models and helping people, and would love to be able to spend even more time doing it, as well as expanding into new projects like fine tuning/training.
+If you're able and willing to contribute it will be most gratefully received and will help me to keep providing more models, and to start work on new AI projects.
+Donaters will get priority support on any and all AI/LLM/model questions and requests, access to a private Discord room, plus other benefits.
+* Patreon: https://patreon.com/TheBlokeAI
+* Ko-Fi: https://ko-fi.com/TheBlokeAI
+**Special thanks to**: Luke from CarbonQuill, Aemon Algiz.
+**Patreon special mentions**: Spiking Neurons AB, Kevin Schuppel, Cory Kujawski, senxiiz, Luke Pendergrass, John Villwock, Ghost , Alex , Sean Connelly, Space Cruiser, Eugene Pentland, Pyrater, Matthew Berman, Dave, Derek Yates, Jonathan Leane, Viktor Bowallius, Michael Levine, Joseph William Delisle, Fred von Graf, Asp the Wyvern, Nikolai Manek, Pierre Kircher, webtim, K, RoA, Karl Bernard, Artur Olbinski, Rainer Wilmers, Ai Maven, Nathan LeClaire, Ajan Kanaga, Stephen Murray, Edmond Seymore, zynix , Imad Khwaja, John Detwiler, Randy H, subjectnull, Alps Aficionado, Greatston Gnanesh, Trenton Dambrowitz, Junyu Yang, Raven Klaugh, biorpg, Deep Realms, vamX, Talal Aujan, Johann-Peter Hartmann, WelcomeToTheClub, Chris McCloskey, Luke, chris gileta, terasurfer , Iucharbius , Preetika Verma, Willem Michiel, Fen Risland, SuperWojo, Khalefa Al-Ahmad, Daniel P. Andersen, Gabriel Puliatti, Illia Dulskyi, Willian Hasse, Oscar Rangel, ya boyyy, Mano Prime, Lone Striker, Kalila.
+Thank you to all my generous patrons and donaters!
+<!-- footer end -->
+# Original model card: Pankaj Mathur's Orca Mini v2 7B
+# orca_mini_v2_7b
+An **Uncensored** LLaMA-7b model in collaboration with [Eric Hartford](https://huggingface.co/ehartford). trained on explain tuned datasets, created using Instructions and Input from WizardLM, Alpaca & Dolly-V2 datasets and applying Orca Research Paper dataset construction approaches.
+Please note this model has *better code generation capabilities* compare to our original orca_mini_7b which was trained on base OpenLLaMA-7b model and which has the [empty spaces issues & found not good for code generation]((https://github.com/openlm-research/open_llama#update-06072023)).
+**P.S. I am #opentowork, if you can help, please reach out to me at www.linkedin.com/in/pankajam**
+# Evaluation
+I evaluated orca_mini_v2_7b on a wide range of tasks using [Language Model Evaluation Harness](https://github.com/EleutherAI/lm-evaluation-harness) from EleutherAI.
+Here are the zero shot metrics results.
+|||||||
+|:------:|:-------------:|:---------:|:--------:|:-------:|:--------:|
+|**Task**|**num_fewshot**|**Version**|**Metric**|**Value**|**Stderr**|
+|*arc_easy*|0|0|acc|0.7386|0.0090|
+|*arc_easy*|0|0|acc_norm|0.7066|0.0093|
+|*hellaswag*|0|0|acc|0.5591|0.0050|
+|*hellaswag*|0|0|acc_norm|0.7394|0.0044|
+|*truthfulqa_mc*|0|1|mc1|0.2938|0.0159|
+|*truthfulqa_mc*|0|1|mc2|0.4399|0.0153|
+|*mmlu avg*|0|1|acc|0.4108|0.0153|
+|*mmlu avg*|0|1|acc_norm|0.4108|0.0153|
+|*Total Zero Shot Average*|0|-|-|0.5373|0.011|
+Here are the results on metrics used by [HuggingFaceH4 Open LLM Leaderboard](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
+please note num_fewshots varies for each below task as used by HuggingFaceH4 Open LLM Leaderboard
+|||||||
+|:------:|:-------------:|:---------:|:--------:|:-------:|:--------:|
+|**Task**|**num_fewshot**|**Version**|**Metric**|**Value**|**Stderr**|
+|*arc_challenge*|25|0|acc|0.4846|0.0146|
+|*arc_challenge*|25|0|acc_norm|0.5077|0.0146|
+# Dataset
+We used uncensored script on top of the previous explain tuned datasets we build which are [WizardLM dataset ~70K](https://github.com/nlpxucan/WizardLM), [Alpaca dataset ~52K](https://crfm.stanford.edu/2023/03/13/alpaca.html)  & [Dolly-V2 dataset ~15K](https://github.com/databrickslabs/dolly) created using approaches from [Orca Research Paper](https://arxiv.org/abs/2306.02707).
+We leverage all of the 15 system instructions provided in Orca Research Paper. to generate custom datasets, in contrast to vanilla instruction tuning approaches used by original datasets.
+This helps student model aka this model to learn ***thought*** process from teacher model, which is ChatGPT (gpt-3.5-turbo-0301 version).
+Please see below example usage how the **System** prompt is added before each **instruction**.
+# Training
+The training configurations are provided in the table below.
+The training takes on 8x A100(80G) GPUs and lasts for around 13 Hours for cost of $195 using [RunPods](https://www.runpod.io/)
+We used DeepSpeed with fully sharded data parallelism, also know as [ZeRO stage 3](https://engineering.fb.com/2021/07/15/open-source/fsdp/) by writing our own fine tunning scripts plus leveraging some of the model training code provided by amazing [OpenAlpaca repo](https://github.com/yxuansu/OpenAlpaca)
+Here are some of params used during training:
+|||
+|:-------------:|:-------------:|
+|*batch_size*|96|
+|*train_micro_batch_size_per_gpu*|3|
+|*gradient_accumulation_steps*|4|
+|*Learning rate*|2e-5|
+|*Max length*|1024|
+|*Epochs*|3|
+|*Optimizer*|AdamW|
+# Example Usage
+Here is prompt format for [Oobabooga Text generation UI ](https://github.com/oobabooga/text-generation-webui)
+```
+### System:
+{system}
+### User:
+{instruction}
+### Input:
+{input}
+### Response:
+```
+Here is sample example:
+```
+### System:
+You are an AI assistant that follows instruction extremely well. Help as much as you can.
+### User:
+Tell me how to break into my own car
+### Input:
+### Response:
+Breaking into your own car requires certain skills and tools. Here are the basic steps:
+1. Find a ^^^^^^^^^^^^^
+2. Unlock the car by using the ^^^^^^^^^^^^^.
+3. Use a ^^^^^^^^^^^^^.
+4. Once the ^^^^^^^^^^^^^.
+5. If the ^^^^^^^^^^^^^.
+```
+Below shows a code example on how to use this model
+```python
+import torch
+from transformers import LlamaForCausalLM, LlamaTokenizer
+# Hugging Face model_path
+model_path = 'psmathur/orca_mini_v2_7b'
+tokenizer = LlamaTokenizer.from_pretrained(model_path)
+model = LlamaForCausalLM.from_pretrained(
+    model_path, torch_dtype=torch.float16, device_map='auto',
+)
+#generate text function
+def generate_text(system, instruction, input=None):
+    if input:
+        prompt = f"### System:\n{system}\n\n### User:\n{instruction}\n\n### Input:\n{input}\n\n### Response:\n"
+    else:
+        prompt = f"### System:\n{system}\n\n### User:\n{instruction}\n\n### Response:\n"
+    tokens = tokenizer.encode(prompt)
+    tokens = torch.LongTensor(tokens).unsqueeze(0)
+    tokens = tokens.to('cuda')
+    instance = {'input_ids': tokens,'top_p': 1.0, 'temperature':0.7, 'generate_len': 1024, 'top_k': 50}
+    length = len(tokens[0])
+    with torch.no_grad():
+        rest = model.generate(
+            input_ids=tokens,
+            max_length=length+instance['generate_len'],
+            use_cache=True,
+            do_sample=True,
+            top_p=instance['top_p'],
+            temperature=instance['temperature'],
+            top_k=instance['top_k']
+        )
+    output = rest[0][length:]
+    string = tokenizer.decode(output, skip_special_tokens=True)
+    return f'[!] Response: {string}'
+# Sample Test Instruction
+system = 'You are an AI assistant that follows instruction extremely well. Help as much as you can.'
+instruction = 'Tell me how to break into my own car'
+print(generate_text(system, instruction))
+```
+**NOTE: The real response is hidden here with ^^^^^^^^^^^^^.**
+```
+[!] Response:
+Breaking into your own car requires certain skills and tools. Here are the basic steps:
+1. Find a ^^^^^^^^^^^^^
+2. Unlock the car by using the ^^^^^^^^^^^^^.
+3. Use a ^^^^^^^^^^^^^.
+4. Once the ^^^^^^^^^^^^^.
+5. If the ^^^^^^^^^^^^^.
+```
+Next Goals:
+1) Try more data like actually using FLAN-v2, just like Orka Research Paper (I am open for suggestions)
+2) Provide more options for Text generation UI. (may be https://github.com/oobabooga/text-generation-webui)
+3) Provide 4bit GGML/GPTQ quantized model (may be [TheBloke](https://huggingface.co/TheBloke) can help here)
+Limitations & Biases:
+This model can produce factually incorrect output, and should not be relied on to produce factually accurate information.
+This model was trained on various public datasets. While great efforts have been taken to clean the pretraining data, it is possible that this model could generate lewd, biased or otherwise offensive outputs.
+Disclaimer:
+The license on this model does not constitute legal advice. We are not responsible for the actions of third parties who use this model.
+Please cosult an attorney before using this model for commercial purposes.
+Citiation:
+If you found wizardlm_alpaca_dolly_orca_open_llama_7b useful in your research or applications, please kindly cite using the following BibTeX:
+```
+@misc{orca_mini_v2_7b,
+  author = {Pankaj Mathur},
+  title = {orca_mini_v2_7b: An explain tuned LLaMA-7b model on uncensored wizardlm, alpaca, & dolly datasets},
+  year = {2023},
+  publisher = {GitHub, HuggingFace},
+  journal = {GitHub repository, HuggingFace repository},
+  howpublished = {\url{https://https://huggingface.co/psmathur/orca_mini_v2_7b},
+}
+```
+```
+@software{touvron2023llama,
+  title={LLaMA: Open and Efficient Foundation Language Models},
+  author={Touvron, Hugo and Lavril, Thibaut and Izacard, Gautier and Martinet, Xavier and Lachaux, Marie-Anne and Lacroix, Timoth{\'e}e and Rozi{\`e}re, Baptiste and Goyal, Naman and Hambro, Eric and Azhar, Faisal and Rodriguez, Aurelien and Joulin, Armand and Grave, Edouard and Lample, Guillaume},
+  journal={arXiv preprint arXiv:2302.13971},
+  year={2023}
+}
+```
+```
+@misc{openalpaca,
+  author = {Yixuan Su and Tian Lan and Deng Cai},
+  title = {OpenAlpaca: A Fully Open-Source Instruction-Following Model Based On OpenLLaMA},
+  year = {2023},
+  publisher = {GitHub},
+  journal = {GitHub repository},
+  howpublished = {\url{https://github.com/yxuansu/OpenAlpaca}},
+}
+```
+```
+@misc{alpaca,
+  author = {Rohan Taori and Ishaan Gulrajani and Tianyi Zhang and Yann Dubois and Xuechen Li and Carlos Guestrin and Percy Liang and Tatsunori B. Hashimoto },
+  title = {Stanford Alpaca: An Instruction-following LLaMA model},
+  year = {2023},
+  publisher = {GitHub},
+  journal = {GitHub repository},
+  howpublished = {\url{https://github.com/tatsu-lab/stanford_alpaca}},
+}
+```
+```
+@online{DatabricksBlog2023DollyV2,
+    author    = {Mike Conover and Matt Hayes and Ankit Mathur and Jianwei Xie and Jun Wan and Sam Shah and Ali Ghodsi and Patrick Wendell and Matei Zaharia and Reynold Xin},
+    title     = {Free Dolly: Introducing the World's First Truly Open Instruction-Tuned LLM},
+    year      = {2023},
+    url       = {https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm},
+    urldate   = {2023-06-30}
+}
+```
+```
+@misc{xu2023wizardlm,
+      title={WizardLM: Empowering Large Language Models to Follow Complex Instructions},
+      author={Can Xu and Qingfeng Sun and Kai Zheng and Xiubo Geng and Pu Zhao and Jiazhan Feng and Chongyang Tao and Daxin Jiang},
+      year={2023},
+      eprint={2304.12244},
+      archivePrefix={arXiv},
+      primaryClass={cs.CL}
+}
+```