Upload merged SCI Assistant model (OpenHermes-2.5-Mistral-7B + SCI LoRA)

Browse files

Files changed (9) hide show

README.md +38 -353
config.json +26 -0
generation_config.json +6 -0
model-00001-of-00003.safetensors +3 -0
model-00002-of-00003.safetensors +3 -0
model-00003-of-00003.safetensors +3 -0
model.safetensors.index.json +298 -0
special_tokens_map.json +0 -7
tokenizer_config.json +1 -1

README.md CHANGED Viewed

@@ -1,380 +1,65 @@
----
-base_model: teknium/OpenHermes-2.5-Mistral-7B
-library_name: peft
-pipeline_tag: text-generation
-tags:
-- base_model:adapter:teknium/OpenHermes-2.5-Mistral-7B
-- lora
-- medical
-- spinal-cord-injury
-- healthcare
-- assistant
----
-# SCI Assistant - Spinal Cord Injury Specialized AI Assistant
-A specialized AI assistant fine-tuned specifically for people with spinal cord injuries (SCI). This model is based on OpenHermes-2.5-Mistral-7B and has been trained using a two-phase approach with LoRA (Low-Rank Adaptation) to provide contextually appropriate and medically-informed responses for the SCI community.
 ## Model Description
-This model was fine-tuned using a two-phase training approach:
-1. **Phase 1**: Domain pretraining on SCI-related medical texts and resources
-2. **Phase 2**: Instruction tuning on conversational SCI-focused Q&A pairs
-The model understands the unique challenges, medical realities, and daily life considerations of individuals living with spinal cord injuries.
-## Training Details
-- **Base Model**: teknium/OpenHermes-2.5-Mistral-7B
-- **Training Method**: QLoRA (4-bit quantization with LoRA adapters)
-- **Training Data**: 119,117 total entries (35,779 domain text + 83,337 instruction pairs)
-- **Hardware**: RTX 4070 Super (8GB VRAM)
-- **Training Time**: ~20 hours total (Phase 1 + Phase 2)
 ## Usage
 ```python
-from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
-from peft import PeftModel
-import torch
-# Load model
-bnb_config = BitsAndBytesConfig(
-    load_in_4bit=True,
-    bnb_4bit_compute_dtype=torch.float16,
-)
-base_model = AutoModelForCausalLM.from_pretrained(
-    "teknium/OpenHermes-2.5-Mistral-7B",
-    quantization_config=bnb_config,
-    device_map="auto"
-)
-model = PeftModel.from_pretrained(base_model, "basiphobe/sci-assistant")
-tokenizer = AutoTokenizer.from_pretrained("basiphobe/sci-assistant")
-# Format prompt with SCI context
-system_context = "You are a specialized medical assistant for people with spinal cord injuries. Your responses should always consider the unique needs, challenges, and medical realities of individuals living with SCI."
-prompt = f"{system_context}\n\n### Instruction:\n{your_question}\n\n### Response:\n"
-# Generate response
 inputs = tokenizer(prompt, return_tensors="pt")
-outputs = model.generate(**inputs, max_new_tokens=256, temperature=0.7)
-response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
 ```
 ## Intended Use
-This model is designed to:
-- Provide SCI-specific information and guidance
-- Answer questions about daily life with spinal cord injuries
-- Offer practical advice for common SCI challenges
-- Support the SCI community with contextually appropriate responses
 ## Limitations
-- This model is for informational purposes only and should not replace professional medical advice
-- Always consult with healthcare providers for medical decisions
-- The model may not have information about the latest medical developments
-- Responses should be verified with medical professionals when making health-related decisions
-## Direct Use
-This model can be used directly for:
-- Educational purposes about spinal cord injuries
-- Providing general information and support to the SCI community
-- Research into specialized medical AI assistants
-- Personal use by individuals seeking SCI-related information
-The model is designed to provide contextually appropriate responses that consider the unique challenges and medical realities of spinal cord injuries.
-### Downstream Use
-This model can be fine-tuned further for:
-- Integration into healthcare applications
-- Specialized medical chatbots for rehabilitation centers
-- Educational platforms for SCI awareness and training
-- Research applications in medical AI
-- Custom applications for SCI support organizations
-When used in downstream applications, implementers should:
-- Maintain the medical disclaimer requirements
-- Ensure proper supervision by medical professionals
-- Implement appropriate safety measures and content filtering
-- Validate outputs for medical accuracy in their specific use case
-### Out-of-Scope Use
-This model should NOT be used for:
-- **Medical diagnosis or treatment decisions** - Always consult healthcare professionals
-- **Emergency medical situations** - Seek immediate professional medical help
-- **Legal or financial advice** related to SCI cases
-- **Replacement for professional medical consultation**
-- **Clinical decision-making** without physician oversight
-- **Applications targeting vulnerable populations** without proper safeguards
-- **Commercial medical applications** without appropriate medical validation and oversight
-## Bias, Risks, and Limitations
-### Medical Limitations
-- **Not a substitute for medical professionals**: All medical advice should be verified with qualified healthcare providers
-- **Training data limitations**: May not include the most recent medical research or treatments
-- **Individual variation**: SCI affects individuals differently; responses may not apply to all cases
-- **Geographic bias**: Training data may be biased toward certain healthcare systems or regions
-### Technical Limitations
-- **Hallucination risk**: Like all language models, may generate plausible-sounding but incorrect information
-- **Context limitations**: Limited by input context window and may not retain information across long conversations
-- **Language limitations**: Primarily trained on English content
-- **Update lag**: Cannot access real-time medical research or current events
-### Bias Considerations
-- **Training data bias**: Reflects biases present in source medical literature and online content
-- **Demographic representation**: May not equally represent all demographics within the SCI community
-- **Healthcare access bias**: May reflect biases toward certain types of healthcare systems
-- **Severity bias**: May be more informed about certain types or severities of SCI
-### Risk Mitigation
-- Always include medical disclaimers when using this model
-- Implement content filtering for harmful or dangerous advice
-- Regular evaluation by medical professionals is recommended
-- Monitor outputs for accuracy and appropriateness
-### Recommendations
-<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
-### Recommendations
-Users should be aware of the following recommendations:
-**For Direct Users:**
-- Always verify medical information with qualified healthcare professionals
-- Use responses as educational/informational starting points, not definitive advice
-- Be aware that individual SCI experiences vary significantly
-- Seek immediate professional help for urgent medical concerns
-**For Developers/Implementers:**
-- Implement clear medical disclaimers in any application using this model
-- Provide easy access to professional medical resources alongside model responses
-- Consider implementing content filtering for potentially harmful advice
-- Regular review by medical professionals is strongly recommended
-- Ensure compliance with relevant healthcare regulations (HIPAA, etc.)
-**For Healthcare Organizations:**
-- Professional medical oversight is essential when implementing in clinical settings
-- Regular validation of model outputs against current medical standards
-- Integration should complement, not replace, professional medical consultation
-- Staff training on AI limitations and appropriate use cases
-## Training Details
-### Training Data
-The training dataset consisted of 119,117 carefully curated entries focused on spinal cord injury information:
-**Domain Pretraining Data (35,779 entries):**
-- Medical literature and research papers on SCI
-- Educational materials from reputable SCI organizations
-- Clinical guidelines and treatment protocols
-- Rehabilitation and therapy documentation
-- Patient education resources
-**Instruction Tuning Data (83,337 entries):**
-- SCI-focused question-answer pairs
-- Conversational examples with appropriate medical context
-- Real-world scenarios and practical advice situations
-- Educational Q&A formatted for instruction following
-All training data was filtered and curated to ensure:
-- Sources from reputable medical organizations and healthcare professionals
-- Content originally created or reviewed by medical professionals in the SCI field
-- Appropriate tone and sensitivity for SCI community
-- Removal of potentially harmful or dangerous advice
-- Proper medical disclaimers and context
-**Note**: While the source materials were created by medical professionals, this model itself has not undergone independent medical validation.
-### Training Procedure
-The model was trained using a two-phase approach with QLoRA (Quantized Low-Rank Adaptation):
-**Phase 1 - Domain Pretraining:**
-- Focus: Medical terminology and SCI-specific knowledge
-- Duration: 2 epochs (~8 hours)
-- Data: 35,779 domain text entries
-- Objective: Adapt base model to SCI medical domain
-**Phase 2 - Instruction Tuning:**
-- Focus: Conversational abilities and response formatting
-- Duration: 2 epochs (~12 hours)
-- Data: 83,337 instruction-response pairs
-- Objective: Teach appropriate response patterns and tone
-#### Preprocessing
-Training data underwent extensive preprocessing:
-- Content sourced from materials created by healthcare professionals
-- Sensitive content filtering and safety checks
-- Standardized formatting for instruction-following
-- Quality filtering to remove low-quality or inappropriate content
-- Tokenization optimization for efficient training
-#### Training Hyperparameters
-- **Training regime:** 4-bit quantization with LoRA adapters (QLoRA)
-- **Learning rate:** 2e-4 with cosine scheduling
-- **LoRA rank:** 16
-- **LoRA alpha:** 32
-- **LoRA dropout:** 0.05
-- **Target modules:** q_proj, v_proj
-- **Batch size:** 4 with gradient accumulation
-- **Max sequence length:** 512 tokens
-- **Optimizer:** AdamW with weight decay
-#### Speeds, Sizes, Times
-- **Total training time:** ~20 hours (8h Phase 1 + 12h Phase 2)
-- **Hardware:** RTX 4070 Super (8GB VRAM)
-- **Final model size:** 30MB (LoRA adapter only)
-- **Base model size:** 7B parameters (not included in adapter)
-- **Training throughput:** ~3.5 samples/second average
-- **Memory usage:** 6-7GB VRAM during training
-## Evaluation
-### Testing Data, Factors & Metrics
-#### Testing Data
-The model was evaluated using:
-- Held-out test set of SCI-related questions (500 samples)
-- Manual review of response quality and appropriateness
-- Comparative analysis against general-purpose models on SCI topics
-- Assessment of domain-specific knowledge retention
-**Note**: Evaluation was conducted by the model developer, not independent medical professionals.
-#### Factors
-Evaluation considered multiple factors:
-- **Medical accuracy**: Correctness of SCI-related information
-- **Appropriateness**: Sensitivity and tone for SCI community
-- **Contextual relevance**: Understanding of SCI-specific challenges
-- **Safety**: Avoidance of harmful or dangerous advice
-- **Completeness**: Comprehensive responses to complex questions
-#### Metrics
-- **Medical accuracy score**: Based on consistency with source medical literature (not independently validated)
-- **Appropriateness rating**: Developer assessment of tone and sensitivity (4.2/5.0 subjective rating)
-- **Response relevance**: SCI-specific context understanding (82% relevance score)
-- **Safety compliance**: No obviously harmful medical advice detected in test samples
-- **Response quality**: Perplexity improvements over base model for SCI domain
-### Results
-**Quantitative Results:**
-- 40% improvement in SCI domain perplexity over base model
-- Responses demonstrate consistency with source medical literature
-- 95% safety compliance (no obviously harmful medical advice detected)
-- 82% average relevance score for SCI-specific contexts
-**Qualitative Results:**
-- Responses demonstrate clear understanding of SCI terminology and concepts
-- Appropriate tone and sensitivity for disability community
-- Consistent inclusion of medical disclaimers
-- Good balance between being helpful and cautious about medical advice
-**Limitations of Evaluation:**
-- Evaluation conducted by model developer, not independent medical experts
-- No formal clinical validation or testing with SCI patients
-- Results based on consistency with training sources, not independent medical verification
-## Environmental Impact
-Training carbon emissions estimated using energy consumption data:
-- **Hardware Type:** RTX 4070 Super (8GB VRAM)
-- **Hours used:** ~20 hours total training time
-- **Cloud Provider:** Local training (personal hardware)
-- **Compute Region:** North America
-- **Carbon Emitted:** Approximately 2.1 kg CO2eq (estimated based on local energy grid)
-The use of QLoRA significantly reduced training time and energy consumption compared to full fine-tuning methods, making this a relatively efficient training approach.
-## Technical Specifications
-### Model Architecture and Objective
-- **Base Architecture:** Mistral 7B transformer model
-- **Adaptation Method:** QLoRA (Quantized Low-Rank Adaptation)
-- **Objective:** Causal language modeling with SCI domain specialization
-- **Quantization:** 4-bit precision for memory efficiency
-- **LoRA Configuration:** Rank-16 adapters on attention projection layers
-### Compute Infrastructure
-#### Hardware
-- **GPU:** NVIDIA RTX 4070 Super (8GB VRAM)
-- **CPU:** Modern multi-core processor
-- **RAM:** 32GB system memory
-- **Storage:** NVMe SSD for fast data loading
-#### Software
-- **Framework:** Transformers 4.36+, PEFT 0.16.0
-- **Training:** QLoRA with bitsandbytes quantization
-- **Environment:** Python 3.10+, PyTorch 2.0+, CUDA 12.1
-## Citation
-If you use this model in your research or applications, please cite:
-**BibTeX:**
-```bibtex
-@misc{sci_assistant_2025,
-  title={SCI Assistant: A Specialized AI Assistant for Spinal Cord Injury Support},
-  author={basiphobe},
-  year={2025},
-  howpublished={Hugging Face Model Repository},
-  url={https://huggingface.co/basiphobe/sci-assistant}
-}
-```
-**APA:**
-basiphobe. (2025). *SCI Assistant: A Specialized AI Assistant for Spinal Cord Injury Support*. Hugging Face. https://huggingface.co/basiphobe/sci-assistant
-## Glossary
-**SCI**: Spinal Cord Injury - damage to the spinal cord that results in temporary or permanent changes in function
-**QLoRA**: Quantized Low-Rank Adaptation - an efficient fine-tuning method that reduces memory requirements
-**Domain Pretraining**: Training phase focused on learning domain-specific terminology and knowledge
-**Instruction Tuning**: Training phase focused on learning conversational patterns and response formatting
-**Perplexity**: A metric measuring how well a language model predicts text (lower is better)
-**LoRA**: Low-Rank Adaptation - parameter-efficient fine-tuning technique
-## Model Card Authors
-**Primary Author:** basiphobe
-**Model Development:** Individual research project for SCI community support
-**Data Sources:** Curated from medical literature and educational materials created by healthcare professionals
-**Validation Status:** Model has not undergone independent medical professional validation
-## Model Card Contact
-For questions, issues, or feedback regarding this model:
-- **Hugging Face:** https://huggingface.co/basiphobe/sci-assistant
-- **Issues:** Please report issues through Hugging Face model repository
-- **Medical Concerns:** Always consult qualified healthcare professionals
-**Important Note:** This model is provided for educational and informational purposes. Always seek professional medical advice for health-related questions and decisions.
-### Framework versions
-- PEFT 0.16.0

+# SCI Assistant 7B
+A specialized language model for spinal cord injury (SCI) information and support, based on OpenHermes-2.5-Mistral-7B with custom LoRA fine-tuning.
 ## Model Description
+This model has been fine-tuned specifically to provide accurate, helpful information about spinal cord injuries, including:
+- **Medical information** about SCI conditions and symptoms
+- **Practical advice** for daily living with SCI
+- **Equipment recommendations** for wheelchairs, adaptive technology, etc.
+- **Exercise and rehabilitation** guidance
+- **Emotional support** and community resources
+## Training Data
+The model was trained on curated SCI-related content including:
+- Medical literature and research papers
+- Patient education materials
+- Community forums and discussions
+- Rehabilitation guides and resources
 ## Usage
 ```python
+from transformers import AutoModelForCausalLM, AutoTokenizer
+model = AutoModelForCausalLM.from_pretrained("your-username/sci-assistant-7b")
+tokenizer = AutoTokenizer.from_pretrained("your-username/sci-assistant-7b")
+# Example usage
+prompt = "What are the signs of autonomic dysreflexia?"
 inputs = tokenizer(prompt, return_tensors="pt")
+outputs = model.generate(**inputs, max_length=200)
+response = tokenizer.decode(outputs[0], skip_special_tokens=True)
 ```
 ## Intended Use
+- **Educational purposes** - Learning about SCI conditions and management
+- **Community support** - Providing accessible information to SCI community
+- **Research** - Supporting SCI-related research and development
 ## Limitations
+- This model provides educational information only
+- Always consult healthcare professionals for medical advice
+- Not a replacement for professional medical care
+- May not reflect the most recent medical developments
+## Technical Details
+- **Base Model**: teknium/OpenHermes-2.5-Mistral-7B
+- **Fine-tuning**: LoRA (Low-Rank Adaptation)
+- **Parameters**: ~7 billion
+- **Precision**: FP16
+## License
+Please respect the original OpenHermes-2.5 license terms.
+## Acknowledgments
+Built on the excellent OpenHermes-2.5-Mistral-7B model by Teknium.
+Training data curated from publicly available SCI educational resources.

config.json ADDED Viewed

	@@ -0,0 +1,26 @@

+{
+  "architectures": [
+    "MistralForCausalLM"
+  ],
+  "attention_dropout": 0.0,
+  "bos_token_id": 1,
+  "eos_token_id": 32000,
+  "head_dim": 128,
+  "hidden_act": "silu",
+  "hidden_size": 4096,
+  "initializer_range": 0.02,
+  "intermediate_size": 14336,
+  "max_position_embeddings": 32768,
+  "model_type": "mistral",
+  "num_attention_heads": 32,
+  "num_hidden_layers": 32,
+  "num_key_value_heads": 8,
+  "rms_norm_eps": 1e-05,
+  "rope_theta": 10000.0,
+  "sliding_window": 4096,
+  "tie_word_embeddings": false,
+  "torch_dtype": "float16",
+  "transformers_version": "4.50.3",
+  "use_cache": false,
+  "vocab_size": 32002
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,6 @@

+{
+  "_from_model_config": true,
+  "bos_token_id": 1,
+  "eos_token_id": 32000,
+  "transformers_version": "4.50.3"
+}

model-00001-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:0af3ba118f0a9418e007b7dfcb2b06cb43c229fd83687c11e9a62739844aeed9
+size 4943178624

model-00002-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:051612a9761d6d906d2317c79b6b51938f98495c73606ba383cb66ba3c98423f
+size 4999819232

model-00003-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8a019c618c1d3c5fe330a4afe1bda6100fcd8bde456aa7ea18e1f937112ea833
+size 4540532640

model.safetensors.index.json ADDED Viewed

	@@ -0,0 +1,298 @@

+{
+  "metadata": {
+    "total_size": 14483496960
+  },
+  "weight_map": {
+    "lm_head.weight": "model-00003-of-00003.safetensors",
+    "model.embed_tokens.weight": "model-00001-of-00003.safetensors",
+    "model.layers.0.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.0.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.0.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.1.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.1.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.1.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.10.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.10.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.10.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.10.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.10.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.10.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.10.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.10.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.10.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.11.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.11.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.11.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.11.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.11.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.11.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.11.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.11.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.11.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.12.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.12.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.12.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.12.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.12.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.12.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.12.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.12.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.12.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.13.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.13.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.13.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.13.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.13.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.13.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.13.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.13.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.13.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.14.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.14.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.14.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.14.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.14.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.14.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.14.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.14.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.14.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.15.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.15.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.15.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.15.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.15.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.15.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.15.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.15.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.15.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.16.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.16.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.16.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.16.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.16.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.16.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.16.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.16.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.16.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.17.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.17.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.17.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.17.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.17.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.17.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.17.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.17.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.17.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.18.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.18.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.18.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.18.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.18.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.18.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.18.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.18.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.18.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.19.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.19.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.19.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.19.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.19.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.19.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.19.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.19.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.19.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.2.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.2.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.2.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.20.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.20.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.20.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.20.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.20.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.20.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.20.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.20.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.20.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.21.input_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.21.mlp.down_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.21.mlp.gate_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.21.mlp.up_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.21.post_attention_layernorm.weight": "model-00002-of-00003.safetensors",
+    "model.layers.21.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.21.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.21.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.21.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.22.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.22.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.22.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.22.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.22.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.22.self_attn.k_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.22.self_attn.o_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.22.self_attn.q_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.22.self_attn.v_proj.weight": "model-00002-of-00003.safetensors",
+    "model.layers.23.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.23.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.23.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.23.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.23.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.23.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.23.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.23.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.23.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.24.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.24.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.24.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.24.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.24.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.24.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.24.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.24.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.24.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.25.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.25.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.25.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.25.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.25.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.25.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.25.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.25.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.25.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.26.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.26.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.26.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.26.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.26.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.26.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.26.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.26.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.26.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.27.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.27.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.27.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.27.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.27.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.27.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.27.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.27.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.27.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.28.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.28.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.28.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.28.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.28.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.28.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.28.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.28.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.28.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.29.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.29.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.29.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.29.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.29.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.29.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.29.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.29.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.29.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.3.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.3.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.3.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.30.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.30.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.30.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.30.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.30.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.30.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.30.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.30.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.30.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.31.input_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.31.mlp.down_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.31.mlp.gate_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.31.mlp.up_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.31.post_attention_layernorm.weight": "model-00003-of-00003.safetensors",
+    "model.layers.31.self_attn.k_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.31.self_attn.o_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.31.self_attn.q_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.31.self_attn.v_proj.weight": "model-00003-of-00003.safetensors",
+    "model.layers.4.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.4.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.4.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.5.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.5.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.5.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.5.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.5.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.6.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.6.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.6.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.6.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.6.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.6.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.6.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.6.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.6.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.7.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.7.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.7.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.7.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.7.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.7.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.7.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.7.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.7.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.8.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.8.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.8.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.8.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.8.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.8.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.8.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.8.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.8.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.9.input_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.9.mlp.down_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.9.mlp.gate_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.9.mlp.up_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.9.post_attention_layernorm.weight": "model-00001-of-00003.safetensors",
+    "model.layers.9.self_attn.k_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.9.self_attn.o_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.9.self_attn.q_proj.weight": "model-00001-of-00003.safetensors",
+    "model.layers.9.self_attn.v_proj.weight": "model-00001-of-00003.safetensors",
+    "model.norm.weight": "model-00003-of-00003.safetensors"
+  }
+}

special_tokens_map.json CHANGED Viewed

@@ -13,13 +13,6 @@
     "rstrip": false,
     "single_word": false
   },
-  "pad_token": {
-    "content": "<|im_end|>",
-    "lstrip": false,
-    "normalized": false,
-    "rstrip": false,
-    "single_word": false
-  },
   "unk_token": {
     "content": "<unk>",
     "lstrip": false,

     "rstrip": false,
     "single_word": false
   },
   "unk_token": {
     "content": "<unk>",
     "lstrip": false,

tokenizer_config.json CHANGED Viewed

@@ -52,7 +52,7 @@
   "extra_special_tokens": {},
   "legacy": true,
   "model_max_length": 1000000000000000019884624838656,
-  "pad_token": "<|im_end|>",
   "sp_model_kwargs": {},
   "spaces_between_special_tokens": false,
   "tokenizer_class": "LlamaTokenizer",

   "extra_special_tokens": {},
   "legacy": true,
   "model_max_length": 1000000000000000019884624838656,
+  "pad_token": null,
   "sp_model_kwargs": {},
   "spaces_between_special_tokens": false,
   "tokenizer_class": "LlamaTokenizer",