Fill-Mask
Transformers
PyTorch
modernbert
orionweller SidTheChillGuy commited on
Commit
13f0442
·
verified ·
1 Parent(s): 621556d

Dataset Link Fix (#6)

Browse files

- Dataset Link Fix (84b3213cb9758d66029e9106322f438fecc67778)


Co-authored-by: Siddhant Mahajan <[email protected]>

Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -91,8 +91,8 @@ mmBERT training data is publicly available across different phases:
91
  | Phase | Dataset | Tokens | Description |
92
  |:------|:--------|:-------|:------------|
93
  | Pre-training P1 | [mmbert-pretrain-p1](https://huggingface.co/datasets/jhu-clsp/mmbert-pretrain-p1-fineweb2-langs) | 2.3T | 60 languages, foundational training |
94
- | Pre-training P2 | [mmbert-pretrain-p2](https://huggingface.co/datasets/jhu-clsp/mmbert-pretrain-p2-fineweb2-langs) | - | Extension data for pre-training phase |
95
- | Pre-training P3 | [mmbert-pretrain-p3](https://huggingface.co/datasets/jhu-clsp/mmbert-pretrain-p3-fineweb2-langs) | - | Final pre-training data |
96
  | Mid-training | [mmbert-midtraining](https://huggingface.co/datasets/jhu-clsp/mmbert-midtraining-data) | 600B | 110 languages, context extension to 8K |
97
  | Decay Phase | [mmbert-decay](https://huggingface.co/datasets/jhu-clsp/mmbert-decay-data) | 100B | 1833 languages, premium quality |
98
 
 
91
  | Phase | Dataset | Tokens | Description |
92
  |:------|:--------|:-------|:------------|
93
  | Pre-training P1 | [mmbert-pretrain-p1](https://huggingface.co/datasets/jhu-clsp/mmbert-pretrain-p1-fineweb2-langs) | 2.3T | 60 languages, foundational training |
94
+ | Pre-training P2 | [mmbert-pretrain-p2](https://huggingface.co/datasets/jhu-clsp/mmBERT-pretrain-p2-fineweb2-remaining) | - | Extension data for pre-training phase |
95
+ | Pre-training P3 | [mmbert-pretrain-p3](https://huggingface.co/datasets/jhu-clsp/mmBERT-pretrain-p3-others) | - | Final pre-training data |
96
  | Mid-training | [mmbert-midtraining](https://huggingface.co/datasets/jhu-clsp/mmbert-midtraining-data) | 600B | 110 languages, context extension to 8K |
97
  | Decay Phase | [mmbert-decay](https://huggingface.co/datasets/jhu-clsp/mmbert-decay-data) | 100B | 1833 languages, premium quality |
98