·
AI & ML interests
I like to fine-tune the small models of the Doge series.
Organizations
Trainable Dynamic Mask Sparse Attention: Bridging Efficiency and Effectiveness in Long-Context Language Models
datasets
15
wubingheng/MixtureOfThoughts-Chinese-tryrun
Viewer
•
Updated
•
10
•
11
wubingheng/Mixture-of-Thoughts-zh-try-run
Viewer
•
Updated
•
10
•
11
wubingheng/Budget-aware-2048
Viewer
•
Updated
•
25k
•
12
wubingheng/Budget-aware-2048-in
Viewer
•
Updated
•
25k
•
113
wubingheng/Budget-aware-2048-in-try-run
Viewer
•
Updated
•
2
•
23
wubingheng/Budget-aware-2048-try-run
Viewer
•
Updated
•
2
•
9
Viewer
•
Updated
•
25k
•
7
Viewer
•
Updated
•
25k
•
5
wubingheng/compressed-openthoughts-50
Viewer
•
Updated
•
25k
•
13
wubingheng/compressed-openthoughts-90
Viewer
•
Updated
•
25k
•
16