Yi Cui

onekq

AI & ML interests

Benchmark, Code Generation Model

Recent Activity

updated a model about 8 hours ago
onekq-ai/Qwen2.5-14B-Instruct-1M-bnb-4bit
updated a model about 8 hours ago
onekq-ai/Qwen2.5-7B-Instruct-1M-bnb-4bit
published a model about 9 hours ago
onekq-ai/Qwen2.5-14B-Instruct-1M-bnb-4bit
View all activity

Articles

Organizations

MLX Community's profile picture ONEKQ AI's profile picture

Posts 11

view post
Post
1779
So πŸ‹DeepSeekπŸ‹ hits the mainstream media. But it has been a star in our little cult for at least 6 months. Its meteoric success is not overnight, but two years in the making.

To learn their history, just look at their πŸ€— repo https://huggingface.co/deepseek-ai

* End of 2023, they launched the first model (pretrained by themselves) following Llama 2 architecture
* June 2024, v2 (MoE architecture) surpassed Gemini 1.5, but behind Mistral
* September, v2.5 surpassed GPT 4o mini
* December, v3 surpassed GPT 4o
* Now R1 surpassed o1

Most importantly, if you think DeepSeek success is singular and unrivaled, that's WRONG. The following models are also near or equal the o1 bar.

* Minimax-01
* Kimi k1.5
* Doubao 1.5 pro

models

None public yet

datasets

None public yet