DeepSeek-R1-Distill-Qwen-32B
PermissiveDeepSeek · DeepSeek · Released 2025-01
32BText
A Qwen2.5-32B base fine-tuned on DeepSeek-R1 reasoning traces, bringing most of R1's reasoning gains to a size that fits on a single GPU.
Strengths
- +Apache 2.0 license
- +Reasoning performance close to much larger models
- +Fits on a single 24-48GB GPU
Limitations
- -Inherits verbose chain-of-thought behavior, higher token cost per answer
- -Still below full DeepSeek-R1 on the hardest benchmarks
License
Apache 2.0
use commercially with attribution niceties
Hardware
wants 24-48GB of VRAM - a 3090/4090-class card or better
Links
Stats
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B— downloads·— likes
via Hugging Face