DeepSeek-R1-Distill-Qwen-32B

Permissive

DeepSeek · DeepSeek · Released 2025-01

32BText

A Qwen2.5-32B base fine-tuned on DeepSeek-R1 reasoning traces, bringing most of R1's reasoning gains to a size that fits on a single GPU.

Strengths

  • +Apache 2.0 license
  • +Reasoning performance close to much larger models
  • +Fits on a single 24-48GB GPU

Limitations

  • -Inherits verbose chain-of-thought behavior, higher token cost per answer
  • -Still below full DeepSeek-R1 on the hardest benchmarks

License

Apache 2.0

use commercially with attribution niceties

Hardware

wants 24-48GB of VRAM - a 3090/4090-class card or better

Links

Stats

deepseek-ai/DeepSeek-R1-Distill-Qwen-32B— downloads·— likes

via Hugging Face

Related models