Qwen: download, install, fine-tune and use
Introduction
Qwen is the open-weight LLM family from Alibaba (QwenLM).
Current generations are Qwen3 / Qwen3.5 (new) and
Qwen2.5 (widely deployed). Weights are published on
Hugging Face (Qwen/Qwen3-8B,
Qwen/Qwen2.5-7B-Instruct, etc.) and on ModelScope.
This manual covers the full loop:
- Pick a size.
- Download and install it (Ollama, Transformers, vLLM, or llama.cpp).
- Use it (chat, Python, HTTP API).
- Fine-tune it on your data (LoRA / QLoRA) and run the result.
«Download» and «install» mean different things for LLMs: you download weights (tens of GB) once, and you install a runtime (Ollama, vLLM, transformers, llama.cpp) that loads them.
Which Qwen to download
Ollama names (easiest):
- qwen3:0.6b
- qwen3:1.7b
- qwen3:4b
- qwen3:8b
- qwen3:14b
- qwen3:32b
Hugging Face names:
- Qwen/Qwen3-0.6B
- Qwen/Qwen3-8B
- Qwen/Qwen2.5-7B-Instruct
Rules of thumb:
- Chat / RAG / tools: take an -Instruct checkpoint.
- Continue pretraining / base for your own SFT: take the base (non-Instruct) checkpoint.
- Coding: take a Qwen2.5-Coder- or Qwen3- checkpoint.
- Vision: take a Qwen2.5-VL- / Qwen3-VL- checkpoint, not the text-only one.
Check the model card license before commercial use. Most Qwen weights use Apache 2.0, but confirm on the exact repository page (huggingface.co/Qwen, github.com/QwenLM/Qwen3).
Hardware requirements
Rough VRAM for full-precision (bf16) inference, context 4-8k:
| Model | bf16 VRAM | 4-bit (GGUF / AWQ / GPTQ) |
|---|---|---|
| 0.6B - 1.7B | 2-5 GB | runs on CPU / laptop |
| 4B - 8B | 10-20 GB | 6-10 GB VRAM |
| 14B | ~30 GB | 10-12 GB VRAM |
| 32B | ~65 GB | 20-24 GB VRAM |
Fine-tuning needs more than inference. Full fine-tune of 7-8B needs ~60 GB VRAM; LoRA needs ~16-24 GB; QLoRA (4-bit base + LoRA adapters) fits 7-8B on one 16 GB card. CPU-only training is not practical beyond toy tests.
You need: Linux / macOS / WSL2 on Windows, Python 3.10+, NVIDIA drivers + CUDA for GPU runs, 30-80 GB free disk (weights + cache).
Fastest start: Ollama
Ollama downloads the weights and installs the runtime in one command. Best for «just chat now».
1. Install Ollama (Linux / macOS / WSL2)
curl -fsSL https://ollama.com/install.sh | sh
2. Download Qwen (pull once, reuse offline later)
ollama pull qwen3:8b
3. Chat
ollama run qwen3:8b
Smaller / bigger alternatives:
ollama pull qwen3:4b
ollama pull qwen3:14b
Ollama also serves an OpenAI-compatible HTTP API on port 11434:
List models:
curl http://localhost:11434/api/tags
Chat via API:
curl http://localhost:11434/api/chat -d '{"model": "qwen3:8b", "messages": [{"role": "user", "content": "Hello, Qwen!"}]}'
Upgrade weights later with ollama pull qwen3:8b again. Remove with ollama rm qwen3:8b.
Install via Hugging Face Transformers
Use this when you want Python-level control (scripts, training, embeddings).
1. Python env
python3 -m venv qwen-env
source qwen-env/bin/activate
pip install -U pip
pip install -U transformers accelerate torch --index-url https://download.pytorch.org/whl/cu121
Optional: flash attention for long context / speed
pip install -U flash-attn --no-build-isolation
2. Login once if you hit gated repos / rate limits
pip install -U huggingface_hub
hf auth login
3. Download is automatic on first use (cached in ~/.cache/huggingface). Or pre-fetch explicitly:
hf download Qwen/Qwen3-8B
from transformers import AutoModelForCausalLM, AutoTokenizer MODEL_ID = "Qwen/Qwen3-8B" # or "Qwen/Qwen2.5-7B-Instruct" tok = AutoTokenizer.from_pretrained(MODEL_ID) model = AutoModelForCausalLM.from_pretrained( MODEL_ID, dtype="auto", # bf16 on GPU, fp32 fallback on CPU device_map="auto", # spread across GPU/CPU automatically ) messages = [{"role": "user", "content": "Write a haiku about the Volga."}] text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tok([text], return_tensors="pt").to(model.device) out = model.generate(**inputs, max_new_tokens=256) # type: ignore[misc] print(tok.decode(out[0][len(inputs.input_ids[0]) :], skip_special_tokens=True))
Qwen3 may print a /think / /no_think switch
in chat templates. Pass enable_thinking=False to
apply_chat_template if you want direct answers without reasoning traces.
The # type: ignore[misc] on the
generate line silences a known mypy false positive in
transformers 5.x stubs (the declared from_pretrained
return type does not match the generate self type);
the call itself is correct at runtime.
Serve with vLLM (OpenAI-compatible API)
Use vLLM for high-throughput serving with an OpenAI-compatible endpoint. First run downloads weights from Hugging Face automatically.
pip install -U vllm
Serve Qwen3-8B on one GPU:
python -m vllm.entrypoints.openai.api_server --model Qwen/Qwen3-8B --dtype auto --max-model-len 8192 --port 8000
Test it:
curl http://localhost:8000/v1/models
curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model": "Qwen/Qwen3-8B", "messages": [{"role": "user", "content": "What is Qwen?"}]}'
Python client (same code works against real OpenAI by changing base_url):
from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") resp = client.chat.completions.create( model="Qwen/Qwen3-8B", messages=[{"role": "user", "content": "Explain LoRA in one paragraph."}], ) print(resp.choices[0].message.content)
Run on CPU with llama.cpp
For machines without a GPU, use a quantized GGUF build:
Ollama already does this for you; manual path:
1. Get a GGUF file, e.g. Qwen3-8B-Q4_K_M.gguf from a GGUF repo
hf download Qwen/Qwen3-8B-GGUF --include "*Q4_K_M*.gguf"
2a. llama.cpp server (OpenAI-compatible on :8080)
llama-server -m Qwen3-8B-Q4_K_M.gguf -c 4096 --port 8080
2b. or plain chat in terminal
llama-cli -m Qwen3-8B-Q4_K_M.gguf -c 4096 -n 512 -p "You are helpful. User: Hi! Assistant:"
How to use Qwen
- Chat template matters. Always format prompts with apply_chat_template (Transformers) or the messages API (Ollama / vLLM). Raw concatenation degrades Instruct models.
- System prompt sets behavior, e.g. You are a concise senior Python reviewer. Answer with code first.
- Sampling: start with temperature 0.7, top_p 0.8 for chat; temperature 0.0-0.2 for factual / coding tasks.
- RAG: point the model at your docs via LangChain / LlamaIndex and the OpenAI-compatible endpoint above - no retraining needed for knowledge cutoff issues.
- Tools / function calling: Qwen supports tool calls via its chat template and via the OpenAI-compatible tools parameter in vLLM / Ollama.
Fine-tune Qwen (LoRA / QLoRA)
Do not full-fine-tune on one GPU - use LoRA (small trainable adapters) or QLoRA (4-bit frozen base + LoRA). Typical recipe for a chat / support / domain-QA style adaptation of Qwen3-8B or Qwen2.5-7B on a single 16-24 GB GPU:
1. Prepare data as chat JSONL (one object per line):
{"messages": [ {"role": "system", "content": "You are a bike-shop assistant."}, {"role": "user", "content": "Which brake pads fit model X?"}, {"role": "assistant", "content": "Resin pads, type B. ..."}]}
2. Install the training stack:
source qwen-env/bin/activate
pip install -U trl peft bitsandbytes datasets
3. Train (QLoRA, ~works for 7-8B on 16 GB VRAM):
import torch from datasets import load_dataset from peft import LoraConfig from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig from trl import SFTConfig, SFTTrainer MODEL_ID = "Qwen/Qwen3-8B" # or "Qwen/Qwen2.5-7B-Instruct" ds = load_dataset("json", data_files="train.jsonl", split="train") tok = AutoTokenizer.from_pretrained(MODEL_ID, use_fast=True) if tok.pad_token is None: tok.pad_token = tok.eos_token bnb = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, ) model = AutoModelForCausalLM.from_pretrained( MODEL_ID, quantization_config=bnb, device_map="auto", ) model.config.use_cache = False peft_cfg = LoraConfig( r=16, lora_alpha=32, lora_dropout=0.05, target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"], task_type="CAUSAL_LM", ) args = SFTConfig( output_dir="qwen-lora", num_train_epochs=2, per_device_train_batch_size=2, gradient_accumulation_steps=8, learning_rate=2e-4, logging_steps=10, save_steps=200, bf16=True, packing=True, max_length=2048, ) trainer = SFTTrainer( model=model, train_dataset=ds, processing_class=tok, peft_config=peft_cfg, args=args, ) trainer.train() trainer.save_model("qwen-lora") tok.save_pretrained("qwen-lora")
4. Practical rules:
- Start with 1-3 epochs, lr 1e-4 - 2e-4, rank 8-32.
- Keep 5-10% of data as a held-out eval set; watch eval loss, not just train loss.
- 500-5000 good examples beat 100k noisy ones.
- Alternatives with less code: LLaMA-Factory (web UI + YAML configs) or Unsloth (fastest single-GPU QLoRA kernels).
- For ModelScope users (China mirror): replace Qwen/Qwen3-8B with the ModelScope id and pip install modelscope.
Export and run your fine-tune
LoRA output is adapters only. Merge them back for serving:
import torch from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer BASE = "Qwen/Qwen3-8B" tok = AutoTokenizer.from_pretrained(BASE) model = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16, device_map="cpu") peft_model = PeftModel.from_pretrained(model, "qwen-lora") merged = peft_model.merge_and_unload() merged.save_pretrained("qwen-merged") tok.save_pretrained("qwen-merged")
Then either serve qwen-merged with vLLM / Transformers, or quantize to GGUF with llama.cpp (llama-quantize / convert_hf_to_gguf.py) and run it with llama-server or import into Ollama via a Modelfile (FROM ./qwen-merged.gguf + ollama create my-qwen -f Modelfile).
Troubleshooting
- CUDA out of memory: drop to a smaller Qwen (8b -> 4b), shorten --max-model-len / max_length, use QLoRA instead of LoRA, set batch size 1 with more accumulation steps.
- Slow CPU inference: expected - use a smaller GGUF (Q4_K_M) and fewer context tokens (-c 2048).
- Gibberish / no answer: you skipped the chat template. Use apply_chat_template / messages, not raw strings.
- Model not found on HF: typo in id, or gated repo - run hf auth login and accept the license on the web page.
- Wrong language / style after SFT: training data too narrow - mix in 10-20% general chat examples to prevent catastrophic forgetting.
Official sources:
- github.com/QwenLM/Qwen3
- huggingface.co/Qwen
- ollama.com/library/qwen3
- docs.vllm.ai
Article author: Andrei Olegovich
| AI | |
| Qwen: download, install, fine-tune and use | |
| OpenRouter: one API key for 300+ AI models (account, API, OpenCode, VSCode) |