Qwen: download, install, fine-tune and use

Contents
Introduction
Which Qwen to download
Hardware requirements
Fastest start: Ollama
Install via Hugging Face Transformers
Serve with vLLM (OpenAI-compatible API)
Run on CPU with llama.cpp
How to use Qwen
Fine-tune Qwen (LoRA / QLoRA)
Export and run your fine-tune
Troubleshooting

Introduction

Qwen is the open-weight LLM family from Alibaba (QwenLM). Current generations are Qwen3 / Qwen3.5 (new) and Qwen2.5 (widely deployed). Weights are published on Hugging Face (Qwen/Qwen3-8B, Qwen/Qwen2.5-7B-Instruct, etc.) and on ModelScope.

This manual covers the full loop:

  1. Pick a size.
  2. Download and install it (Ollama, Transformers, vLLM, or llama.cpp).
  3. Use it (chat, Python, HTTP API).
  4. Fine-tune it on your data (LoRA / QLoRA) and run the result.

«Download» and «install» mean different things for LLMs: you download weights (tens of GB) once, and you install a runtime (Ollama, vLLM, transformers, llama.cpp) that loads them.

Which Qwen to download

Ollama names (easiest):

Hugging Face names:

Rules of thumb:

Check the model card license before commercial use. Most Qwen weights use Apache 2.0, but confirm on the exact repository page (huggingface.co/Qwen, github.com/QwenLM/Qwen3).

Hardware requirements

Rough VRAM for full-precision (bf16) inference, context 4-8k:

Modelbf16 VRAM4-bit (GGUF / AWQ / GPTQ)
0.6B - 1.7B2-5 GBruns on CPU / laptop
4B - 8B10-20 GB6-10 GB VRAM
14B~30 GB10-12 GB VRAM
32B~65 GB20-24 GB VRAM

Fine-tuning needs more than inference. Full fine-tune of 7-8B needs ~60 GB VRAM; LoRA needs ~16-24 GB; QLoRA (4-bit base + LoRA adapters) fits 7-8B on one 16 GB card. CPU-only training is not practical beyond toy tests.

You need: Linux / macOS / WSL2 on Windows, Python 3.10+, NVIDIA drivers + CUDA for GPU runs, 30-80 GB free disk (weights + cache).

Fastest start: Ollama

Ollama downloads the weights and installs the runtime in one command. Best for «just chat now».

1. Install Ollama (Linux / macOS / WSL2)

curl -fsSL https://ollama.com/install.sh | sh

2. Download Qwen (pull once, reuse offline later)

ollama pull qwen3:8b

3. Chat

ollama run qwen3:8b

Smaller / bigger alternatives:

ollama pull qwen3:4b

ollama pull qwen3:14b

Ollama also serves an OpenAI-compatible HTTP API on port 11434:

List models:

curl http://localhost:11434/api/tags

Chat via API:

curl http://localhost:11434/api/chat -d '{"model": "qwen3:8b", "messages": [{"role": "user", "content": "Hello, Qwen!"}]}'

Upgrade weights later with ollama pull qwen3:8b again. Remove with ollama rm qwen3:8b.

Install via Hugging Face Transformers

Use this when you want Python-level control (scripts, training, embeddings).

1. Python env

python3 -m venv qwen-env

source qwen-env/bin/activate

pip install -U pip

pip install -U transformers accelerate torch --index-url https://download.pytorch.org/whl/cu121

Optional: flash attention for long context / speed

pip install -U flash-attn --no-build-isolation

2. Login once if you hit gated repos / rate limits

pip install -U huggingface_hub

hf auth login

3. Download is automatic on first use (cached in ~/.cache/huggingface). Or pre-fetch explicitly:

hf download Qwen/Qwen3-8B

from transformers import AutoModelForCausalLM, AutoTokenizer MODEL_ID = "Qwen/Qwen3-8B" # or "Qwen/Qwen2.5-7B-Instruct" tok = AutoTokenizer.from_pretrained(MODEL_ID) model = AutoModelForCausalLM.from_pretrained( MODEL_ID, dtype="auto", # bf16 on GPU, fp32 fallback on CPU device_map="auto", # spread across GPU/CPU automatically ) messages = [{"role": "user", "content": "Write a haiku about the Volga."}] text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tok([text], return_tensors="pt").to(model.device) out = model.generate(**inputs, max_new_tokens=256) # type: ignore[misc] print(tok.decode(out[0][len(inputs.input_ids[0]) :], skip_special_tokens=True))

Qwen3 may print a /think / /no_think switch in chat templates. Pass enable_thinking=False to apply_chat_template if you want direct answers without reasoning traces.

The # type: ignore[misc] on the generate line silences a known mypy false positive in transformers 5.x stubs (the declared from_pretrained return type does not match the generate self type); the call itself is correct at runtime.

Serve with vLLM (OpenAI-compatible API)

Use vLLM for high-throughput serving with an OpenAI-compatible endpoint. First run downloads weights from Hugging Face automatically.

pip install -U vllm

Serve Qwen3-8B on one GPU:

python -m vllm.entrypoints.openai.api_server --model Qwen/Qwen3-8B --dtype auto --max-model-len 8192 --port 8000

Test it:

curl http://localhost:8000/v1/models

curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model": "Qwen/Qwen3-8B", "messages": [{"role": "user", "content": "What is Qwen?"}]}'

Python client (same code works against real OpenAI by changing base_url):

from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") resp = client.chat.completions.create( model="Qwen/Qwen3-8B", messages=[{"role": "user", "content": "Explain LoRA in one paragraph."}], ) print(resp.choices[0].message.content)

Run on CPU with llama.cpp

For machines without a GPU, use a quantized GGUF build:

Ollama already does this for you; manual path:

1. Get a GGUF file, e.g. Qwen3-8B-Q4_K_M.gguf from a GGUF repo

hf download Qwen/Qwen3-8B-GGUF --include "*Q4_K_M*.gguf"

2a. llama.cpp server (OpenAI-compatible on :8080)

llama-server -m Qwen3-8B-Q4_K_M.gguf -c 4096 --port 8080

2b. or plain chat in terminal

llama-cli -m Qwen3-8B-Q4_K_M.gguf -c 4096 -n 512 -p "You are helpful. User: Hi! Assistant:"

How to use Qwen

Fine-tune Qwen (LoRA / QLoRA)

Do not full-fine-tune on one GPU - use LoRA (small trainable adapters) or QLoRA (4-bit frozen base + LoRA). Typical recipe for a chat / support / domain-QA style adaptation of Qwen3-8B or Qwen2.5-7B on a single 16-24 GB GPU:

1. Prepare data as chat JSONL (one object per line):

{"messages": [ {"role": "system", "content": "You are a bike-shop assistant."}, {"role": "user", "content": "Which brake pads fit model X?"}, {"role": "assistant", "content": "Resin pads, type B. ..."}]}

2. Install the training stack:

source qwen-env/bin/activate

pip install -U trl peft bitsandbytes datasets

3. Train (QLoRA, ~works for 7-8B on 16 GB VRAM):

import torch from datasets import load_dataset from peft import LoraConfig from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig from trl import SFTConfig, SFTTrainer MODEL_ID = "Qwen/Qwen3-8B" # or "Qwen/Qwen2.5-7B-Instruct" ds = load_dataset("json", data_files="train.jsonl", split="train") tok = AutoTokenizer.from_pretrained(MODEL_ID, use_fast=True) if tok.pad_token is None: tok.pad_token = tok.eos_token bnb = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, ) model = AutoModelForCausalLM.from_pretrained( MODEL_ID, quantization_config=bnb, device_map="auto", ) model.config.use_cache = False peft_cfg = LoraConfig( r=16, lora_alpha=32, lora_dropout=0.05, target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"], task_type="CAUSAL_LM", ) args = SFTConfig( output_dir="qwen-lora", num_train_epochs=2, per_device_train_batch_size=2, gradient_accumulation_steps=8, learning_rate=2e-4, logging_steps=10, save_steps=200, bf16=True, packing=True, max_length=2048, ) trainer = SFTTrainer( model=model, train_dataset=ds, processing_class=tok, peft_config=peft_cfg, args=args, ) trainer.train() trainer.save_model("qwen-lora") tok.save_pretrained("qwen-lora")

4. Practical rules:

Export and run your fine-tune

LoRA output is adapters only. Merge them back for serving:

import torch from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer BASE = "Qwen/Qwen3-8B" tok = AutoTokenizer.from_pretrained(BASE) model = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16, device_map="cpu") peft_model = PeftModel.from_pretrained(model, "qwen-lora") merged = peft_model.merge_and_unload() merged.save_pretrained("qwen-merged") tok.save_pretrained("qwen-merged")

Then either serve qwen-merged with vLLM / Transformers, or quantize to GGUF with llama.cpp (llama-quantize / convert_hf_to_gguf.py) and run it with llama-server or import into Ollama via a Modelfile (FROM ./qwen-merged.gguf + ollama create my-qwen -f Modelfile).

Troubleshooting

Official sources:

Article author: Andrei Olegovich

Related articles
AI
Qwen: download, install, fine-tune and use
OpenRouter: one API key for 300+ AI models (account, API, OpenCode, VSCode)

Search this site

Channel @aofeed Chat @aofeedchat

Contacts and cooperation:
I recommend our hosting beget.ru
Write to info@urn.su if you:
1. Want to write an article for our site or translate an article into your native language.
2. Want to place thematically relevant ads on the site.
3. Ads on my site pass maximum censorship. If you see an ad block unsuitable for school-age children, shocking or misleading - please contact us by e-mail
4. Found a mistake, inaccuracy, bug, etc. on the site. ... .......
5. Articles can be shared on social media by clicking a network icon: