Qwen3-VL-2B-Instruct-W4A16-AutoRound-AWQ

Model Overview

This is a 4-bit quantized version of the powerful Qwen/Qwen3-VL-2B-Instruct vision-language model.

It was optimized using Intel's AutoRound algorithm, which calibrates weights for 800 iterations to minimize quantization loss. This version retains the original FP16 vision tower, ensuring that visual capabilities (OCR, spatial reasoning, chart analysis) remain degradation-free.

Quantization Specifications

  • Method: AutoRound (Advanced Weight-Only Quantization)
  • Scheme: W4A16 (4-bit weights, 16-bit activations)
  • Symmetric: True
  • Group Size: 128
  • Vision Tower: Kept in FP16 (Unquantized for max accuracy)
  • Calibration: 512 samples, 800 iterations

Quickstart

1. Installation

To use this model in its native AutoRound format, you need the auto-round library.

pip install auto-round transformers torch

2. Inference Code

from transformers import AutoModelForCausalLM, AutoTokenizer, AutoProcessor
from auto_round import AutoRoundConfig

model_id = "Vishva007/Qwen3-VL-2B-Instruct-W4A16-AutoRound-AWQ"

# Load Model
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    trust_remote_code=True
)
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)

# Prepare Input
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg"},
            {"type": "text", "text": "Describe this image detailly."},
        ],
    }
]

text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)

inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
).to(model.device)

# Generate
generated_ids = model.generate(**inputs, max_new_tokens=128)
print(processor.batch_decode(generated_ids, skip_special_tokens=True))

🚀 Deploy on RunPod

One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.

🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.

PyTorch 2.13

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.13 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.13-runpod gmlupxnxfk Deploy to RunPod
PyTorch 2.13 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.13-runpod y3j8xvk4f4 Deploy to RunPod
PyTorch 2.13 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.13-runpod vigpissn5w Deploy to RunPod

PyTorch 2.12

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.12 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.12-runpod ctmz86zmf0 Deploy to RunPod
PyTorch 2.12 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.12-runpod qjko5yiwzi Deploy to RunPod
PyTorch 2.12 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.12-runpod ifg6xmye0f Deploy to RunPod

Citation

@article{cheng2023optimize,
  title={Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs},
  author={Cheng, Wenhua and Zhang, Weiwei and Shen, Haihao and Cai, Yiyang and He, Xin and Lv, Kaokao},
  journal={arXiv preprint arXiv:2309.05516},
  year={2023}
}
Downloads last month
28
Safetensors
Model size
2B params
Tensor type
I32
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vishva007/Qwen3-VL-2B-Instruct-W4A16-AutoRound-AWQ

Quantized
(91)
this model

Collection including Vishva007/Qwen3-VL-2B-Instruct-W4A16-AutoRound-AWQ

Paper for Vishva007/Qwen3-VL-2B-Instruct-W4A16-AutoRound-AWQ