Instructions to use Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix") model = AutoModelForMultimodalLM.from_pretrained("Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix
- SGLang
How to use Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix with Docker Model Runner:
docker model run hf.co/Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix
Quantization Dataset & How-to
Hello, thank you for your work. I am perceiving your quantized model way more robust than the one from Qwen directly in some specific tasks (e.g. RAG). Could you give an info on how you quantized the model? So which dataset you used and how you processed it (parameters, maybe script if possible)? Best :-)
Are you asking about Benasd/Qwen2.5-VL-72B-Instruct-AWQ(without trailing -fix) instead of Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix?
Benasd/Qwen2.5-VL-72B-Instruct-AWQ-fix uses their official AWQ weights, so the model is exactly the same. This is a temporary fix for their official preprocessor_config.json which is not needed any more, because they have fixed the official model.
Thank you, got it. And Benasd/Qwen2.5-VL-72B-Instruct-AWQ? Because in their official AWQ weights I observed some character issues when generating e.g. German text. The Benasd/Qwen2.5-VL-72B-Instruct-AWQ model worked better on instruction following too. That is why I was curious :-)
Sorry for the late reply.
I used the simplest quantization script for Benasd/Qwen2.5-VL-72B-Instruct-AWQ, which relies on the default mit-han-lab/pile-val-backup as its calibration dataset (a text-only dataset).
Below is the script I used:
from AutoAWQ.awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model_path = "Qwen/Qwen2.5-VL-72B-Instruct"
quant_path = "Qwen2.5-VL-72B-Instruct-AWQ"
quant_config = { "zero_point": True, "q_group_size": 64, "w_bit": 4, "version": "GEMM" }
# Load model
model = AutoAWQForCausalLM.from_pretrained(model_path)
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
# Quantize
model.quantize(tokenizer, quant_config=quant_config)
# Save quantized model
model.save_quantized(quant_path)
tokenizer.save_pretrained(quant_path)
print(f'Model is quantized and saved at "{quant_path}"')
I think the official AWQ model uses the script recommended by this contributor, which is supposed to be better than using mit-han-lab/pile-val-backup, since it uses a vision-and-text dataset as the calibration set. However, in both my experiment and yours, it actually performs worse — and I don't have a theory as to why.
Here is the script that uses the image + text dataset (make sure to set Qwen2_5_VLProcessor.from_pretrained(model_path, padding_side='left')): https://github.com/casper-hansen/AutoAWQ/blob/main/docs/examples.md#custom-quantizer-qwen2-vl-example