Instructions to use LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct
- SGLang
How to use LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct with Docker Model Runner:
docker model run hf.co/LGAI-EXAONE/EXAONE-3.5-2.4B-Instruct
Model upload incomplete?
Noticed that when converting to GGUF we end up with a missing output_weight, only for the 2.4B model
Also noticed that the other two models have a final safetensors file containing lm_head.weight where this one doesn't, so wondering if it somehow got missed. Your GGUF files seem to have the proper output_weight so must have been made with the proper original file
"tie_word_embeddings" is true! The output_weight is just the embedding matrix....
Strange then that the produced GGUF I get is different from the one they posted π€
Hello, thank you for your question.
We basically use tie_word_embeddings to be True on 2.4B models, for memory efficiency on inference stage.
However, we found that conversion steps for AWQ and GGUF do not fit to the model with tie_word_embeddings, so we need to convert "Tied" model into "Un-tied" model before quantization.
When you try making another GGUF model on yourself, please copy the model.transformer.wte to model.lm_head and save it for quantization/conversion.
I'll look into how to do that later today and get back to you
Do you have any reference material on how to do this? From quick searching online I can't find any kind of reference to converting a "tied" model to an "un-tied" model..
Load w/ huggingface transformers, and copy the tensor there?
It's a bit unclear to me too. But, I made this tiny script and it worked, make sure to replace the directories:
import torch
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained('/model/exaone_folder/')
model.lm_head.weight = torch.nn.Parameter(model.transformer.wte.weight.clone().detach())
print(model.lm_head.weight.requires_grad)
model.tie_word_embeddings = False
model.save_pretrained('/model/exaone_folder/')