Instructions to use nvidia/Llama-3_1-Nemotron-51B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nvidia/Llama-3_1-Nemotron-51B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nvidia/Llama-3_1-Nemotron-51B-Instruct", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("nvidia/Llama-3_1-Nemotron-51B-Instruct", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nvidia/Llama-3_1-Nemotron-51B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nvidia/Llama-3_1-Nemotron-51B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama-3_1-Nemotron-51B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nvidia/Llama-3_1-Nemotron-51B-Instruct
- SGLang
How to use nvidia/Llama-3_1-Nemotron-51B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nvidia/Llama-3_1-Nemotron-51B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama-3_1-Nemotron-51B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nvidia/Llama-3_1-Nemotron-51B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nvidia/Llama-3_1-Nemotron-51B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use nvidia/Llama-3_1-Nemotron-51B-Instruct with Docker Model Runner:
docker model run hf.co/nvidia/Llama-3_1-Nemotron-51B-Instruct
AttributeError: 'DeciLMConfig'
Looks like something is missing in this repo. This is what running it with VLLM does:
AttributeError: 'DeciLMConfig' object has no attribute 'num_key_value_heads_per_layer'
And with lmsys’ model_worker from fastchat, the output is jibberish:
NoQuestion and dried#def035def0Question the but which said WELL)def1TitleQuestion from.........NETQuestion l cellsRQuestion that she pursuedQuestion Question from if0def0Question Question in def0Question or is we.from event's selectsdef1#ForQuestionQuestion and he.Question he,Question we,Question but he [def3Question and securing we,##import The waiting
Question atGenQuestion (Question as she.##Question1#InternationalQuestion Question Question and if0Question anddef0Question ppQuestion and ScottishQuestion1def2Question althoughQuestion but and question and leadingQuestion Question and which she girl pageQuestion and |#Question as sincedef0Question andMichaeldef0
#Question createQuestion and Evedef10Adef1-Question 196def 'importThis and have constantQuestion #Question This of including even as to they:Question import......MicQuestion or well and people andQuestion it RegionalQuestion the CaliforniaQuestion and experienced thedef0Question000Question or picks3As COMQuestion such.
#def0def01Question000Question Question and employingQuestion OpticalQuestion where Question from especially between which In the PoQuestion onQuestion and WordQuestion seQuestion and ##Question and Lisadef0Question particularly I:We.Question 199Question which says tools and pressingQuestion 201import6Question where such in being and IndependentQuestion MaryQuestion onQuestion andQuestion that This 201#from{import :12#Question and SecAimport1def2Question Question import #Question and trusted smaller.##Question I.Question the VisashiQuestion which says FL#Question are thedef0def1Question and ‘I.#Question importbdef15Question000#Question 11Question import1Question but
Question and Serve##One2def0def1def0Question that he
importdef0def0Question 201Question theQuestion because1#Question0A#FQuestion but I
Question how my largest storyTh#def0Question Question said QuestionQuestion and she,def0As cub##UnitedQuestion &Question and perseverance They 13I:def0Question letQuestion anddef0def03Question andQuestion Question2Question from3I.5def0Question here.43#def1Question for picksAQuestion and choices andThedef0#Question andQuestion where but and she,Thedef100Question they:def30Question and.WaterQuestion def0F#It,Question 202def0Question 201Question as if2Question then skilled SchoolQuestion given and combines she:Should1RWe-This 198def3TheQuestion which Most QuestionQuestion one then thedef0def0TheQuestion but to stdoutQuestion anddef0def1Ldef1
def0def0def35-Question and silly InstituteQuestion importdef0Question such |__VerQuestion which said とIn idef0
Question there,Question so important from '{0def0def0Question def0Question and QuestionQuestion and we (Question and question and027Question of orOneQuestion and local hádef1Question StateQuestion as how She &#def0Question the 201def0Question Question and choices with major
Question who:def0Question and NOThe#def2def0Question Question but but CQuestion to we will the ‘Question and explained000Question Question for but picksQuestion that she
I,Question then
import2F#def1Question including participantsdef1def0def0def0Question 11Question and feminist.,Question you:def1Question if3Mdef18Question and cáQuestion #ByQuestion to we,Question which stateQuestion and fitness is selectsdef0
This ofdef1#Question a
importifQuestion for while backgroundQuestion and
Question which
#Question and.Question or DevQuestionAnTheQuestion in the shridef03Question theQuestion Texasdef2Question are the UnitedQuestion and interestQuestion 1##"def2def0Question by they
NQuestion of but challenged is
We/Question and questionnaire and will said while9Question Question or MQuestion atWe,Question that for#!/201This in StandardQuestion and selectsQuestion from LisaQuestion and electronicch#sQuestion importIt.def0
Question but and aboutQuestion in DeskQuestion delivering because5def0Question we
#Does he,def1#Question it,Question but ydef0def2Question anddef0Question a we
#Question (How to he
Question its if1
def0def0This/importTagdef0def0Question and said1Question aboutMicrosoftdef02#Analysis#def0Question nimportLQuestion 201Question HomeQuestion MarilynQuestion to polyQuestion most But _Question and EuropeanQuestion W#def0Question def
and trust remote code doesn't fix it.
help appreciated
yup endup with above error and any work arounds please
yep same issue. DeciLMConfig' object has no attribute 'num_key_value_heads_per_layer'