joerowell commited on
Commit
0bab366
·
verified ·
1 Parent(s): 0f4935b

Document SGLang support (sgl-project/sglang#24204)

Browse files
Files changed (1) hide show
  1. README.md +5 -1
README.md CHANGED
@@ -124,7 +124,7 @@ Submit feedback with `/feedback` and read the [full documentation on GitHub](htt
124
 
125
  ### Local deployment
126
 
127
- Laguna XS.2 is supported in vLLM and Transformers, and TRT-LLM thanks to the support of the team at NVIDIA. Use Laguna-XS.2 with Ollama (with MLX support) and the mlx-lm framework for the best experience on your local machine.
128
 
129
  #### vLLM
130
 
@@ -147,6 +147,10 @@ vllm serve \
147
 
148
  See the [vLLM recipes page](https://recipes.vllm.ai/poolside/Laguna-XS.2) for additional deployment guidance.
149
 
 
 
 
 
150
  #### Speculative decoding (DFlash)
151
 
152
  For lower latency, serve Laguna XS.2 with the [Laguna-XS.2 DFlash speculator](https://huggingface.co/poolside/Laguna-XS.2-speculator.dflash) — a 5-layer Llama-style draft model that proposes up to 7 tokens per step at ~70% per-position acceptance on coding tasks.
 
124
 
125
  ### Local deployment
126
 
127
+ Laguna XS.2 is supported in vLLM, SGLang, and Transformers, and TRT-LLM thanks to the support of the team at NVIDIA. Use Laguna-XS.2 with Ollama (with MLX support) and the mlx-lm framework for the best experience on your local machine.
128
 
129
  #### vLLM
130
 
 
147
 
148
  See the [vLLM recipes page](https://recipes.vllm.ai/poolside/Laguna-XS.2) for additional deployment guidance.
149
 
150
+ #### SGLang
151
+
152
+ Laguna XS.2 is supported in SGLang via [sgl-project/sglang#24204](https://github.com/sgl-project/sglang/pull/24204). See the [SGLang cookbook entry](https://github.com/sgl-project/sglang/pull/24730) for a serving recipe.
153
+
154
  #### Speculative decoding (DFlash)
155
 
156
  For lower latency, serve Laguna XS.2 with the [Laguna-XS.2 DFlash speculator](https://huggingface.co/poolside/Laguna-XS.2-speculator.dflash) — a 5-layer Llama-style draft model that proposes up to 7 tokens per step at ~70% per-position acceptance on coding tasks.