Video-Text-to-Text
Transformers
Safetensors
English
Chinese
moss_vl
feature-extraction
Realtime
Streaming
Video-Understanding
Image-Understanding
MOSS-VL
OpenMOSS
multimodal
video
vision-language
custom_code
sglang
Instructions to use OpenMOSS-Team/MOSS-VL-Realtime-SGLANG with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OpenMOSS-Team/MOSS-VL-Realtime-SGLANG with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("OpenMOSS-Team/MOSS-VL-Realtime-SGLANG", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
docs: add bilingual model cards and streamline inference guidance
Browse files- README.md +49 -367
- README_zh.md +105 -0
- inference.md +256 -0
README.md
CHANGED
|
@@ -22,405 +22,87 @@ tags:
|
|
| 22 |
- sglang
|
| 23 |
---
|
| 24 |
|
| 25 |
-
<p align="center">
|
| 26 |
-
<img src="assets/logo.png" width="300" alt="MOSS-VL"/>
|
| 27 |
-
</p>
|
| 28 |
-
|
| 29 |
# MOSS-VL-Realtime-SGLANG
|
| 30 |
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
| Resource | Purpose |
|
| 34 |
-
| --- | --- |
|
| 35 |
-
| [OpenMOSS-Team/MOSS-VL-Realtime-SGLANG](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-SGLANG) | This Transformers 5.12.1-compatible checkpoint and custom code |
|
| 36 |
-
| [OpenMOSS-Team/MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) | Original checkpoint and Transformers 4.57-series reference code |
|
| 37 |
-
| [fnlp-vision/sglang-omni-realtime](https://github.com/fnlp-vision/sglang-omni-realtime) | Specialized realtime serving backend, developed on [SGLang-Omni](https://github.com/sgl-project/sglang-omni) |
|
| 38 |
-
| [fnlp-vision/MOSS-VL-Realtime_Demo](https://github.com/fnlp-vision/MOSS-VL-Realtime_Demo) | Separate Demo and gateway integration |
|
| 39 |
-
|
| 40 |
-
The custom Python files and configuration in this repository must be used together. Do not replace them with the original repository's 4.57-series files or assume that installing the official `sglang-omni` PyPI package includes the specialized backend.
|
| 41 |
-
|
| 42 |
-
## Compatibility and RoPE Fixes
|
| 43 |
-
|
| 44 |
-
The validated stack uses Transformers **5.12.1**, SGLang **0.5.16**, and PyTorch **2.11.0** on NVIDIA CUDA. The package keeps the 5.12.1 adaptations for configuration/RoPE APIs, output recording, chat-template return values, and attention-mask/cache interfaces.
|
| 45 |
-
|
| 46 |
-
Two targeted fixes are included:
|
| 47 |
-
|
| 48 |
-
- **Cross-attention Query RoPE:** every newly computed text query is rotated, including steps that reuse cached visual KV without new visual input. Visual keys are rotated only when newly computed; cached keys are not rotated again. The equivalent narrow fix is also published in the [original reference repository](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime/commit/1e6a45b292eeaf02aa733bd3aa7b6c85214ddc86).
|
| 49 |
-
- **Vision rotary frequencies under Transformers 5.12.1:** non-persistent frequency buffers may be rematerialized during model loading. This implementation reconstructs the canonical FP32 frequencies on the active device rather than trusting potentially invalid buffer contents. This is a 5.12.1 compatibility fix, not a claim that the original 4.57 environment has the same loading problem.
|
| 50 |
|
| 51 |
-
|
| 52 |
|
| 53 |
-
MOSS-VL
|
| 54 |
|
| 55 |
-
|
| 56 |
|
| 57 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 58 |
|
| 59 |
-
|
| 60 |
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
- Realtime streaming understanding: processes incoming frames continuously instead of waiting for a complete video.
|
| 68 |
-
- Interruptible interaction: users can ask questions at any timestamp in a running stream, and the model answers based on the frames observed so far.
|
| 69 |
-
- Proactive silence: the model can emit `<|silence|>` and continue observing when there is no meaningful visual update or the context is not sufficient.
|
| 70 |
-
- Dynamic correction: as new frames arrive, the model can revise earlier responses instead of being locked to an initial interpretation.
|
| 71 |
-
- Timestamp-aware frames: each streamed frame is associated with an absolute timestamp, helping the model reason about event order, duration, pacing, and fine-grained temporal localization.
|
| 72 |
-
- Unified MOSS-VL family: released together with MOSS-VL-Instruct and MOSS-VL-Base for offline use, continued pretraining, fine-tuning, and applied research.
|
| 73 |
-
|
| 74 |
-
## Model Design
|
| 75 |
-
|
| 76 |
-
### Architecture
|
| 77 |
-
|
| 78 |
-
MOSS-VL-Realtime adopts a cross-attention-based vision-language architecture that decouples visual encoding from language reasoning. This design is important for realtime usage because incoming visual content can be integrated into the running generation context without forcing the model into a strictly offline "load all frames, then answer" workflow.
|
| 79 |
|
| 80 |
-
|
| 81 |
-
<img src="assets/architecture.png" alt="MOSS-VL Architecture" width="100%"/>
|
| 82 |
-
</p>
|
| 83 |
|
| 84 |
-
|
| 85 |
|
| 86 |
-
|
| 87 |
|
| 88 |
-
MOSS-VL
|
| 89 |
|
| 90 |
-
|
| 91 |
|
| 92 |
| Item | Value |
|
| 93 |
| --- | --- |
|
| 94 |
| Parameters | 11B |
|
| 95 |
-
|
|
| 96 |
-
|
|
| 97 |
-
| Vision patch size | 16 |
|
| 98 |
-
|
|
| 99 |
-
|
|
| 100 |
-
|
|
| 101 |
-
| Realtime frame format | PIL-compatible image plus timestamp |
|
| 102 |
-
| Direct Transformers session scope | One active realtime session per model instance |
|
| 103 |
-
| SGLang-Omni session scope | Configurable with `--max-running-requests`; subject to KV capacity |
|
| 104 |
-
|
| 105 |
-
## Performance
|
| 106 |
-
|
| 107 |
-
MOSS-VL-Realtime is designed for streaming video understanding benchmarks where questions can arrive before a full video has been observed and correct answers may change as the scene evolves. It targets realtime interaction quality, proactive silence, and dynamic response updates in addition to standard video understanding accuracy.
|
| 108 |
-
|
| 109 |
-
<p align="center">
|
| 110 |
-
<img src="assets/benchmark-streaming.png" alt="MOSS-VL Streaming Benchmark" width="100%"/>
|
| 111 |
-
</p>
|
| 112 |
-
|
| 113 |
-
Detailed benchmark tables and comparisons for this release will be maintained in the MOSS-VL project resources.
|
| 114 |
-
|
| 115 |
-
## Quickstart
|
| 116 |
-
|
| 117 |
-
### Installation
|
| 118 |
-
|
| 119 |
-
For the specialized backend, install its source and pinned dependencies in a fresh, compatible CUDA environment. Install `uv` first if it is not already available:
|
| 120 |
-
|
| 121 |
-
```bash
|
| 122 |
-
git clone https://github.com/fnlp-vision/sglang-omni-realtime.git
|
| 123 |
-
cd sglang-omni-realtime
|
| 124 |
-
uv venv .venv -p 3.12
|
| 125 |
-
source .venv/bin/activate
|
| 126 |
-
uv pip install -e .
|
| 127 |
-
```
|
| 128 |
-
|
| 129 |
-
See the [backend README](https://github.com/fnlp-vision/sglang-omni-realtime#readme) for CUDA/toolchain prerequisites and deployment details. This is not a universal installation recipe for arbitrary CUDA versions. For direct Transformers inference, use the same compatible Transformers 5.12.1 environment and the example below; an SGLang server does not need to be running.
|
| 130 |
-
|
| 131 |
-
### Load the Model
|
| 132 |
-
|
| 133 |
-
```python
|
| 134 |
-
import torch
|
| 135 |
-
from transformers import AutoModelForCausalLM, AutoProcessor
|
| 136 |
-
|
| 137 |
-
checkpoint = "OpenMOSS-Team/MOSS-VL-Realtime-SGLANG"
|
| 138 |
-
|
| 139 |
-
processor = AutoProcessor.from_pretrained(
|
| 140 |
-
checkpoint,
|
| 141 |
-
trust_remote_code=True,
|
| 142 |
-
frame_extract_num_threads=1,
|
| 143 |
-
)
|
| 144 |
-
model = AutoModelForCausalLM.from_pretrained(
|
| 145 |
-
checkpoint,
|
| 146 |
-
trust_remote_code=True,
|
| 147 |
-
device_map="auto",
|
| 148 |
-
dtype=torch.bfloat16,
|
| 149 |
-
attn_implementation="eager",
|
| 150 |
-
)
|
| 151 |
-
model.eval()
|
| 152 |
-
```
|
| 153 |
-
|
| 154 |
-
This direct-Transformers example uses eager attention. The SGLang-Omni backend has its own attention configuration; do not confuse the two execution paths.
|
| 155 |
-
|
| 156 |
-
### Start the SGLang-Omni Backend
|
| 157 |
|
| 158 |
-
|
| 159 |
|
| 160 |
-
|
| 161 |
-
hf download OpenMOSS-Team/MOSS-VL-Realtime-SGLANG \
|
| 162 |
-
--local-dir /path/to/moss-vl-realtime-sglang
|
| 163 |
-
|
| 164 |
-
python examples/run_moss_vl_realtime_server.py \
|
| 165 |
-
--model-path /path/to/moss-vl-realtime-sglang \
|
| 166 |
-
--gpu 0 --host 127.0.0.1 --port 8000 \
|
| 167 |
-
--context-length 131072 \
|
| 168 |
-
--mem-fraction-static 0.60 \
|
| 169 |
-
--max-running-requests 1
|
| 170 |
-
```
|
| 171 |
-
|
| 172 |
-
Access requires authorization while the repository is private. The command uses a 128K context as an explicit deployment example; the model's 256K capacity and launcher defaults do not guarantee sufficient KV memory on every device. Tune context, memory fraction and concurrency for your hardware. The backend endpoint is `/v1/video/realtime`; its WebSocket protocol and TP options are documented in the [Realtime Cookbook](https://github.com/fnlp-vision/sglang-omni-realtime/blob/main/docs/cookbook/moss_vl_realtime.md).
|
| 173 |
-
|
| 174 |
-
This model repository contains weights and custom model/processor code, not the backend scheduler, Demo frontend, customer authentication or billing service.
|
| 175 |
-
|
| 176 |
-
## Inference Examples
|
| 177 |
-
|
| 178 |
-
### Online Inference
|
| 179 |
-
|
| 180 |
-
<details>
|
| 181 |
-
<summary><b>Session-style Online Inference</b></summary>
|
| 182 |
-
|
| 183 |
-
The recommended direct API is `create_realtime_session(...)`. A service or application owns the video capture pipeline, converts camera, screen, or video-file input into PIL-compatible frames, and pushes each frame with a non-decreasing timestamp.
|
| 184 |
-
|
| 185 |
-
Common session operations:
|
| 186 |
-
|
| 187 |
-
- `session.push_frame(image, timestamp=...)` appends one visual frame.
|
| 188 |
-
- `session.push_prompt("...")` appends a user question while the stream is running.
|
| 189 |
-
- `session.push_prompt_frame(prompt, image, timestamp=...)` aligns a prompt with a specific frame.
|
| 190 |
-
- `session.poll_output(...)` or `session.stream_outputs(...)` returns incremental text chunks.
|
| 191 |
-
|
| 192 |
-
`system_prompt` and `initial_prompt` are tokenized as the initial system/user turns before the first frame arrives. Subsequent user turns can be appended with `push_prompt(...)` while the same session continues observing frames.
|
| 193 |
-
|
| 194 |
-
For complete real-time inference usage, including local-video replay and service deployment, see [`realtime_inference`](https://github.com/OpenMOSS/MOSS-VL/tree/main/realtime_inference) in the MOSS-VL GitHub repository.
|
| 195 |
-
|
| 196 |
-
```python
|
| 197 |
-
import time
|
| 198 |
-
from PIL import Image
|
| 199 |
-
|
| 200 |
-
session = model.create_realtime_session(
|
| 201 |
-
processor,
|
| 202 |
-
initial_prompt=(
|
| 203 |
-
"As the video streams frame by frame, describe important changes as they happen. "
|
| 204 |
-
"Stay silent when there is no relevant update."
|
| 205 |
-
),
|
| 206 |
-
frame_queue_size=256,
|
| 207 |
-
max_tokens_per_turn=12,
|
| 208 |
-
max_new_tokens=4096,
|
| 209 |
-
do_sample=False,
|
| 210 |
-
)
|
| 211 |
-
|
| 212 |
-
frame_paths = [
|
| 213 |
-
"data/frame_0001.jpg",
|
| 214 |
-
"data/frame_0002.jpg",
|
| 215 |
-
"data/frame_0003.jpg",
|
| 216 |
-
]
|
| 217 |
-
|
| 218 |
-
try:
|
| 219 |
-
session.start()
|
| 220 |
-
|
| 221 |
-
for index, frame_path in enumerate(frame_paths):
|
| 222 |
-
image = Image.open(frame_path).convert("RGB")
|
| 223 |
-
session.push_frame(image, timestamp=index / 1.0)
|
| 224 |
-
|
| 225 |
-
while True:
|
| 226 |
-
chunk = session.poll_output(timeout=0.0)
|
| 227 |
-
if chunk is None:
|
| 228 |
-
break
|
| 229 |
-
print(chunk, end="", flush=True)
|
| 230 |
-
|
| 231 |
-
time.sleep(1.0)
|
| 232 |
-
|
| 233 |
-
session.push_prompt("What changed in the latest frames?")
|
| 234 |
-
|
| 235 |
-
# Realtime sessions stay alive waiting for future input, so use a bounded
|
| 236 |
-
# drain window and close the session explicitly when the producer is done.
|
| 237 |
-
drain_deadline = time.monotonic() + 5.0
|
| 238 |
-
while time.monotonic() < drain_deadline:
|
| 239 |
-
chunk = session.poll_output(timeout=0.1)
|
| 240 |
-
if chunk is not None:
|
| 241 |
-
print(chunk, end="", flush=True)
|
| 242 |
-
finally:
|
| 243 |
-
session.close()
|
| 244 |
-
```
|
| 245 |
|
| 246 |
-
|
| 247 |
-
|
| 248 |
-
</details>
|
| 249 |
-
|
| 250 |
-
<details>
|
| 251 |
-
<summary><b>Queue-style Online Inference</b></summary>
|
| 252 |
-
|
| 253 |
-
`online_generate(...)` is useful for backend systems that separate frame production and model inference through queues. It accepts dictionaries containing frames, prompts, events, reset controls, and stop controls.
|
| 254 |
-
|
| 255 |
-
```python
|
| 256 |
-
import queue
|
| 257 |
-
import threading
|
| 258 |
-
from PIL import Image
|
| 259 |
-
|
| 260 |
-
input_queue = queue.Queue()
|
| 261 |
-
output_queue = queue.Queue()
|
| 262 |
-
|
| 263 |
-
worker = threading.Thread(
|
| 264 |
-
target=model.online_generate,
|
| 265 |
-
args=(processor, input_queue, output_queue),
|
| 266 |
-
kwargs={
|
| 267 |
-
"frame_queue_size": 256,
|
| 268 |
-
"max_tokens_per_turn": 12,
|
| 269 |
-
"max_new_tokens": 4096,
|
| 270 |
-
"do_sample": False,
|
| 271 |
-
},
|
| 272 |
-
daemon=True,
|
| 273 |
-
)
|
| 274 |
-
worker.start()
|
| 275 |
-
|
| 276 |
-
input_queue.put({
|
| 277 |
-
"initial_prompt": "Answer only when the streamed video provides enough evidence.",
|
| 278 |
-
})
|
| 279 |
-
|
| 280 |
-
input_queue.put({"frame": Image.open("data/frame_0001.jpg").convert("RGB"), "timestamp": 0.0})
|
| 281 |
-
input_queue.put({"frame": Image.open("data/frame_0002.jpg").convert("RGB"), "timestamp": 1.0})
|
| 282 |
-
input_queue.put({"prompt": "What is happening now?"})
|
| 283 |
-
|
| 284 |
-
try:
|
| 285 |
-
while True:
|
| 286 |
-
chunk = output_queue.get(timeout=0.5)
|
| 287 |
-
print(chunk, end="", flush=True)
|
| 288 |
-
except queue.Empty:
|
| 289 |
-
pass
|
| 290 |
-
|
| 291 |
-
input_queue.put({"stop_online_generate": True})
|
| 292 |
-
worker.join()
|
| 293 |
-
```
|
| 294 |
|
| 295 |
-
|
| 296 |
-
|
| 297 |
-
</details>
|
| 298 |
-
|
| 299 |
-
### Offline Inference
|
| 300 |
-
|
| 301 |
-
MOSS-VL-Realtime also keeps the offline helper APIs for image and video prompts. For purely offline use, MOSS-VL-Instruct is usually the preferred checkpoint, but the realtime checkpoint can still process complete image and video inputs.
|
| 302 |
-
|
| 303 |
-
<details>
|
| 304 |
-
<summary><b>Single-video Offline Inference</b></summary>
|
| 305 |
-
|
| 306 |
-
```python
|
| 307 |
-
video_path = "data/example_video.mp4"
|
| 308 |
-
prompt = "Describe this video."
|
| 309 |
-
|
| 310 |
-
text = model.offline_video_generate(
|
| 311 |
-
processor,
|
| 312 |
-
prompt=prompt,
|
| 313 |
-
video=video_path,
|
| 314 |
-
shortest_edge=4096,
|
| 315 |
-
longest_edge=16777216,
|
| 316 |
-
video_max_pixels=201326592,
|
| 317 |
-
patch_size=16,
|
| 318 |
-
temporal_patch_size=1,
|
| 319 |
-
merge_size=2,
|
| 320 |
-
video_fps=1.0,
|
| 321 |
-
min_frames=1,
|
| 322 |
-
max_frames=256,
|
| 323 |
-
num_extract_threads=4,
|
| 324 |
-
image_mean=[0.5, 0.5, 0.5],
|
| 325 |
-
image_std=[0.5, 0.5, 0.5],
|
| 326 |
-
max_new_tokens=256,
|
| 327 |
-
temperature=1.0,
|
| 328 |
-
top_k=50,
|
| 329 |
-
top_p=1.0,
|
| 330 |
-
repetition_penalty=1.0,
|
| 331 |
-
do_sample=False,
|
| 332 |
-
vision_chunked_length=64,
|
| 333 |
-
)
|
| 334 |
-
|
| 335 |
-
print(text)
|
| 336 |
-
```
|
| 337 |
|
| 338 |
-
|
| 339 |
-
|
| 340 |
-
<details>
|
| 341 |
-
<summary><b>Batched Offline Inference</b></summary>
|
| 342 |
-
|
| 343 |
-
`offline_batch_generate` accepts independent image/video/text queries. Queries in the same batch should share the same `media_kwargs` and `generate_kwargs`.
|
| 344 |
-
|
| 345 |
-
```python
|
| 346 |
-
queries = [
|
| 347 |
-
{
|
| 348 |
-
"prompt": "Describe sample A.",
|
| 349 |
-
"images": [],
|
| 350 |
-
"videos": ["data/sample_a.mp4"],
|
| 351 |
-
"media_kwargs": {
|
| 352 |
-
"video_fps": 1.0,
|
| 353 |
-
"min_frames": 8,
|
| 354 |
-
"max_frames": 256,
|
| 355 |
-
},
|
| 356 |
-
"generate_kwargs": {
|
| 357 |
-
"temperature": 1.0,
|
| 358 |
-
"top_k": 50,
|
| 359 |
-
"top_p": 1.0,
|
| 360 |
-
"max_new_tokens": 256,
|
| 361 |
-
"repetition_penalty": 1.0,
|
| 362 |
-
"do_sample": False,
|
| 363 |
-
},
|
| 364 |
-
},
|
| 365 |
-
{
|
| 366 |
-
"prompt": "Describe sample B.",
|
| 367 |
-
"images": [],
|
| 368 |
-
"videos": ["data/sample_b.mp4"],
|
| 369 |
-
"media_kwargs": {
|
| 370 |
-
"video_fps": 1.0,
|
| 371 |
-
"min_frames": 8,
|
| 372 |
-
"max_frames": 256,
|
| 373 |
-
},
|
| 374 |
-
"generate_kwargs": {
|
| 375 |
-
"temperature": 1.0,
|
| 376 |
-
"top_k": 50,
|
| 377 |
-
"top_p": 1.0,
|
| 378 |
-
"max_new_tokens": 256,
|
| 379 |
-
"repetition_penalty": 1.0,
|
| 380 |
-
"do_sample": False,
|
| 381 |
-
},
|
| 382 |
-
},
|
| 383 |
-
]
|
| 384 |
-
|
| 385 |
-
with torch.no_grad():
|
| 386 |
-
result = model.offline_batch_generate(
|
| 387 |
-
processor,
|
| 388 |
-
queries,
|
| 389 |
-
vision_chunked_length=64,
|
| 390 |
-
)
|
| 391 |
-
|
| 392 |
-
texts = [item["text"] for item in result["results"]]
|
| 393 |
-
print(texts)
|
| 394 |
-
```
|
| 395 |
-
|
| 396 |
-
</details>
|
| 397 |
|
| 398 |
-
|
| 399 |
|
| 400 |
-
|
| 401 |
-
| --- | ---: | ---: | --- | --- |
|
| 402 |
-
| MOSS-VL-Realtime-SGLANG | 11B | 256K | Transformers 5.12.1 / specialized SGLang-Omni compatibility package | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-SGLANG |
|
| 403 |
-
| MOSS-VL-Realtime | 11B | 256K | Realtime streaming video interaction | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime |
|
| 404 |
-
| MOSS-VL-Instruct | 11B | 256K | Offline multimodal instruction following | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct |
|
| 405 |
-
| MOSS-VL-Base | 11B | 256K | Continued pretraining and fine-tuning | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Base |
|
| 406 |
-
| MOSS-VL-Instruct-0408 | 11B | 256K | Previous instruction-tuned checkpoint | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct-0408 |
|
| 407 |
-
| MOSS-VL-Base-0408 | 11B | 256K | Previous base checkpoint | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Base-0408 |
|
| 408 |
|
| 409 |
-
|
|
|
|
| 410 |
|
| 411 |
-
|
| 412 |
|
| 413 |
-
|
| 414 |
|
| 415 |
-
|
|
|
|
|
|
|
|
|
|
| 416 |
|
| 417 |
-
|
| 418 |
|
| 419 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 420 |
|
| 421 |
-
|
| 422 |
|
| 423 |
-
|
| 424 |
|
| 425 |
## Citation
|
| 426 |
|
|
|
|
| 22 |
- sglang
|
| 23 |
---
|
| 24 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
# MOSS-VL-Realtime-SGLANG
|
| 26 |
|
| 27 |
+
English | [简体中文](./README_zh.md)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
|
| 29 |
+
<p align="center"><img src="assets/logo.png" width="300" alt="MOSS-VL"/></p>
|
| 30 |
|
| 31 |
+
MOSS-VL-Realtime-SGLANG provides the MOSS-VL realtime checkpoint with Transformers 5.12.1-compatible model and processor code for the specialized SGLang-Omni backend. This is a compatibility package, not a retrained or quantized model.
|
| 32 |
|
| 33 |
+
## Features
|
| 34 |
|
| 35 |
+
- Continuous frame-by-frame video understanding and incremental text generation.
|
| 36 |
+
- Questions at any point in a stream, with response interruption.
|
| 37 |
+
- Proactive silence when no response is needed or evidence is insufficient.
|
| 38 |
+
- Updated responses as new visual information arrives.
|
| 39 |
+
- Timestamp-aware reasoning about event order and duration.
|
| 40 |
|
| 41 |
+
## Quick Start
|
| 42 |
|
| 43 |
+
| Goal | Entry |
|
| 44 |
+
| --- | --- |
|
| 45 |
+
| Browser video/voice interaction and optional memory | [MOSS-VL-Realtime Demo](https://github.com/fnlp-vision/MOSS-VL-Realtime_Demo#quick-start) |
|
| 46 |
+
| Standalone streaming inference service | [sglang-omni-realtime](https://github.com/fnlp-vision/sglang-omni-realtime#installation) |
|
| 47 |
+
| Direct Transformers inference | [Python examples](./inference.md) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 48 |
|
| 49 |
+
This repository contains weights, configuration, tokenizer, processor, and custom code. It does not include a running service. The model is public; download the complete repository rather than individual weight files.
|
|
|
|
|
|
|
| 50 |
|
| 51 |
+
ASR, TTS, and memory are provided by the Demo. The model's offline Python APIs do not mean that the Demo's realtime deployment enables offline chat.
|
| 52 |
|
| 53 |
+
## Architecture and Configuration
|
| 54 |
|
| 55 |
+
MOSS-VL separates visual encoding from language reasoning with cross-attention. Timestamped frames and Cross-attention Rotary Position Embedding (XRoPE) align text and visual patches across time, height, and width.
|
| 56 |
|
| 57 |
+
<p align="center"><img src="assets/architecture.png" alt="MOSS-VL Architecture" width="100%"/></p>
|
| 58 |
|
| 59 |
| Item | Value |
|
| 60 |
| --- | --- |
|
| 61 |
| Parameters | 11B |
|
| 62 |
+
| Weights | BF16 |
|
| 63 |
+
| Model context | 256K |
|
| 64 |
+
| Vision patch size / temporal patch size | 16 / 1 |
|
| 65 |
+
| Default video FPS / maximum sampled frames | 1.0 / 256 |
|
| 66 |
+
| Direct Transformers runtime | One active realtime session per model instance |
|
| 67 |
+
| SGLang-Omni runtime | Configurable session capacity, subject to GPU memory |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
|
| 69 |
+
The model context is not a per-session memory allocation guarantee. The recommended service configuration uses 131072 context; adjust context and concurrency for the available GPU memory.
|
| 70 |
|
| 71 |
+
## Performance
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
|
| 73 |
+
MOSS-VL-Realtime targets streaming understanding, proactive silence, and dynamic response updates. Benchmark results are summarized below; see the [technical report](https://arxiv.org/abs/2608.15045) for model research.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 74 |
|
| 75 |
+
<p align="center"><img src="assets/benchmark-streaming.png" alt="MOSS-VL Streaming Benchmark" width="100%"/></p>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 76 |
|
| 77 |
+
## Compatibility
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
+
Use this package's custom code and configuration together with the [specialized backend](https://github.com/fnlp-vision/sglang-omni-realtime). Installing a generic `sglang-omni` package does not supply the MOSS-VL integration.
|
| 80 |
|
| 81 |
+
The backend uses Transformers 5.12.1, SGLang 0.5.16, and PyTorch 2.11.0. This package adapts configuration, RoPE, and generation interfaces to Transformers 5.12.1 while keeping the original five BF16 weight shards, tokenizer, and vocabulary unchanged.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
|
| 83 |
+
- **Cross-attention Query RoPE:** rotate each newly computed text query, including when visual KV is reused; cached visual keys are not rotated again. The corresponding fix is also available in the [original reference implementation](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime/commit/1e6a45b292eeaf02aa733bd3aa7b6c85214ddc86).
|
| 84 |
+
- **Vision rotary frequencies:** reconstruct canonical FP32 frequencies on the active device to handle buffer rematerialization in Transformers 5.12.1. This does not imply the same loading issue exists in the original 4.57 environment.
|
| 85 |
|
| 86 |
+
The original [MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) uses separate Transformers 4.57-series code. Do not mix its custom files into this package. Select matching model and backend versions; CUDA compatibility does not imply NPU support.
|
| 87 |
|
| 88 |
+
## Limitations
|
| 89 |
|
| 90 |
+
- Response latency depends on hardware, frame rate, transport, and decoding settings.
|
| 91 |
+
- The direct Python runtime and the SGLang-Omni service have different session and backpressure behavior.
|
| 92 |
+
- Frame dropping and visual KV windows limit the visible history. Continuing beyond the context limit requires application-level memory handling.
|
| 93 |
+
- Applications should handle control tokens such as `<|silence|>`, `<|round_start|>`, and `<|round_end|>`.
|
| 94 |
|
| 95 |
+
## Related Models
|
| 96 |
|
| 97 |
+
| Model | Purpose |
|
| 98 |
+
| --- | --- |
|
| 99 |
+
| [MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) | Original realtime checkpoint and reference implementation |
|
| 100 |
+
| [MOSS-VL-Instruct](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct) | Offline multimodal instruction following |
|
| 101 |
+
| [MOSS-VL-Base](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Base) | Continued pretraining and fine-tuning |
|
| 102 |
|
| 103 |
+
## License
|
| 104 |
|
| 105 |
+
Apache-2.0. See [OpenMOSS/MOSS-VL](https://github.com/OpenMOSS/MOSS-VL) for the model family.
|
| 106 |
|
| 107 |
## Citation
|
| 108 |
|
README_zh.md
ADDED
|
@@ -0,0 +1,105 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# MOSS-VL-Realtime-SGLANG
|
| 2 |
+
|
| 3 |
+
[English](./README.md) | 简体中文
|
| 4 |
+
|
| 5 |
+
<p align="center"><img src="assets/logo.png" width="300" alt="MOSS-VL"/></p>
|
| 6 |
+
|
| 7 |
+
MOSS-VL-Realtime-SGLANG 提供 MOSS-VL 实时模型权重,以及面向特化 SGLang-Omni 后端的 Transformers 5.12.1 兼容模型与 processor 代码。这是兼容版本,不是重新训练或量化的模型。
|
| 8 |
+
|
| 9 |
+
## 功能
|
| 10 |
+
|
| 11 |
+
- 连续逐帧视频理解与增量文本生成。
|
| 12 |
+
- 在视频流任意时刻提问,并打断当前回答。
|
| 13 |
+
- 无需回答或证据不足时主动保持静默。
|
| 14 |
+
- 随新增视觉信息更新回答。
|
| 15 |
+
- 基于时间戳理解事件顺序和持续时间。
|
| 16 |
+
|
| 17 |
+
## 快速开始
|
| 18 |
+
|
| 19 |
+
| 目标 | 入口 |
|
| 20 |
+
| --- | --- |
|
| 21 |
+
| 浏览器视频/语音交互与可选 memory | [MOSS-VL-Realtime Demo](https://github.com/fnlp-vision/MOSS-VL-Realtime_Demo/blob/main/README_zh.md#快速开始) |
|
| 22 |
+
| 独立流式推理服务 | [sglang-omni-realtime](https://github.com/fnlp-vision/sglang-omni-realtime/blob/main/README_zh.md#安装) |
|
| 23 |
+
| 直接使用 Transformers 推理 | [Python 示例](./inference.md) |
|
| 24 |
+
|
| 25 |
+
本仓库包含权重、配置、tokenizer、processor 和自定义代码,不包含正在运行的服务。模型公开可访问,请下载完整仓库,而不是只下载个别权重文件。
|
| 26 |
+
|
| 27 |
+
ASR、TTS 和 memory 由 Demo 提供。模型保留离线 Python API,不代表 Demo 的实时部署已经启用离线聊天。
|
| 28 |
+
|
| 29 |
+
## 架构与配置
|
| 30 |
+
|
| 31 |
+
MOSS-VL 使用交叉注意力分离视觉编码与语言推理。带时间戳的帧及 Cross-attention Rotary Position Embedding(XRoPE)在时间、高度和宽度维度上对齐文本与视觉 patch。
|
| 32 |
+
|
| 33 |
+
<p align="center"><img src="assets/architecture.png" alt="MOSS-VL 架构" width="100%"/></p>
|
| 34 |
+
|
| 35 |
+
| 项目 | 值 |
|
| 36 |
+
| --- | --- |
|
| 37 |
+
| 参数量 | 11B |
|
| 38 |
+
| 权重类型 | BF16 |
|
| 39 |
+
| 模型上下文 | 256K |
|
| 40 |
+
| 视觉 patch 大小 / 时间 patch 大小 | 16 / 1 |
|
| 41 |
+
| 默认视频 FPS / 最大采样帧数 | 1.0 / 256 |
|
| 42 |
+
| 直接 Transformers 推理 | 每个模型实例一个活跃实时会话 |
|
| 43 |
+
| SGLang-Omni 推理 | 会话容量可配置,受 GPU 显存限制 |
|
| 44 |
+
|
| 45 |
+
模型上下文不等于部署时每路保证分配的容量。推荐服务配置使用 131072 context,请按可用显存调整上下文和并发数。
|
| 46 |
+
|
| 47 |
+
## 性能
|
| 48 |
+
|
| 49 |
+
MOSS-VL-Realtime 面向流式理解、主动静默及动态回答更新。下图汇总相关基准结果,模型研究详见[技术报告](https://arxiv.org/abs/2608.15045)。
|
| 50 |
+
|
| 51 |
+
<p align="center"><img src="assets/benchmark-streaming.png" alt="MOSS-VL 流式基准" width="100%"/></p>
|
| 52 |
+
|
| 53 |
+
## 兼容性
|
| 54 |
+
|
| 55 |
+
本仓库的自定义代码和配置应配套使用,并连接[特化后端](https://github.com/fnlp-vision/sglang-omni-realtime)。安装通用 `sglang-omni` 包不等于获得 MOSS-VL 的接入实现。
|
| 56 |
+
|
| 57 |
+
后端使用 Transformers 5.12.1、SGLang 0.5.16 和 PyTorch 2.11.0。本版本适配 Transformers 5.12.1 的配置、RoPE 和生成接口,原有五个 BF16 权重分片、tokenizer 和词表保持不变。
|
| 58 |
+
|
| 59 |
+
- **交叉注意力 Query RoPE:** 每次新计算的文本 query 都应用旋转,包括复用视觉 KV 的情况;已缓存的视觉 key 不重复旋转。对应修复也已应用于[原版参考实现](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime/commit/1e6a45b292eeaf02aa733bd3aa7b6c85214ddc86)。
|
| 60 |
+
- **视觉旋转频率:** 在当前设备重建标准 FP32 频率,以处理 Transformers 5.12.1 加载时重新创建 buffer 的情况,不代表原版 4.57 环境存在同样的加载问题。
|
| 61 |
+
|
| 62 |
+
原版 [MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) 使用独立的 Transformers 4.57 系列代码,不要将其自定义文件混入此版本。模型与后端应配套选择版本;CUDA 兼容不等于支持 NPU.
|
| 63 |
+
|
| 64 |
+
## 限制
|
| 65 |
+
|
| 66 |
+
- 响应延迟受硬件、帧率、传输和解码配置影响。
|
| 67 |
+
- 直接 Python 推理与 SGLang-Omni 服务具有不同的会话及背压行为。
|
| 68 |
+
- 丢帧和视觉 KV 窗口会限制可访问的视觉历史;跨上下文延续需要应用层记忆管理。
|
| 69 |
+
- 应用需处理 `<|silence|>`、`<|round_start|>`、`<|round_end|>` 等控制 token。
|
| 70 |
+
|
| 71 |
+
## 相关模型
|
| 72 |
+
|
| 73 |
+
| 模型 | 用途 |
|
| 74 |
+
| --- | --- |
|
| 75 |
+
| [MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) | 原版实时模型与参考实现 |
|
| 76 |
+
| [MOSS-VL-Instruct](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct) | 离线多模态指令理解 |
|
| 77 |
+
| [MOSS-VL-Base](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Base) | 继续预训练与微调 |
|
| 78 |
+
|
| 79 |
+
## 许可证
|
| 80 |
+
|
| 81 |
+
Apache-2.0。模型系列见 [OpenMOSS/MOSS-VL](https://github.com/OpenMOSS/MOSS-VL)。
|
| 82 |
+
|
| 83 |
+
## 引用
|
| 84 |
+
|
| 85 |
+
```bibtex
|
| 86 |
+
@misc{mossvl,
|
| 87 |
+
title = {MOSS-VL Technical Report},
|
| 88 |
+
author = {Wang, Pengyu and Tan, Chenkun and Zhou, Shaojun and Zhou, Qirui and Chen, Yanxin and He, Xingyang and Zeng, Huazheng and Cheng, Jijun and Wang, Chenghao and Qian, Xiaomeng and Wang, Pengfei and Huang, Zhan and Gao, Shanqing and Huang, Wei and Cao, Longjun and Ran, Wu and Liu, Jie and Zhu, Changtai and Wang, Hongkai and Tian, Yixian and Liu, Chenghao and Ye, Zhen and Wang, Xinghao and Jiang, Botian and Feng, Guoguo and Fei, Zhaoye and Li, Ruixiao and Chen, Mingshu and Gao, Yang and Cheng, Qinyuan and Li, Shimin and Qiu, Xipeng},
|
| 89 |
+
year = {2026},
|
| 90 |
+
eprint = {2608.15045},
|
| 91 |
+
archivePrefix = {arXiv},
|
| 92 |
+
primaryClass = {cs.CV},
|
| 93 |
+
url = {https://arxiv.org/abs/2608.15045}
|
| 94 |
+
}
|
| 95 |
+
|
| 96 |
+
@misc{mossvideopreview,
|
| 97 |
+
title = {{MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention}},
|
| 98 |
+
author = {Pengyu Wang and Chenkun Tan and Shaojun Zhou and Wei Huang and Qirui Zhou and Zhan Huang and Zhen Ye and Jijun Cheng and Xiaomeng Qian and Yanxin Chen and Xingyang He and Huazheng Zeng and Chenghao Wang and Pengfei Wang and Hongkai Wang and Shanqing Gao and Yixian Tian and Chenghao Liu and Xinghao Wang and Botian Jiang and Xipeng Qiu},
|
| 99 |
+
year = {2026},
|
| 100 |
+
eprint = {2606.07639},
|
| 101 |
+
archivePrefix = {arXiv},
|
| 102 |
+
primaryClass = {cs.CV},
|
| 103 |
+
url = {https://arxiv.org/abs/2606.07639}
|
| 104 |
+
}
|
| 105 |
+
```
|
inference.md
ADDED
|
@@ -0,0 +1,256 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Direct Transformers Inference
|
| 2 |
+
|
| 3 |
+
These examples use this repository's Transformers 5.12.1-compatible implementation, not the browser Demo or its WebSocket protocol. Prepare the [compatible backend environment](https://github.com/fnlp-vision/sglang-omni-realtime/blob/main/docs/get_started/installation.md); a running SGLang server is not required for direct Python inference. Do not use the Demo CPU environment to load this checkpoint.
|
| 4 |
+
|
| 5 |
+
Pin the model revision selected by the [compatibility manifest](https://github.com/fnlp-vision/MOSS-VL-Realtime_Demo/blob/main/deployment/repro/manifest.json). Sample paths below refer to media supplied by the caller.
|
| 6 |
+
|
| 7 |
+
## Load the Model
|
| 8 |
+
|
| 9 |
+
```python
|
| 10 |
+
import torch
|
| 11 |
+
from transformers import AutoModelForCausalLM, AutoProcessor
|
| 12 |
+
|
| 13 |
+
checkpoint = "OpenMOSS-Team/MOSS-VL-Realtime-SGLANG"
|
| 14 |
+
revision = "bcfd9ccf1e9db2896ad852301cc8dde4a6349c78"
|
| 15 |
+
|
| 16 |
+
processor = AutoProcessor.from_pretrained(
|
| 17 |
+
checkpoint,
|
| 18 |
+
revision=revision,
|
| 19 |
+
trust_remote_code=True,
|
| 20 |
+
frame_extract_num_threads=1,
|
| 21 |
+
)
|
| 22 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 23 |
+
checkpoint,
|
| 24 |
+
revision=revision,
|
| 25 |
+
trust_remote_code=True,
|
| 26 |
+
device_map="auto",
|
| 27 |
+
dtype=torch.bfloat16,
|
| 28 |
+
attn_implementation="eager",
|
| 29 |
+
)
|
| 30 |
+
model.eval()
|
| 31 |
+
```
|
| 32 |
+
|
| 33 |
+
This direct-Transformers example uses eager attention. The SGLang-Omni backend has its own attention configuration; do not confuse the two execution paths.
|
| 34 |
+
|
| 35 |
+
|
| 36 |
+
## Inference Examples
|
| 37 |
+
|
| 38 |
+
### Online Inference
|
| 39 |
+
|
| 40 |
+
<details>
|
| 41 |
+
<summary><b>Session-style Online Inference</b></summary>
|
| 42 |
+
|
| 43 |
+
The recommended direct API is `create_realtime_session(...)`. A service or application owns the video capture pipeline, converts camera, screen, or video-file input into PIL-compatible frames, and pushes each frame with a non-decreasing timestamp.
|
| 44 |
+
|
| 45 |
+
Common session operations:
|
| 46 |
+
|
| 47 |
+
- `session.push_frame(image, timestamp=...)` appends one visual frame.
|
| 48 |
+
- `session.push_prompt("...")` appends a user question while the stream is running.
|
| 49 |
+
- `session.push_prompt_frame(prompt, image, timestamp=...)` aligns a prompt with a specific frame.
|
| 50 |
+
- `session.poll_output(...)` or `session.stream_outputs(...)` returns incremental text chunks.
|
| 51 |
+
|
| 52 |
+
`system_prompt` and `initial_prompt` are tokenized as the initial system/user turns before the first frame arrives. Subsequent user turns can be appended with `push_prompt(...)` while the same session continues observing frames.
|
| 53 |
+
|
| 54 |
+
For complete real-time inference usage, including local-video replay and service deployment, see [`realtime_inference`](https://github.com/OpenMOSS/MOSS-VL/tree/main/realtime_inference) in the MOSS-VL GitHub repository.
|
| 55 |
+
|
| 56 |
+
```python
|
| 57 |
+
import time
|
| 58 |
+
from PIL import Image
|
| 59 |
+
|
| 60 |
+
session = model.create_realtime_session(
|
| 61 |
+
processor,
|
| 62 |
+
initial_prompt=(
|
| 63 |
+
"As the video streams frame by frame, describe important changes as they happen. "
|
| 64 |
+
"Stay silent when there is no relevant update."
|
| 65 |
+
),
|
| 66 |
+
frame_queue_size=256,
|
| 67 |
+
max_tokens_per_turn=12,
|
| 68 |
+
max_new_tokens=4096,
|
| 69 |
+
do_sample=False,
|
| 70 |
+
)
|
| 71 |
+
|
| 72 |
+
frame_paths = [
|
| 73 |
+
"data/frame_0001.jpg",
|
| 74 |
+
"data/frame_0002.jpg",
|
| 75 |
+
"data/frame_0003.jpg",
|
| 76 |
+
]
|
| 77 |
+
|
| 78 |
+
try:
|
| 79 |
+
session.start()
|
| 80 |
+
|
| 81 |
+
for index, frame_path in enumerate(frame_paths):
|
| 82 |
+
image = Image.open(frame_path).convert("RGB")
|
| 83 |
+
session.push_frame(image, timestamp=index / 1.0)
|
| 84 |
+
|
| 85 |
+
while True:
|
| 86 |
+
chunk = session.poll_output(timeout=0.0)
|
| 87 |
+
if chunk is None:
|
| 88 |
+
break
|
| 89 |
+
print(chunk, end="", flush=True)
|
| 90 |
+
|
| 91 |
+
time.sleep(1.0)
|
| 92 |
+
|
| 93 |
+
session.push_prompt("What changed in the latest frames?")
|
| 94 |
+
|
| 95 |
+
# Realtime sessions stay alive waiting for future input, so use a bounded
|
| 96 |
+
# drain window and close the session explicitly when the producer is done.
|
| 97 |
+
drain_deadline = time.monotonic() + 5.0
|
| 98 |
+
while time.monotonic() < drain_deadline:
|
| 99 |
+
chunk = session.poll_output(timeout=0.1)
|
| 100 |
+
if chunk is not None:
|
| 101 |
+
print(chunk, end="", flush=True)
|
| 102 |
+
finally:
|
| 103 |
+
session.close()
|
| 104 |
+
```
|
| 105 |
+
|
| 106 |
+
Frame timestamps are measured in seconds and must be non-decreasing within a session. The input producer can be a camera, screen capture, decoded video file, browser frame sampler, or any other source that yields images with timestamps.
|
| 107 |
+
|
| 108 |
+
</details>
|
| 109 |
+
|
| 110 |
+
<details>
|
| 111 |
+
<summary><b>Queue-style Online Inference</b></summary>
|
| 112 |
+
|
| 113 |
+
`online_generate(...)` is useful for backend systems that separate frame production and model inference through queues. It accepts dictionaries containing frames, prompts, events, reset controls, and stop controls.
|
| 114 |
+
|
| 115 |
+
```python
|
| 116 |
+
import queue
|
| 117 |
+
import threading
|
| 118 |
+
from PIL import Image
|
| 119 |
+
|
| 120 |
+
input_queue = queue.Queue()
|
| 121 |
+
output_queue = queue.Queue()
|
| 122 |
+
|
| 123 |
+
worker = threading.Thread(
|
| 124 |
+
target=model.online_generate,
|
| 125 |
+
args=(processor, input_queue, output_queue),
|
| 126 |
+
kwargs={
|
| 127 |
+
"frame_queue_size": 256,
|
| 128 |
+
"max_tokens_per_turn": 12,
|
| 129 |
+
"max_new_tokens": 4096,
|
| 130 |
+
"do_sample": False,
|
| 131 |
+
},
|
| 132 |
+
daemon=True,
|
| 133 |
+
)
|
| 134 |
+
worker.start()
|
| 135 |
+
|
| 136 |
+
input_queue.put({
|
| 137 |
+
"initial_prompt": "Answer only when the streamed video provides enough evidence.",
|
| 138 |
+
})
|
| 139 |
+
|
| 140 |
+
input_queue.put({"frame": Image.open("data/frame_0001.jpg").convert("RGB"), "timestamp": 0.0})
|
| 141 |
+
input_queue.put({"frame": Image.open("data/frame_0002.jpg").convert("RGB"), "timestamp": 1.0})
|
| 142 |
+
input_queue.put({"prompt": "What is happening now?"})
|
| 143 |
+
|
| 144 |
+
try:
|
| 145 |
+
while True:
|
| 146 |
+
chunk = output_queue.get(timeout=0.5)
|
| 147 |
+
print(chunk, end="", flush=True)
|
| 148 |
+
except queue.Empty:
|
| 149 |
+
pass
|
| 150 |
+
|
| 151 |
+
input_queue.put({"stop_online_generate": True})
|
| 152 |
+
worker.join()
|
| 153 |
+
```
|
| 154 |
+
|
| 155 |
+
Each queue item can contain `frame` or `image`, `timestamp`, `prompt`, `frames`, `event`, `events`, `initial_prompt`, `system_prompt`, `generate_kwargs`, `reset_session`, or stop controls such as `stop_online_generate`.
|
| 156 |
+
|
| 157 |
+
</details>
|
| 158 |
+
|
| 159 |
+
### Offline Inference
|
| 160 |
+
|
| 161 |
+
MOSS-VL-Realtime also keeps the offline helper APIs for image and video prompts. For purely offline use, MOSS-VL-Instruct is usually the preferred checkpoint, but the realtime checkpoint can still process complete image and video inputs.
|
| 162 |
+
|
| 163 |
+
<details>
|
| 164 |
+
<summary><b>Single-video Offline Inference</b></summary>
|
| 165 |
+
|
| 166 |
+
```python
|
| 167 |
+
video_path = "data/example_video.mp4"
|
| 168 |
+
prompt = "Describe this video."
|
| 169 |
+
|
| 170 |
+
text = model.offline_video_generate(
|
| 171 |
+
processor,
|
| 172 |
+
prompt=prompt,
|
| 173 |
+
video=video_path,
|
| 174 |
+
shortest_edge=4096,
|
| 175 |
+
longest_edge=16777216,
|
| 176 |
+
video_max_pixels=201326592,
|
| 177 |
+
patch_size=16,
|
| 178 |
+
temporal_patch_size=1,
|
| 179 |
+
merge_size=2,
|
| 180 |
+
video_fps=1.0,
|
| 181 |
+
min_frames=1,
|
| 182 |
+
max_frames=256,
|
| 183 |
+
num_extract_threads=4,
|
| 184 |
+
image_mean=[0.5, 0.5, 0.5],
|
| 185 |
+
image_std=[0.5, 0.5, 0.5],
|
| 186 |
+
max_new_tokens=256,
|
| 187 |
+
temperature=1.0,
|
| 188 |
+
top_k=50,
|
| 189 |
+
top_p=1.0,
|
| 190 |
+
repetition_penalty=1.0,
|
| 191 |
+
do_sample=False,
|
| 192 |
+
vision_chunked_length=64,
|
| 193 |
+
)
|
| 194 |
+
|
| 195 |
+
print(text)
|
| 196 |
+
```
|
| 197 |
+
|
| 198 |
+
</details>
|
| 199 |
+
|
| 200 |
+
<details>
|
| 201 |
+
<summary><b>Batched Offline Inference</b></summary>
|
| 202 |
+
|
| 203 |
+
`offline_batch_generate` accepts independent image/video/text queries. Queries in the same batch should share the same `media_kwargs` and `generate_kwargs`.
|
| 204 |
+
|
| 205 |
+
```python
|
| 206 |
+
queries = [
|
| 207 |
+
{
|
| 208 |
+
"prompt": "Describe sample A.",
|
| 209 |
+
"images": [],
|
| 210 |
+
"videos": ["data/sample_a.mp4"],
|
| 211 |
+
"media_kwargs": {
|
| 212 |
+
"video_fps": 1.0,
|
| 213 |
+
"min_frames": 8,
|
| 214 |
+
"max_frames": 256,
|
| 215 |
+
},
|
| 216 |
+
"generate_kwargs": {
|
| 217 |
+
"temperature": 1.0,
|
| 218 |
+
"top_k": 50,
|
| 219 |
+
"top_p": 1.0,
|
| 220 |
+
"max_new_tokens": 256,
|
| 221 |
+
"repetition_penalty": 1.0,
|
| 222 |
+
"do_sample": False,
|
| 223 |
+
},
|
| 224 |
+
},
|
| 225 |
+
{
|
| 226 |
+
"prompt": "Describe sample B.",
|
| 227 |
+
"images": [],
|
| 228 |
+
"videos": ["data/sample_b.mp4"],
|
| 229 |
+
"media_kwargs": {
|
| 230 |
+
"video_fps": 1.0,
|
| 231 |
+
"min_frames": 8,
|
| 232 |
+
"max_frames": 256,
|
| 233 |
+
},
|
| 234 |
+
"generate_kwargs": {
|
| 235 |
+
"temperature": 1.0,
|
| 236 |
+
"top_k": 50,
|
| 237 |
+
"top_p": 1.0,
|
| 238 |
+
"max_new_tokens": 256,
|
| 239 |
+
"repetition_penalty": 1.0,
|
| 240 |
+
"do_sample": False,
|
| 241 |
+
},
|
| 242 |
+
},
|
| 243 |
+
]
|
| 244 |
+
|
| 245 |
+
with torch.no_grad():
|
| 246 |
+
result = model.offline_batch_generate(
|
| 247 |
+
processor,
|
| 248 |
+
queries,
|
| 249 |
+
vision_chunked_length=64,
|
| 250 |
+
)
|
| 251 |
+
|
| 252 |
+
texts = [item["text"] for item in result["results"]]
|
| 253 |
+
print(texts)
|
| 254 |
+
```
|
| 255 |
+
|
| 256 |
+
</details>
|