CCCCyx commited on
Commit
ebf9fc5
·
verified ·
1 Parent(s): bcfd9cc

docs: add bilingual model cards and streamline inference guidance

Browse files
Files changed (3) hide show
  1. README.md +49 -367
  2. README_zh.md +105 -0
  3. inference.md +256 -0
README.md CHANGED
@@ -22,405 +22,87 @@ tags:
22
  - sglang
23
  ---
24
 
25
- <p align="center">
26
- <img src="assets/logo.png" width="300" alt="MOSS-VL"/>
27
- </p>
28
-
29
  # MOSS-VL-Realtime-SGLANG
30
 
31
- This repository packages **MOSS-VL-Realtime for Transformers 5.12.1 and the MOSS-VL SGLang-Omni realtime backend**. It is a compatibility release of the original streaming checkpoint, not a newly trained or quantized model.
32
-
33
- | Resource | Purpose |
34
- | --- | --- |
35
- | [OpenMOSS-Team/MOSS-VL-Realtime-SGLANG](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-SGLANG) | This Transformers 5.12.1-compatible checkpoint and custom code |
36
- | [OpenMOSS-Team/MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) | Original checkpoint and Transformers 4.57-series reference code |
37
- | [fnlp-vision/sglang-omni-realtime](https://github.com/fnlp-vision/sglang-omni-realtime) | Specialized realtime serving backend, developed on [SGLang-Omni](https://github.com/sgl-project/sglang-omni) |
38
- | [fnlp-vision/MOSS-VL-Realtime_Demo](https://github.com/fnlp-vision/MOSS-VL-Realtime_Demo) | Separate Demo and gateway integration |
39
-
40
- The custom Python files and configuration in this repository must be used together. Do not replace them with the original repository's 4.57-series files or assume that installing the official `sglang-omni` PyPI package includes the specialized backend.
41
-
42
- ## Compatibility and RoPE Fixes
43
-
44
- The validated stack uses Transformers **5.12.1**, SGLang **0.5.16**, and PyTorch **2.11.0** on NVIDIA CUDA. The package keeps the 5.12.1 adaptations for configuration/RoPE APIs, output recording, chat-template return values, and attention-mask/cache interfaces.
45
-
46
- Two targeted fixes are included:
47
-
48
- - **Cross-attention Query RoPE:** every newly computed text query is rotated, including steps that reuse cached visual KV without new visual input. Visual keys are rotated only when newly computed; cached keys are not rotated again. The equivalent narrow fix is also published in the [original reference repository](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime/commit/1e6a45b292eeaf02aa733bd3aa7b6c85214ddc86).
49
- - **Vision rotary frequencies under Transformers 5.12.1:** non-persistent frequency buffers may be rematerialized during model loading. This implementation reconstructs the canonical FP32 frequencies on the active device rather than trusting potentially invalid buffer contents. This is a 5.12.1 compatibility fix, not a claim that the original 4.57 environment has the same loading problem.
50
 
51
- The five BF16 weight shards, tokenizer and vocabulary are unchanged from the original checkpoint. The compatibility changes are in code and configuration, not model training. Pin the model revision and backend commit together when reproducing results. CUDA validation does not imply NPU support, nor should the latest 4.57-series source changes be assumed to be present in this separate compatibility branch.
52
 
53
- MOSS-VL is an open vision-language model family from OpenMOSS, supporting image understanding, long-video understanding, and realtime streaming interaction.
54
 
55
- **Technical Report**: [https://arxiv.org/pdf/2608.15045](https://arxiv.org/pdf/2608.15045)
56
 
57
- ## Overview
 
 
 
 
58
 
59
- MOSS-VL-Realtime is the realtime streaming checkpoint of the MOSS-VL release, part of the OpenMOSS ecosystem for open visual understanding.
60
 
61
- Unlike offline video-language models that first read a complete video and then answer, MOSS-VL-Realtime is designed for continuous video streams. It perceives incoming frames and generates text in parallel, supports questions at arbitrary moments in the stream, and can decide whether to respond or keep observing when the visual evidence is insufficient.
62
-
63
- This release keeps the MOSS-VL cross-attention design and a 256K text context window while adding realtime streaming data and an inference interface for timestamped frame-by-frame input.
64
-
65
- ## Key Features
66
-
67
- - Realtime streaming understanding: processes incoming frames continuously instead of waiting for a complete video.
68
- - Interruptible interaction: users can ask questions at any timestamp in a running stream, and the model answers based on the frames observed so far.
69
- - Proactive silence: the model can emit `<|silence|>` and continue observing when there is no meaningful visual update or the context is not sufficient.
70
- - Dynamic correction: as new frames arrive, the model can revise earlier responses instead of being locked to an initial interpretation.
71
- - Timestamp-aware frames: each streamed frame is associated with an absolute timestamp, helping the model reason about event order, duration, pacing, and fine-grained temporal localization.
72
- - Unified MOSS-VL family: released together with MOSS-VL-Instruct and MOSS-VL-Base for offline use, continued pretraining, fine-tuning, and applied research.
73
-
74
- ## Model Design
75
-
76
- ### Architecture
77
-
78
- MOSS-VL-Realtime adopts a cross-attention-based vision-language architecture that decouples visual encoding from language reasoning. This design is important for realtime usage because incoming visual content can be integrated into the running generation context without forcing the model into a strictly offline "load all frames, then answer" workflow.
79
 
80
- <p align="center">
81
- <img src="assets/architecture.png" alt="MOSS-VL Architecture" width="100%"/>
82
- </p>
83
 
84
- ### Timestamp-aware Video Encoding
85
 
86
- For video and realtime frame inputs, MOSS-VL injects absolute timestamps alongside sampled frames. This helps the model reason about when an event happens, how long it lasts, and how the scene changes over time instead of relying only on frame order.
87
 
88
- MOSS-VL also uses Cross-attention Rotary Position Embedding (XRoPE), which maps text tokens and visual patches into a unified three-dimensional coordinate space defined by Time (t), Height (h), and Width (w). This gives the model a consistent positional representation for image, offline video, and realtime streaming video reasoning.
89
 
90
- ### Configuration
91
 
92
  | Item | Value |
93
  | --- | --- |
94
  | Parameters | 11B |
95
- | Tensor type | BF16 |
96
- | Context length | 256K |
97
- | Vision patch size | 16 |
98
- | Temporal patch size | 1 |
99
- | Default video FPS | 1.0 |
100
- | Default max video frames | 256 |
101
- | Realtime frame format | PIL-compatible image plus timestamp |
102
- | Direct Transformers session scope | One active realtime session per model instance |
103
- | SGLang-Omni session scope | Configurable with `--max-running-requests`; subject to KV capacity |
104
-
105
- ## Performance
106
-
107
- MOSS-VL-Realtime is designed for streaming video understanding benchmarks where questions can arrive before a full video has been observed and correct answers may change as the scene evolves. It targets realtime interaction quality, proactive silence, and dynamic response updates in addition to standard video understanding accuracy.
108
-
109
- <p align="center">
110
- <img src="assets/benchmark-streaming.png" alt="MOSS-VL Streaming Benchmark" width="100%"/>
111
- </p>
112
-
113
- Detailed benchmark tables and comparisons for this release will be maintained in the MOSS-VL project resources.
114
-
115
- ## Quickstart
116
-
117
- ### Installation
118
-
119
- For the specialized backend, install its source and pinned dependencies in a fresh, compatible CUDA environment. Install `uv` first if it is not already available:
120
-
121
- ```bash
122
- git clone https://github.com/fnlp-vision/sglang-omni-realtime.git
123
- cd sglang-omni-realtime
124
- uv venv .venv -p 3.12
125
- source .venv/bin/activate
126
- uv pip install -e .
127
- ```
128
-
129
- See the [backend README](https://github.com/fnlp-vision/sglang-omni-realtime#readme) for CUDA/toolchain prerequisites and deployment details. This is not a universal installation recipe for arbitrary CUDA versions. For direct Transformers inference, use the same compatible Transformers 5.12.1 environment and the example below; an SGLang server does not need to be running.
130
-
131
- ### Load the Model
132
-
133
- ```python
134
- import torch
135
- from transformers import AutoModelForCausalLM, AutoProcessor
136
-
137
- checkpoint = "OpenMOSS-Team/MOSS-VL-Realtime-SGLANG"
138
-
139
- processor = AutoProcessor.from_pretrained(
140
- checkpoint,
141
- trust_remote_code=True,
142
- frame_extract_num_threads=1,
143
- )
144
- model = AutoModelForCausalLM.from_pretrained(
145
- checkpoint,
146
- trust_remote_code=True,
147
- device_map="auto",
148
- dtype=torch.bfloat16,
149
- attn_implementation="eager",
150
- )
151
- model.eval()
152
- ```
153
-
154
- This direct-Transformers example uses eager attention. The SGLang-Omni backend has its own attention configuration; do not confuse the two execution paths.
155
-
156
- ### Start the SGLang-Omni Backend
157
 
158
- Download this checkpoint to a directory outside the backend source tree, then launch from the backend repository root:
159
 
160
- ```bash
161
- hf download OpenMOSS-Team/MOSS-VL-Realtime-SGLANG \
162
- --local-dir /path/to/moss-vl-realtime-sglang
163
-
164
- python examples/run_moss_vl_realtime_server.py \
165
- --model-path /path/to/moss-vl-realtime-sglang \
166
- --gpu 0 --host 127.0.0.1 --port 8000 \
167
- --context-length 131072 \
168
- --mem-fraction-static 0.60 \
169
- --max-running-requests 1
170
- ```
171
-
172
- Access requires authorization while the repository is private. The command uses a 128K context as an explicit deployment example; the model's 256K capacity and launcher defaults do not guarantee sufficient KV memory on every device. Tune context, memory fraction and concurrency for your hardware. The backend endpoint is `/v1/video/realtime`; its WebSocket protocol and TP options are documented in the [Realtime Cookbook](https://github.com/fnlp-vision/sglang-omni-realtime/blob/main/docs/cookbook/moss_vl_realtime.md).
173
-
174
- This model repository contains weights and custom model/processor code, not the backend scheduler, Demo frontend, customer authentication or billing service.
175
-
176
- ## Inference Examples
177
-
178
- ### Online Inference
179
-
180
- <details>
181
- <summary><b>Session-style Online Inference</b></summary>
182
-
183
- The recommended direct API is `create_realtime_session(...)`. A service or application owns the video capture pipeline, converts camera, screen, or video-file input into PIL-compatible frames, and pushes each frame with a non-decreasing timestamp.
184
-
185
- Common session operations:
186
-
187
- - `session.push_frame(image, timestamp=...)` appends one visual frame.
188
- - `session.push_prompt("...")` appends a user question while the stream is running.
189
- - `session.push_prompt_frame(prompt, image, timestamp=...)` aligns a prompt with a specific frame.
190
- - `session.poll_output(...)` or `session.stream_outputs(...)` returns incremental text chunks.
191
-
192
- `system_prompt` and `initial_prompt` are tokenized as the initial system/user turns before the first frame arrives. Subsequent user turns can be appended with `push_prompt(...)` while the same session continues observing frames.
193
-
194
- For complete real-time inference usage, including local-video replay and service deployment, see [`realtime_inference`](https://github.com/OpenMOSS/MOSS-VL/tree/main/realtime_inference) in the MOSS-VL GitHub repository.
195
-
196
- ```python
197
- import time
198
- from PIL import Image
199
-
200
- session = model.create_realtime_session(
201
- processor,
202
- initial_prompt=(
203
- "As the video streams frame by frame, describe important changes as they happen. "
204
- "Stay silent when there is no relevant update."
205
- ),
206
- frame_queue_size=256,
207
- max_tokens_per_turn=12,
208
- max_new_tokens=4096,
209
- do_sample=False,
210
- )
211
-
212
- frame_paths = [
213
- "data/frame_0001.jpg",
214
- "data/frame_0002.jpg",
215
- "data/frame_0003.jpg",
216
- ]
217
-
218
- try:
219
- session.start()
220
-
221
- for index, frame_path in enumerate(frame_paths):
222
- image = Image.open(frame_path).convert("RGB")
223
- session.push_frame(image, timestamp=index / 1.0)
224
-
225
- while True:
226
- chunk = session.poll_output(timeout=0.0)
227
- if chunk is None:
228
- break
229
- print(chunk, end="", flush=True)
230
-
231
- time.sleep(1.0)
232
-
233
- session.push_prompt("What changed in the latest frames?")
234
-
235
- # Realtime sessions stay alive waiting for future input, so use a bounded
236
- # drain window and close the session explicitly when the producer is done.
237
- drain_deadline = time.monotonic() + 5.0
238
- while time.monotonic() < drain_deadline:
239
- chunk = session.poll_output(timeout=0.1)
240
- if chunk is not None:
241
- print(chunk, end="", flush=True)
242
- finally:
243
- session.close()
244
- ```
245
 
246
- Frame timestamps are measured in seconds and must be non-decreasing within a session. The input producer can be a camera, screen capture, decoded video file, browser frame sampler, or any other source that yields images with timestamps.
247
-
248
- </details>
249
-
250
- <details>
251
- <summary><b>Queue-style Online Inference</b></summary>
252
-
253
- `online_generate(...)` is useful for backend systems that separate frame production and model inference through queues. It accepts dictionaries containing frames, prompts, events, reset controls, and stop controls.
254
-
255
- ```python
256
- import queue
257
- import threading
258
- from PIL import Image
259
-
260
- input_queue = queue.Queue()
261
- output_queue = queue.Queue()
262
-
263
- worker = threading.Thread(
264
- target=model.online_generate,
265
- args=(processor, input_queue, output_queue),
266
- kwargs={
267
- "frame_queue_size": 256,
268
- "max_tokens_per_turn": 12,
269
- "max_new_tokens": 4096,
270
- "do_sample": False,
271
- },
272
- daemon=True,
273
- )
274
- worker.start()
275
-
276
- input_queue.put({
277
- "initial_prompt": "Answer only when the streamed video provides enough evidence.",
278
- })
279
-
280
- input_queue.put({"frame": Image.open("data/frame_0001.jpg").convert("RGB"), "timestamp": 0.0})
281
- input_queue.put({"frame": Image.open("data/frame_0002.jpg").convert("RGB"), "timestamp": 1.0})
282
- input_queue.put({"prompt": "What is happening now?"})
283
-
284
- try:
285
- while True:
286
- chunk = output_queue.get(timeout=0.5)
287
- print(chunk, end="", flush=True)
288
- except queue.Empty:
289
- pass
290
-
291
- input_queue.put({"stop_online_generate": True})
292
- worker.join()
293
- ```
294
 
295
- Each queue item can contain `frame` or `image`, `timestamp`, `prompt`, `frames`, `event`, `events`, `initial_prompt`, `system_prompt`, `generate_kwargs`, `reset_session`, or stop controls such as `stop_online_generate`.
296
-
297
- </details>
298
-
299
- ### Offline Inference
300
-
301
- MOSS-VL-Realtime also keeps the offline helper APIs for image and video prompts. For purely offline use, MOSS-VL-Instruct is usually the preferred checkpoint, but the realtime checkpoint can still process complete image and video inputs.
302
-
303
- <details>
304
- <summary><b>Single-video Offline Inference</b></summary>
305
-
306
- ```python
307
- video_path = "data/example_video.mp4"
308
- prompt = "Describe this video."
309
-
310
- text = model.offline_video_generate(
311
- processor,
312
- prompt=prompt,
313
- video=video_path,
314
- shortest_edge=4096,
315
- longest_edge=16777216,
316
- video_max_pixels=201326592,
317
- patch_size=16,
318
- temporal_patch_size=1,
319
- merge_size=2,
320
- video_fps=1.0,
321
- min_frames=1,
322
- max_frames=256,
323
- num_extract_threads=4,
324
- image_mean=[0.5, 0.5, 0.5],
325
- image_std=[0.5, 0.5, 0.5],
326
- max_new_tokens=256,
327
- temperature=1.0,
328
- top_k=50,
329
- top_p=1.0,
330
- repetition_penalty=1.0,
331
- do_sample=False,
332
- vision_chunked_length=64,
333
- )
334
-
335
- print(text)
336
- ```
337
 
338
- </details>
339
-
340
- <details>
341
- <summary><b>Batched Offline Inference</b></summary>
342
-
343
- `offline_batch_generate` accepts independent image/video/text queries. Queries in the same batch should share the same `media_kwargs` and `generate_kwargs`.
344
-
345
- ```python
346
- queries = [
347
- {
348
- "prompt": "Describe sample A.",
349
- "images": [],
350
- "videos": ["data/sample_a.mp4"],
351
- "media_kwargs": {
352
- "video_fps": 1.0,
353
- "min_frames": 8,
354
- "max_frames": 256,
355
- },
356
- "generate_kwargs": {
357
- "temperature": 1.0,
358
- "top_k": 50,
359
- "top_p": 1.0,
360
- "max_new_tokens": 256,
361
- "repetition_penalty": 1.0,
362
- "do_sample": False,
363
- },
364
- },
365
- {
366
- "prompt": "Describe sample B.",
367
- "images": [],
368
- "videos": ["data/sample_b.mp4"],
369
- "media_kwargs": {
370
- "video_fps": 1.0,
371
- "min_frames": 8,
372
- "max_frames": 256,
373
- },
374
- "generate_kwargs": {
375
- "temperature": 1.0,
376
- "top_k": 50,
377
- "top_p": 1.0,
378
- "max_new_tokens": 256,
379
- "repetition_penalty": 1.0,
380
- "do_sample": False,
381
- },
382
- },
383
- ]
384
-
385
- with torch.no_grad():
386
- result = model.offline_batch_generate(
387
- processor,
388
- queries,
389
- vision_chunked_length=64,
390
- )
391
-
392
- texts = [item["text"] for item in result["results"]]
393
- print(texts)
394
- ```
395
-
396
- </details>
397
 
398
- ## Related Checkpoints
399
 
400
- | Model | Parameters | Context | Usage | Hugging Face |
401
- | --- | ---: | ---: | --- | --- |
402
- | MOSS-VL-Realtime-SGLANG | 11B | 256K | Transformers 5.12.1 / specialized SGLang-Omni compatibility package | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-SGLANG |
403
- | MOSS-VL-Realtime | 11B | 256K | Realtime streaming video interaction | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime |
404
- | MOSS-VL-Instruct | 11B | 256K | Offline multimodal instruction following | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct |
405
- | MOSS-VL-Base | 11B | 256K | Continued pretraining and fine-tuning | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Base |
406
- | MOSS-VL-Instruct-0408 | 11B | 256K | Previous instruction-tuned checkpoint | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct-0408 |
407
- | MOSS-VL-Base-0408 | 11B | 256K | Previous base checkpoint | https://huggingface.co/OpenMOSS-Team/MOSS-VL-Base-0408 |
408
 
409
- ## Validation of This Compatibility Package
 
410
 
411
- On 2026-09-06, the Query RoPE and vision-frequency regression suite passed **17 tests**, including CUDA dtype/autocast and non-persistent-buffer rematerialization cases. An isolated H200 smoke test with Transformers 5.12.1 / PyTorch 2.11.0 loaded the full checkpoint, completed an image forward pass and generated a finite natural-language answer with eager attention.
412
 
413
- These checks cover the targeted fixes and basic model execution. They do not certify every live-stream schedule, full NPU compatibility, or a production latency/concurrency SLA. The current Transformers runtime still emits deprecation warnings for processor aliases and `cache_position`; future Transformers releases require separate adaptation rather than an unqualified dependency upgrade.
414
 
415
- ## Limitations and Roadmap
 
 
 
416
 
417
- MOSS-VL-Realtime is optimized for timestamped frame-by-frame streaming, but production latency depends on GPU hardware, frame sampling rate, transport overhead, and decoding speed. The direct Transformers API supports one active realtime session per model instance, and its bounded frame queue may drop older pending frames when overloaded. The specialized SGLang-Omni backend supports configured multi-session serving and uses protocol backpressure; these are different runtime contracts.
418
 
419
- The 1 FPS workflow is the semantic validation baseline. Real-time input arrival, sampling, session stop rules and optional vision-KV eviction can change outputs even when fixed-boundary comparisons agree. Neither the RoPE fixes nor a successful smoke test imply bitwise equality for every live HF/SGLang session or a production SLA. Optional frame pooling is not required for this checkpoint and remains disabled by default in the specialized backend.
 
 
 
 
420
 
421
- The model may emit realtime control tokens such as `<|silence|>`, `<|round_start|>`, and `<|round_end|>` depending on the application protocol. Downstream services should filter or render these tokens according to their UI needs.
422
 
423
- We are continuing to improve realtime response timing, dynamic correction, broader streaming evaluations, RL post-training, and task-specific deployment recipes for future MOSS-VL releases.
424
 
425
  ## Citation
426
 
 
22
  - sglang
23
  ---
24
 
 
 
 
 
25
  # MOSS-VL-Realtime-SGLANG
26
 
27
+ English | [简体中文](./README_zh.md)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
28
 
29
+ <p align="center"><img src="assets/logo.png" width="300" alt="MOSS-VL"/></p>
30
 
31
+ MOSS-VL-Realtime-SGLANG provides the MOSS-VL realtime checkpoint with Transformers 5.12.1-compatible model and processor code for the specialized SGLang-Omni backend. This is a compatibility package, not a retrained or quantized model.
32
 
33
+ ## Features
34
 
35
+ - Continuous frame-by-frame video understanding and incremental text generation.
36
+ - Questions at any point in a stream, with response interruption.
37
+ - Proactive silence when no response is needed or evidence is insufficient.
38
+ - Updated responses as new visual information arrives.
39
+ - Timestamp-aware reasoning about event order and duration.
40
 
41
+ ## Quick Start
42
 
43
+ | Goal | Entry |
44
+ | --- | --- |
45
+ | Browser video/voice interaction and optional memory | [MOSS-VL-Realtime Demo](https://github.com/fnlp-vision/MOSS-VL-Realtime_Demo#quick-start) |
46
+ | Standalone streaming inference service | [sglang-omni-realtime](https://github.com/fnlp-vision/sglang-omni-realtime#installation) |
47
+ | Direct Transformers inference | [Python examples](./inference.md) |
 
 
 
 
 
 
 
 
 
 
 
 
 
48
 
49
+ This repository contains weights, configuration, tokenizer, processor, and custom code. It does not include a running service. The model is public; download the complete repository rather than individual weight files.
 
 
50
 
51
+ ASR, TTS, and memory are provided by the Demo. The model's offline Python APIs do not mean that the Demo's realtime deployment enables offline chat.
52
 
53
+ ## Architecture and Configuration
54
 
55
+ MOSS-VL separates visual encoding from language reasoning with cross-attention. Timestamped frames and Cross-attention Rotary Position Embedding (XRoPE) align text and visual patches across time, height, and width.
56
 
57
+ <p align="center"><img src="assets/architecture.png" alt="MOSS-VL Architecture" width="100%"/></p>
58
 
59
  | Item | Value |
60
  | --- | --- |
61
  | Parameters | 11B |
62
+ | Weights | BF16 |
63
+ | Model context | 256K |
64
+ | Vision patch size / temporal patch size | 16 / 1 |
65
+ | Default video FPS / maximum sampled frames | 1.0 / 256 |
66
+ | Direct Transformers runtime | One active realtime session per model instance |
67
+ | SGLang-Omni runtime | Configurable session capacity, subject to GPU memory |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
68
 
69
+ The model context is not a per-session memory allocation guarantee. The recommended service configuration uses 131072 context; adjust context and concurrency for the available GPU memory.
70
 
71
+ ## Performance
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
72
 
73
+ MOSS-VL-Realtime targets streaming understanding, proactive silence, and dynamic response updates. Benchmark results are summarized below; see the [technical report](https://arxiv.org/abs/2608.15045) for model research.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
74
 
75
+ <p align="center"><img src="assets/benchmark-streaming.png" alt="MOSS-VL Streaming Benchmark" width="100%"/></p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
76
 
77
+ ## Compatibility
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78
 
79
+ Use this package's custom code and configuration together with the [specialized backend](https://github.com/fnlp-vision/sglang-omni-realtime). Installing a generic `sglang-omni` package does not supply the MOSS-VL integration.
80
 
81
+ The backend uses Transformers 5.12.1, SGLang 0.5.16, and PyTorch 2.11.0. This package adapts configuration, RoPE, and generation interfaces to Transformers 5.12.1 while keeping the original five BF16 weight shards, tokenizer, and vocabulary unchanged.
 
 
 
 
 
 
 
82
 
83
+ - **Cross-attention Query RoPE:** rotate each newly computed text query, including when visual KV is reused; cached visual keys are not rotated again. The corresponding fix is also available in the [original reference implementation](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime/commit/1e6a45b292eeaf02aa733bd3aa7b6c85214ddc86).
84
+ - **Vision rotary frequencies:** reconstruct canonical FP32 frequencies on the active device to handle buffer rematerialization in Transformers 5.12.1. This does not imply the same loading issue exists in the original 4.57 environment.
85
 
86
+ The original [MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) uses separate Transformers 4.57-series code. Do not mix its custom files into this package. Select matching model and backend versions; CUDA compatibility does not imply NPU support.
87
 
88
+ ## Limitations
89
 
90
+ - Response latency depends on hardware, frame rate, transport, and decoding settings.
91
+ - The direct Python runtime and the SGLang-Omni service have different session and backpressure behavior.
92
+ - Frame dropping and visual KV windows limit the visible history. Continuing beyond the context limit requires application-level memory handling.
93
+ - Applications should handle control tokens such as `<|silence|>`, `<|round_start|>`, and `<|round_end|>`.
94
 
95
+ ## Related Models
96
 
97
+ | Model | Purpose |
98
+ | --- | --- |
99
+ | [MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) | Original realtime checkpoint and reference implementation |
100
+ | [MOSS-VL-Instruct](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct) | Offline multimodal instruction following |
101
+ | [MOSS-VL-Base](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Base) | Continued pretraining and fine-tuning |
102
 
103
+ ## License
104
 
105
+ Apache-2.0. See [OpenMOSS/MOSS-VL](https://github.com/OpenMOSS/MOSS-VL) for the model family.
106
 
107
  ## Citation
108
 
README_zh.md ADDED
@@ -0,0 +1,105 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # MOSS-VL-Realtime-SGLANG
2
+
3
+ [English](./README.md) | 简体中文
4
+
5
+ <p align="center"><img src="assets/logo.png" width="300" alt="MOSS-VL"/></p>
6
+
7
+ MOSS-VL-Realtime-SGLANG 提供 MOSS-VL 实时模型权重,以及面向特化 SGLang-Omni 后端的 Transformers 5.12.1 兼容模型与 processor 代码。这是兼容版本,不是重新训练或量化的模型。
8
+
9
+ ## 功能
10
+
11
+ - 连续逐帧视频理解与增量文本生成。
12
+ - 在视频流任意时刻提问,并打断当前回答。
13
+ - 无需回答或证据不足时主动保持静默。
14
+ - 随新增视觉信息更新回答。
15
+ - 基于时间戳理解事件顺序和持续时间。
16
+
17
+ ## 快速开始
18
+
19
+ | 目标 | 入口 |
20
+ | --- | --- |
21
+ | 浏览器视频/语音交互与可选 memory | [MOSS-VL-Realtime Demo](https://github.com/fnlp-vision/MOSS-VL-Realtime_Demo/blob/main/README_zh.md#快速开始) |
22
+ | 独立流式推理服务 | [sglang-omni-realtime](https://github.com/fnlp-vision/sglang-omni-realtime/blob/main/README_zh.md#安装) |
23
+ | 直接使用 Transformers 推理 | [Python 示例](./inference.md) |
24
+
25
+ 本仓库包含权重、配置、tokenizer、processor 和自定义代码,不包含正在运行的服务。模型公开可访问,请下载完整仓库,而不是只下载个别权重文件。
26
+
27
+ ASR、TTS 和 memory 由 Demo 提供。模型保留离线 Python API,不代表 Demo 的实时部署已经启用离线聊天。
28
+
29
+ ## 架构与配置
30
+
31
+ MOSS-VL 使用交叉注意力分离视觉编码与语言推理。带时间戳的帧及 Cross-attention Rotary Position Embedding(XRoPE)在时间、高度和宽度维度上对齐文本与视觉 patch。
32
+
33
+ <p align="center"><img src="assets/architecture.png" alt="MOSS-VL 架构" width="100%"/></p>
34
+
35
+ | 项目 | 值 |
36
+ | --- | --- |
37
+ | 参数量 | 11B |
38
+ | 权重类型 | BF16 |
39
+ | 模型上下文 | 256K |
40
+ | 视觉 patch 大小 / 时间 patch 大小 | 16 / 1 |
41
+ | 默认视频 FPS / 最大采样帧数 | 1.0 / 256 |
42
+ | 直接 Transformers 推理 | 每个模型实例一个活跃实时会话 |
43
+ | SGLang-Omni 推理 | 会话容量可配置,受 GPU 显存限制 |
44
+
45
+ 模型上下文不等于部署时每路保证分配的容量。推荐服务配置使用 131072 context,请按可用显存调整上下文和并发数。
46
+
47
+ ## 性能
48
+
49
+ MOSS-VL-Realtime 面向流式理解、主动静默及动态回答更新。下图汇总相关基准结果,模型研究详见[技术报告](https://arxiv.org/abs/2608.15045)。
50
+
51
+ <p align="center"><img src="assets/benchmark-streaming.png" alt="MOSS-VL 流式基准" width="100%"/></p>
52
+
53
+ ## 兼容性
54
+
55
+ 本仓库的自定义代码和配置应配套使用,并连接[特化后端](https://github.com/fnlp-vision/sglang-omni-realtime)。安装通用 `sglang-omni` 包不等于获得 MOSS-VL 的接入实现。
56
+
57
+ 后端使用 Transformers 5.12.1、SGLang 0.5.16 和 PyTorch 2.11.0。本版本适配 Transformers 5.12.1 的配置、RoPE 和生成接口,原有五个 BF16 权重分片、tokenizer 和词表保持不变。
58
+
59
+ - **交叉注意力 Query RoPE:** 每次新计算的文本 query 都应用旋转,包括复用视觉 KV 的情况;已缓存的视觉 key 不重复旋转。对应修复也已应用于[原版参考实现](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime/commit/1e6a45b292eeaf02aa733bd3aa7b6c85214ddc86)。
60
+ - **视觉旋转频率:** 在当前设备重建标准 FP32 频率,以处理 Transformers 5.12.1 加载时重新创建 buffer 的情况,不代表原版 4.57 环境存在同样的加载问题。
61
+
62
+ 原版 [MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) 使用独立的 Transformers 4.57 系列代码,不要将其自定义文件混入此版本。模型与后端应配套选择版本;CUDA 兼容不等于支持 NPU.
63
+
64
+ ## 限制
65
+
66
+ - 响应延迟受硬件、帧率、传输和解码配置影响。
67
+ - 直接 Python 推理与 SGLang-Omni 服务具有不同的会话及背压行为。
68
+ - 丢帧和视觉 KV 窗口会限制可访问的视觉历史;跨上下文延续需要应用层记忆管理。
69
+ - 应用需处理 `<|silence|>`、`<|round_start|>`、`<|round_end|>` 等控制 token。
70
+
71
+ ## 相关模型
72
+
73
+ | 模型 | 用途 |
74
+ | --- | --- |
75
+ | [MOSS-VL-Realtime](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime) | 原版实时模型与参考实现 |
76
+ | [MOSS-VL-Instruct](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct) | 离线多模态指令理解 |
77
+ | [MOSS-VL-Base](https://huggingface.co/OpenMOSS-Team/MOSS-VL-Base) | 继续预训练与微调 |
78
+
79
+ ## 许可证
80
+
81
+ Apache-2.0。模型系列见 [OpenMOSS/MOSS-VL](https://github.com/OpenMOSS/MOSS-VL)。
82
+
83
+ ## 引用
84
+
85
+ ```bibtex
86
+ @misc{mossvl,
87
+ title = {MOSS-VL Technical Report},
88
+ author = {Wang, Pengyu and Tan, Chenkun and Zhou, Shaojun and Zhou, Qirui and Chen, Yanxin and He, Xingyang and Zeng, Huazheng and Cheng, Jijun and Wang, Chenghao and Qian, Xiaomeng and Wang, Pengfei and Huang, Zhan and Gao, Shanqing and Huang, Wei and Cao, Longjun and Ran, Wu and Liu, Jie and Zhu, Changtai and Wang, Hongkai and Tian, Yixian and Liu, Chenghao and Ye, Zhen and Wang, Xinghao and Jiang, Botian and Feng, Guoguo and Fei, Zhaoye and Li, Ruixiao and Chen, Mingshu and Gao, Yang and Cheng, Qinyuan and Li, Shimin and Qiu, Xipeng},
89
+ year = {2026},
90
+ eprint = {2608.15045},
91
+ archivePrefix = {arXiv},
92
+ primaryClass = {cs.CV},
93
+ url = {https://arxiv.org/abs/2608.15045}
94
+ }
95
+
96
+ @misc{mossvideopreview,
97
+ title = {{MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention}},
98
+ author = {Pengyu Wang and Chenkun Tan and Shaojun Zhou and Wei Huang and Qirui Zhou and Zhan Huang and Zhen Ye and Jijun Cheng and Xiaomeng Qian and Yanxin Chen and Xingyang He and Huazheng Zeng and Chenghao Wang and Pengfei Wang and Hongkai Wang and Shanqing Gao and Yixian Tian and Chenghao Liu and Xinghao Wang and Botian Jiang and Xipeng Qiu},
99
+ year = {2026},
100
+ eprint = {2606.07639},
101
+ archivePrefix = {arXiv},
102
+ primaryClass = {cs.CV},
103
+ url = {https://arxiv.org/abs/2606.07639}
104
+ }
105
+ ```
inference.md ADDED
@@ -0,0 +1,256 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Direct Transformers Inference
2
+
3
+ These examples use this repository's Transformers 5.12.1-compatible implementation, not the browser Demo or its WebSocket protocol. Prepare the [compatible backend environment](https://github.com/fnlp-vision/sglang-omni-realtime/blob/main/docs/get_started/installation.md); a running SGLang server is not required for direct Python inference. Do not use the Demo CPU environment to load this checkpoint.
4
+
5
+ Pin the model revision selected by the [compatibility manifest](https://github.com/fnlp-vision/MOSS-VL-Realtime_Demo/blob/main/deployment/repro/manifest.json). Sample paths below refer to media supplied by the caller.
6
+
7
+ ## Load the Model
8
+
9
+ ```python
10
+ import torch
11
+ from transformers import AutoModelForCausalLM, AutoProcessor
12
+
13
+ checkpoint = "OpenMOSS-Team/MOSS-VL-Realtime-SGLANG"
14
+ revision = "bcfd9ccf1e9db2896ad852301cc8dde4a6349c78"
15
+
16
+ processor = AutoProcessor.from_pretrained(
17
+ checkpoint,
18
+ revision=revision,
19
+ trust_remote_code=True,
20
+ frame_extract_num_threads=1,
21
+ )
22
+ model = AutoModelForCausalLM.from_pretrained(
23
+ checkpoint,
24
+ revision=revision,
25
+ trust_remote_code=True,
26
+ device_map="auto",
27
+ dtype=torch.bfloat16,
28
+ attn_implementation="eager",
29
+ )
30
+ model.eval()
31
+ ```
32
+
33
+ This direct-Transformers example uses eager attention. The SGLang-Omni backend has its own attention configuration; do not confuse the two execution paths.
34
+
35
+
36
+ ## Inference Examples
37
+
38
+ ### Online Inference
39
+
40
+ <details>
41
+ <summary><b>Session-style Online Inference</b></summary>
42
+
43
+ The recommended direct API is `create_realtime_session(...)`. A service or application owns the video capture pipeline, converts camera, screen, or video-file input into PIL-compatible frames, and pushes each frame with a non-decreasing timestamp.
44
+
45
+ Common session operations:
46
+
47
+ - `session.push_frame(image, timestamp=...)` appends one visual frame.
48
+ - `session.push_prompt("...")` appends a user question while the stream is running.
49
+ - `session.push_prompt_frame(prompt, image, timestamp=...)` aligns a prompt with a specific frame.
50
+ - `session.poll_output(...)` or `session.stream_outputs(...)` returns incremental text chunks.
51
+
52
+ `system_prompt` and `initial_prompt` are tokenized as the initial system/user turns before the first frame arrives. Subsequent user turns can be appended with `push_prompt(...)` while the same session continues observing frames.
53
+
54
+ For complete real-time inference usage, including local-video replay and service deployment, see [`realtime_inference`](https://github.com/OpenMOSS/MOSS-VL/tree/main/realtime_inference) in the MOSS-VL GitHub repository.
55
+
56
+ ```python
57
+ import time
58
+ from PIL import Image
59
+
60
+ session = model.create_realtime_session(
61
+ processor,
62
+ initial_prompt=(
63
+ "As the video streams frame by frame, describe important changes as they happen. "
64
+ "Stay silent when there is no relevant update."
65
+ ),
66
+ frame_queue_size=256,
67
+ max_tokens_per_turn=12,
68
+ max_new_tokens=4096,
69
+ do_sample=False,
70
+ )
71
+
72
+ frame_paths = [
73
+ "data/frame_0001.jpg",
74
+ "data/frame_0002.jpg",
75
+ "data/frame_0003.jpg",
76
+ ]
77
+
78
+ try:
79
+ session.start()
80
+
81
+ for index, frame_path in enumerate(frame_paths):
82
+ image = Image.open(frame_path).convert("RGB")
83
+ session.push_frame(image, timestamp=index / 1.0)
84
+
85
+ while True:
86
+ chunk = session.poll_output(timeout=0.0)
87
+ if chunk is None:
88
+ break
89
+ print(chunk, end="", flush=True)
90
+
91
+ time.sleep(1.0)
92
+
93
+ session.push_prompt("What changed in the latest frames?")
94
+
95
+ # Realtime sessions stay alive waiting for future input, so use a bounded
96
+ # drain window and close the session explicitly when the producer is done.
97
+ drain_deadline = time.monotonic() + 5.0
98
+ while time.monotonic() < drain_deadline:
99
+ chunk = session.poll_output(timeout=0.1)
100
+ if chunk is not None:
101
+ print(chunk, end="", flush=True)
102
+ finally:
103
+ session.close()
104
+ ```
105
+
106
+ Frame timestamps are measured in seconds and must be non-decreasing within a session. The input producer can be a camera, screen capture, decoded video file, browser frame sampler, or any other source that yields images with timestamps.
107
+
108
+ </details>
109
+
110
+ <details>
111
+ <summary><b>Queue-style Online Inference</b></summary>
112
+
113
+ `online_generate(...)` is useful for backend systems that separate frame production and model inference through queues. It accepts dictionaries containing frames, prompts, events, reset controls, and stop controls.
114
+
115
+ ```python
116
+ import queue
117
+ import threading
118
+ from PIL import Image
119
+
120
+ input_queue = queue.Queue()
121
+ output_queue = queue.Queue()
122
+
123
+ worker = threading.Thread(
124
+ target=model.online_generate,
125
+ args=(processor, input_queue, output_queue),
126
+ kwargs={
127
+ "frame_queue_size": 256,
128
+ "max_tokens_per_turn": 12,
129
+ "max_new_tokens": 4096,
130
+ "do_sample": False,
131
+ },
132
+ daemon=True,
133
+ )
134
+ worker.start()
135
+
136
+ input_queue.put({
137
+ "initial_prompt": "Answer only when the streamed video provides enough evidence.",
138
+ })
139
+
140
+ input_queue.put({"frame": Image.open("data/frame_0001.jpg").convert("RGB"), "timestamp": 0.0})
141
+ input_queue.put({"frame": Image.open("data/frame_0002.jpg").convert("RGB"), "timestamp": 1.0})
142
+ input_queue.put({"prompt": "What is happening now?"})
143
+
144
+ try:
145
+ while True:
146
+ chunk = output_queue.get(timeout=0.5)
147
+ print(chunk, end="", flush=True)
148
+ except queue.Empty:
149
+ pass
150
+
151
+ input_queue.put({"stop_online_generate": True})
152
+ worker.join()
153
+ ```
154
+
155
+ Each queue item can contain `frame` or `image`, `timestamp`, `prompt`, `frames`, `event`, `events`, `initial_prompt`, `system_prompt`, `generate_kwargs`, `reset_session`, or stop controls such as `stop_online_generate`.
156
+
157
+ </details>
158
+
159
+ ### Offline Inference
160
+
161
+ MOSS-VL-Realtime also keeps the offline helper APIs for image and video prompts. For purely offline use, MOSS-VL-Instruct is usually the preferred checkpoint, but the realtime checkpoint can still process complete image and video inputs.
162
+
163
+ <details>
164
+ <summary><b>Single-video Offline Inference</b></summary>
165
+
166
+ ```python
167
+ video_path = "data/example_video.mp4"
168
+ prompt = "Describe this video."
169
+
170
+ text = model.offline_video_generate(
171
+ processor,
172
+ prompt=prompt,
173
+ video=video_path,
174
+ shortest_edge=4096,
175
+ longest_edge=16777216,
176
+ video_max_pixels=201326592,
177
+ patch_size=16,
178
+ temporal_patch_size=1,
179
+ merge_size=2,
180
+ video_fps=1.0,
181
+ min_frames=1,
182
+ max_frames=256,
183
+ num_extract_threads=4,
184
+ image_mean=[0.5, 0.5, 0.5],
185
+ image_std=[0.5, 0.5, 0.5],
186
+ max_new_tokens=256,
187
+ temperature=1.0,
188
+ top_k=50,
189
+ top_p=1.0,
190
+ repetition_penalty=1.0,
191
+ do_sample=False,
192
+ vision_chunked_length=64,
193
+ )
194
+
195
+ print(text)
196
+ ```
197
+
198
+ </details>
199
+
200
+ <details>
201
+ <summary><b>Batched Offline Inference</b></summary>
202
+
203
+ `offline_batch_generate` accepts independent image/video/text queries. Queries in the same batch should share the same `media_kwargs` and `generate_kwargs`.
204
+
205
+ ```python
206
+ queries = [
207
+ {
208
+ "prompt": "Describe sample A.",
209
+ "images": [],
210
+ "videos": ["data/sample_a.mp4"],
211
+ "media_kwargs": {
212
+ "video_fps": 1.0,
213
+ "min_frames": 8,
214
+ "max_frames": 256,
215
+ },
216
+ "generate_kwargs": {
217
+ "temperature": 1.0,
218
+ "top_k": 50,
219
+ "top_p": 1.0,
220
+ "max_new_tokens": 256,
221
+ "repetition_penalty": 1.0,
222
+ "do_sample": False,
223
+ },
224
+ },
225
+ {
226
+ "prompt": "Describe sample B.",
227
+ "images": [],
228
+ "videos": ["data/sample_b.mp4"],
229
+ "media_kwargs": {
230
+ "video_fps": 1.0,
231
+ "min_frames": 8,
232
+ "max_frames": 256,
233
+ },
234
+ "generate_kwargs": {
235
+ "temperature": 1.0,
236
+ "top_k": 50,
237
+ "top_p": 1.0,
238
+ "max_new_tokens": 256,
239
+ "repetition_penalty": 1.0,
240
+ "do_sample": False,
241
+ },
242
+ },
243
+ ]
244
+
245
+ with torch.no_grad():
246
+ result = model.offline_batch_generate(
247
+ processor,
248
+ queries,
249
+ vision_chunked_length=64,
250
+ )
251
+
252
+ texts = [item["text"] for item in result["results"]]
253
+ print(texts)
254
+ ```
255
+
256
+ </details>