Text Generation
Transformers
Safetensors
English
Chinese
Russian
yue2
music-generation
orbitquant
quantization
4-bit precision
custom-code
8-bit precision
Instructions to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="WaveCut/YuE2-3B-OrbitQuant-W4A4")# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("WaveCut/YuE2-3B-OrbitQuant-W4A4", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WaveCut/YuE2-3B-OrbitQuant-W4A4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/YuE2-3B-OrbitQuant-W4A4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/WaveCut/YuE2-3B-OrbitQuant-W4A4
- SGLang
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WaveCut/YuE2-3B-OrbitQuant-W4A4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/YuE2-3B-OrbitQuant-W4A4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WaveCut/YuE2-3B-OrbitQuant-W4A4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/YuE2-3B-OrbitQuant-W4A4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with Docker Model Runner:
docker model run hf.co/WaveCut/YuE2-3B-OrbitQuant-W4A4
Document packed runtime and expose optional AR offload
Browse files- README.md +65 -4
- evaluation/ar-fused-fixed.json +33 -0
- evaluation/ar-gemv-fixed.json +33 -0
- evaluation/ar-original-fixed.json +33 -0
- evaluation/asr-source.json +4 -0
- evaluation/harness/ar_bench.py +37 -0
- evaluation/harness/ar_quality.py +41 -0
- evaluation/harness/asr_evaluate.py +42 -0
- evaluation/harness/benchmark.py +121 -0
- evaluation/harness/clean_validate.py +59 -0
- evaluation/harness/satellite.py +29 -0
- evaluation/harness/supervise.py +36 -0
- evaluation/harness/test_fused.py +27 -0
- evaluation/harness/test_gemv.py +35 -0
- evaluation/harness/test_packed_graph.py +54 -0
- evaluation/harness/vae_bench.py +36 -0
- evaluation/ru-ar-quality.json +11 -0
- evaluation/ru-asr.json +21 -0
- evaluation/ru-final-w4a4-full/run-00/metrics.json +82 -0
- evaluation/ru-final-w4a4-full/run-01/metrics.json +82 -0
- evaluation/ru-offload-w4a4-full/run-00/metrics.json +83 -0
- evaluation/ru-offload-w4a4-full/run-01/metrics.json +83 -0
- evaluation/ru-original-full/run-00/metrics.json +64 -0
- evaluation/ru-original-full/run-01/metrics.json +64 -0
- evaluation/ru-original-replay1500/run-00/metrics.json +36 -0
- evaluation/ru-original-replay1500/run-01/metrics.json +36 -0
- evaluation/vae-chunk-benchmark.json +101 -0
- run.py +4 -4
README.md
CHANGED
|
@@ -1,9 +1,70 @@
|
|
| 1 |
---
|
|
|
|
| 2 |
license: cc-by-nc-4.0
|
| 3 |
base_model: m-a-p/YuE2-3B
|
| 4 |
-
|
| 5 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
---
|
| 7 |
-
# YuE2-3B OrbitQuant W4A4
|
| 8 |
|
| 9 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
library_name: transformers
|
| 3 |
license: cc-by-nc-4.0
|
| 4 |
base_model: m-a-p/YuE2-3B
|
| 5 |
+
tags:
|
| 6 |
+
- music-generation
|
| 7 |
+
- orbitquant
|
| 8 |
+
- quantization
|
| 9 |
+
- 4-bit
|
| 10 |
+
- custom-code
|
| 11 |
+
language:
|
| 12 |
+
- en
|
| 13 |
+
- zh
|
| 14 |
+
- ru
|
| 15 |
---
|
| 16 |
+
# YuE2-3B · OrbitQuant W4A4
|
| 17 |
|
| 18 |
+
Packed 4-bit weights and activations for all 392 transformer projections in YuE2-3B, with a directly loadable runtime and a dedicated CUDA GEMV kernel. The default example is **«До утра»**, an original Russian synth-pop song. Russian is an experiment: the upstream model advertises Chinese and English, and this example does not establish general Russian-language reliability.
|
| 19 |
+
|
| 20 |
+
This is a **partial-model W4A4 release**. Embeddings, output heads, normalization, auxiliary projections and the separate FP32 audio VAE retain their original precision. No fine-tuning or distillation was performed.
|
| 21 |
+
|
| 22 |
+
## Architecture
|
| 23 |
+
|
| 24 |
+
YuE2 has 28 mixture-of-transformers layers with fixed AR/NAR token-type routing, hidden width 2048, MLP width 6144, 16 query heads and 8 KV heads. Each branch has Q/K/V/O attention and gate/up/down SwiGLU projections. RMSNorm and rotary position embeddings remain unchanged. This is not learned sparse MoE routing.
|
| 25 |
+
|
| 26 |
+
The autoregressive branch plans an ABC score and generates semantic audio tokens. The non-autoregressive transformer uses flow matching with 32 midpoint integration steps to produce 64-dimensional latents at 25 Hz. The separate Oobleck-style VAE decodes these to 48 kHz stereo using convolutions, transposed convolutions and SnakeBeta activations; it is not a transformer. The upstream semantic encoder is not included in the released generation checkpoint.
|
| 27 |
+
|
| 28 |
+
## Storage and execution
|
| 29 |
+
|
| 30 |
+
2,818,572,288 transformer weights occupy 1,409,286,144 packed bytes. The complete model checkpoint is **3,035,926,840 bytes**, versus **7,261,441,640 bytes** upstream: 2.39× smaller (58.2% reduction). The separate 530,512,720-byte VAE is unchanged and downloaded at a pinned revision.
|
| 31 |
+
|
| 32 |
+
Compatible Q/K/V and gate/up projections are grouped: 224 physical packed modules represent 392 logical projections without duplicated packed storage. A dedicated bias-free CUDA W4A4 GEMV handles 1–8 rows; larger inputs retain OrbitQuant's packed path. No persistent decoded BF16 or INT8 weight caches are used. Original AR CUDA graphs and prefix caching were already present upstream and are not credited as new optimizations.
|
| 33 |
+
|
| 34 |
+
The default VAE core is 512 frames with the original 16-frame halo. On the recorded Russian song this was waveform-identical to 1024 frames, while lowering the decoder's isolated memory peak. This is a tested example, not a universal bitwise-equivalence guarantee.
|
| 35 |
+
|
| 36 |
+
## Supported runtime
|
| 37 |
+
|
| 38 |
+
**Validated only on NVIDIA RTX 5090 (SM120), Linux x86_64, glibc ≥ 2.34, Python 3.12, PyTorch 2.10.0 / CUDA 12.8.** Packaged kernels use Python ABI3. Other GPUs and Torch/CUDA combinations require rebuilding and validation; the runner rejects unsupported GPU capability. This is not a generic `AutoModel.from_pretrained` integration.
|
| 39 |
+
|
| 40 |
+
```bash
|
| 41 |
+
hf download WaveCut/YuE2-3B-OrbitQuant-W4A4 --local-dir YuE2-W4A4
|
| 42 |
+
cd YuE2-W4A4
|
| 43 |
+
python3.12 -m venv .venv
|
| 44 |
+
source .venv/bin/activate
|
| 45 |
+
pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cu128
|
| 46 |
+
pip install -r requirements.txt
|
| 47 |
+
python run.py --output song
|
| 48 |
+
```
|
| 49 |
+
|
| 50 |
+
The output includes `audio.wav` and generation artifacts. `--prompt` accepts a JSON file containing `style`, `lyrics`, `seed`, and `cot`. `--low-memory` disables grouped projections; it is an alternative profile, not a guaranteed speedup. `--offload-ar` enables upstream AR offload: our same-song test was slower (44.33 vs 42.93 seconds) with unchanged overall peak before the VAE chunk-size change, so it is not the default. Acoustic CUDA graph capture was tested but not selected as a default due to its small benefit and additional memory.
|
| 51 |
+
|
| 52 |
+
## Measurement protocol
|
| 53 |
+
|
| 54 |
+
Measurements use one RTX 5090 rented at $0.69/hour. Reported generation time excludes checkpoint loading unless explicitly named otherwise. First-request and warm same-process results are separate. NVML peak is device memory sampled during generation; PyTorch allocated/reserved values are different metrics. Full songs may have different semantic sequences and durations even with the same prompt and seed. The controlled 60-second acoustic comparison replays identical 1500 semantic tokens and initial noise.
|
| 55 |
+
|
| 56 |
+
See `evaluation/` for exact metrics, hashes, source pins and the clean-load proof. The published runtime was downloaded into a separate HF cache and executed with the original checkpoint directory unavailable. All packed module counts and absence of decoded caches were asserted. Model checkpoint SHA256: `e169774b8e79f423e6e8586221558acd28a35bacc67aae912d9a8a905231134a`.
|
| 57 |
+
|
| 58 |
+
## Quality and limitations
|
| 59 |
+
|
| 60 |
+
For the controlled Russian acoustic pair, latent cosine similarity was 0.97822, waveform correlation 0.90434, and spectral convergence 0.16642. On 128 teacher-forced AR positions, mean KL was 0.03131 nats and top-1 agreement 82.81%. These are diagnostics, not perceptual scores.
|
| 61 |
+
|
| 62 |
+
Whisper-large-v3-turbo measured WER 13.87% for the original Russian song and 5.84% for the quantized song against the supplied lyrics. This is one example with different generated music and an imperfect automatic recognizer; it does not show a general quality improvement. Listen to both samples. MP3 previews are lossy encodes without loudness normalization or other postprocessing.
|
| 63 |
+
|
| 64 |
+
Earlier unoptimized packed execution was substantially slower than BF16, illustrating why packing alone is insufficient. The included kernel/runtime changes are required for the reported profile. Low-memory hardware feasibility is not established solely by the measured peak; loading and graph capture can have separate requirements.
|
| 65 |
+
|
| 66 |
+
## Provenance and license
|
| 67 |
+
|
| 68 |
+
Source pins are in `source-lock.json`; runtime changes are in `runtime/OPTIMIZATIONS.md`; standalone kernel sources and tests are in `kernel-source/`. OrbitQuant 0.9.2 is included as a wheel built from its pinned source revision. Original licenses and third-party notices are retained.
|
| 69 |
+
|
| 70 |
+
Model weights are **CC BY-NC 4.0**, inherited from YuE2. This release does not grant commercial rights. Runtime and kernel components retain their respective code licenses. Credit M-A-P / YuE2 and OrbitQuant when using this derivative.
|
evaluation/ar-fused-fixed.json
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model": "models/YuE2-OrbitQuant-W4A4",
|
| 3 |
+
"gemv": true,
|
| 4 |
+
"fuse_packed": true,
|
| 5 |
+
"replay": "benchmark/ru-original-full/run-00",
|
| 6 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 7 |
+
"results": [
|
| 8 |
+
{
|
| 9 |
+
"repeat": 0,
|
| 10 |
+
"prefill_seconds": 1.1080292710103095,
|
| 11 |
+
"wall_seconds": 2.194309923099354,
|
| 12 |
+
"cuda_ms": 2194.27587890625,
|
| 13 |
+
"tokens": 512,
|
| 14 |
+
"tps": 233.33075907382553
|
| 15 |
+
},
|
| 16 |
+
{
|
| 17 |
+
"repeat": 1,
|
| 18 |
+
"prefill_seconds": 0.14538272097706795,
|
| 19 |
+
"wall_seconds": 2.1817037297878414,
|
| 20 |
+
"cuda_ms": 2181.6689453125,
|
| 21 |
+
"tokens": 512,
|
| 22 |
+
"tps": 234.67897726415362
|
| 23 |
+
},
|
| 24 |
+
{
|
| 25 |
+
"repeat": 2,
|
| 26 |
+
"prefill_seconds": 0.14446253678761423,
|
| 27 |
+
"wall_seconds": 2.199943969026208,
|
| 28 |
+
"cuda_ms": 2199.907958984375,
|
| 29 |
+
"tokens": 512,
|
| 30 |
+
"tps": 232.7332001217439
|
| 31 |
+
}
|
| 32 |
+
]
|
| 33 |
+
}
|
evaluation/ar-gemv-fixed.json
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model": "models/YuE2-OrbitQuant-W4A4",
|
| 3 |
+
"gemv": true,
|
| 4 |
+
"fuse_packed": false,
|
| 5 |
+
"replay": "benchmark/ru-original-full/run-00",
|
| 6 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 7 |
+
"results": [
|
| 8 |
+
{
|
| 9 |
+
"repeat": 0,
|
| 10 |
+
"prefill_seconds": 1.698816369054839,
|
| 11 |
+
"wall_seconds": 2.863719446817413,
|
| 12 |
+
"cuda_ms": 2863.684326171875,
|
| 13 |
+
"tokens": 512,
|
| 14 |
+
"tps": 178.78846357278812
|
| 15 |
+
},
|
| 16 |
+
{
|
| 17 |
+
"repeat": 1,
|
| 18 |
+
"prefill_seconds": 0.17642259295098484,
|
| 19 |
+
"wall_seconds": 2.8788871839642525,
|
| 20 |
+
"cuda_ms": 2878.85400390625,
|
| 21 |
+
"tokens": 512,
|
| 22 |
+
"tps": 177.84649667826565
|
| 23 |
+
},
|
| 24 |
+
{
|
| 25 |
+
"repeat": 2,
|
| 26 |
+
"prefill_seconds": 0.17390872887335718,
|
| 27 |
+
"wall_seconds": 2.814904622035101,
|
| 28 |
+
"cuda_ms": 2814.87451171875,
|
| 29 |
+
"tokens": 512,
|
| 30 |
+
"tps": 181.88893363990346
|
| 31 |
+
}
|
| 32 |
+
]
|
| 33 |
+
}
|
evaluation/ar-original-fixed.json
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model": "models/YuE2-3B",
|
| 3 |
+
"gemv": false,
|
| 4 |
+
"fuse_packed": false,
|
| 5 |
+
"replay": "benchmark/ru-original-full/run-00",
|
| 6 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 7 |
+
"results": [
|
| 8 |
+
{
|
| 9 |
+
"repeat": 0,
|
| 10 |
+
"prefill_seconds": 0.4288525169249624,
|
| 11 |
+
"wall_seconds": 2.5710276649333537,
|
| 12 |
+
"cuda_ms": 2570.99365234375,
|
| 13 |
+
"tokens": 512,
|
| 14 |
+
"tps": 199.14215898305866
|
| 15 |
+
},
|
| 16 |
+
{
|
| 17 |
+
"repeat": 1,
|
| 18 |
+
"prefill_seconds": 0.09946558880619705,
|
| 19 |
+
"wall_seconds": 2.5770173692144454,
|
| 20 |
+
"cuda_ms": 2576.986328125,
|
| 21 |
+
"tokens": 512,
|
| 22 |
+
"tps": 198.6792972823747
|
| 23 |
+
},
|
| 24 |
+
{
|
| 25 |
+
"repeat": 2,
|
| 26 |
+
"prefill_seconds": 0.0979860769584775,
|
| 27 |
+
"wall_seconds": 2.5706858979538083,
|
| 28 |
+
"cuda_ms": 2570.656005859375,
|
| 29 |
+
"tokens": 512,
|
| 30 |
+
"tps": 199.16863449071596
|
| 31 |
+
}
|
| 32 |
+
]
|
| 33 |
+
}
|
evaluation/asr-source.json
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"repo": "openai/whisper-large-v3-turbo",
|
| 3 |
+
"revision": "41f01f3fe87f28c78e2fbf8b568835947dd65ed9"
|
| 4 |
+
}
|
evaluation/harness/ar_bench.py
ADDED
|
@@ -0,0 +1,37 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Identical prefix and forced token trace for fair AR runtime measurements."""
|
| 2 |
+
import argparse,json,time,gc
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
import numpy as np
|
| 5 |
+
import torch
|
| 6 |
+
from yue2.modeling_yue2 import YuE2ForCausalLM
|
| 7 |
+
from yue2.cuda_graph import GraphAR
|
| 8 |
+
from yue2.protocol import CODEC_OFFSET
|
| 9 |
+
from yue2_orbit import load_packed
|
| 10 |
+
|
| 11 |
+
@torch.inference_mode()
|
| 12 |
+
def main():
|
| 13 |
+
p=argparse.ArgumentParser();p.add_argument('--model',default='models/YuE2-3B');p.add_argument('--replay',default='benchmark/ru-original-full/run-00');p.add_argument('--output',required=True);p.add_argument('--gemv',action='store_true');p.add_argument('--fuse-packed',action='store_true');p.add_argument('--tokens',type=int,default=512);a=p.parse_args()
|
| 14 |
+
if a.gemv:
|
| 15 |
+
from enable_gemv import enable
|
| 16 |
+
enable()
|
| 17 |
+
if (Path(a.model)/'orbitquant_manifest.json').exists():
|
| 18 |
+
model,manifest=load_packed(a.model)
|
| 19 |
+
if a.fuse_packed:
|
| 20 |
+
from fuse_packed import fuse_model
|
| 21 |
+
fuse_model(model,manifest)
|
| 22 |
+
else:model=YuE2ForCausalLM.from_pretrained(a.model,torch_dtype=torch.bfloat16).cuda().eval()
|
| 23 |
+
prefix=np.load(Path(a.replay)/'prefix.npy').tolist();tokens=np.load(Path(a.replay)/'semantic.npy').tolist()[:a.tokens]
|
| 24 |
+
results=[]
|
| 25 |
+
for repeat in range(3):
|
| 26 |
+
g=GraphAR(model,[prefix],len(tokens)+1,capture=True)
|
| 27 |
+
t=time.perf_counter();g.prefill();torch.cuda.synchronize();prefill=time.perf_counter()-t
|
| 28 |
+
start,end=torch.cuda.Event(enable_timing=True),torch.cuda.Event(enable_timing=True)
|
| 29 |
+
torch.cuda.synchronize();t=time.perf_counter();start.record()
|
| 30 |
+
for token in tokens:g.step(int(token)+CODEC_OFFSET)
|
| 31 |
+
end.record();end.synchronize();wall=time.perf_counter()-t
|
| 32 |
+
results.append(dict(repeat=repeat,prefill_seconds=prefill,wall_seconds=wall,cuda_ms=start.elapsed_time(end),tokens=len(tokens),tps=len(tokens)/wall))
|
| 33 |
+
g.close()
|
| 34 |
+
value=dict(model=a.model,gemv=a.gemv,fuse_packed=a.fuse_packed,replay=a.replay,gpu=torch.cuda.get_device_name(),results=results)
|
| 35 |
+
Path(a.output).write_text(json.dumps(value,indent=2));print(json.dumps(value),flush=True)
|
| 36 |
+
|
| 37 |
+
if __name__=='__main__':main()
|
evaluation/harness/ar_quality.py
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Teacher-forced next-token KL on a saved baseline sequence, never sampled paths."""
|
| 3 |
+
import argparse
|
| 4 |
+
import gc
|
| 5 |
+
import json
|
| 6 |
+
from pathlib import Path
|
| 7 |
+
import numpy as np
|
| 8 |
+
import torch
|
| 9 |
+
from yue2.modeling_yue2 import YuE2ForCausalLM
|
| 10 |
+
from yue2.protocol import CODEC_OFFSET,CODEC_SIZE,MUSIC_END
|
| 11 |
+
from yue2_orbit import load_packed
|
| 12 |
+
|
| 13 |
+
|
| 14 |
+
@torch.inference_mode()
|
| 15 |
+
def logits(model,prefix,tokens):
|
| 16 |
+
# Position 0 predicts first semantic token; every later input is from baseline.
|
| 17 |
+
ids=torch.tensor([prefix+[t+CODEC_OFFSET for t in tokens[:127]]],device='cuda')
|
| 18 |
+
result=model(ids,logits_to_keep=128,use_cache=False).logits[0].float()
|
| 19 |
+
return torch.cat([result[:,MUSIC_END:MUSIC_END+1],result[:,CODEC_OFFSET:CODEC_OFFSET+CODEC_SIZE]],dim=1).cpu()
|
| 20 |
+
|
| 21 |
+
|
| 22 |
+
@torch.inference_mode()
|
| 23 |
+
def main():
|
| 24 |
+
p=argparse.ArgumentParser(); p.add_argument('--replay',default='benchmark/original-full/run-00'); p.add_argument('--candidate',default='models/YuE2-OrbitQuant-W4A4'); p.add_argument('--output',default='reports/ar-quality.json'); p.add_argument('--activation-bits',type=int); a=p.parse_args()
|
| 25 |
+
root=Path(a.replay); prefix=np.load(root/'prefix.npy').tolist(); tokens=np.load(root/'semantic.npy').tolist()
|
| 26 |
+
if len(tokens)<128: raise ValueError('At least 128 baseline semantic tokens required')
|
| 27 |
+
model=YuE2ForCausalLM.from_pretrained('models/YuE2-3B',torch_dtype=torch.bfloat16,local_files_only=True).cuda().eval()
|
| 28 |
+
ref=logits(model,prefix,tokens); del model; gc.collect(); torch.cuda.empty_cache()
|
| 29 |
+
model,manifest=load_packed(a.candidate,activation_bits=a.activation_bits)
|
| 30 |
+
quant=logits(model,prefix,tokens)
|
| 31 |
+
lr=ref.log_softmax(-1); lq=quant.log_softmax(-1)
|
| 32 |
+
kl=(lr.exp()*(lr-lq)).sum(-1)
|
| 33 |
+
targets=torch.tensor(tokens[:128])+1
|
| 34 |
+
result=dict(candidate=a.candidate,activation_bits=a.activation_bits or manifest['config']['activation_bits'],positions=128,
|
| 35 |
+
teacher_forced_kl_mean=float(kl.mean()),teacher_forced_kl_max=float(kl.max()),
|
| 36 |
+
top1_agreement=float((ref.argmax(-1)==quant.argmax(-1)).float().mean()),
|
| 37 |
+
source_nll=float(-lr.gather(1,targets[:,None]).mean()),candidate_nll=float(-lq.gather(1,targets[:,None]).mean()),
|
| 38 |
+
interpretation='128 positions from one held baseline prefix; numerical diagnostic, not a general quality benchmark')
|
| 39 |
+
Path(a.output).write_text(json.dumps(result,indent=2)); print(json.dumps(result),flush=True)
|
| 40 |
+
|
| 41 |
+
if __name__=='__main__': main()
|
evaluation/harness/asr_evaluate.py
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Whisper lyric transcription diagnostic; no perceptual-quality claims."""
|
| 2 |
+
import argparse,json,re,time
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
import numpy as np
|
| 5 |
+
import soundfile as sf
|
| 6 |
+
from scipy.signal import resample_poly
|
| 7 |
+
import torch
|
| 8 |
+
from transformers import AutoModelForSpeechSeq2Seq,AutoProcessor
|
| 9 |
+
|
| 10 |
+
def words(text):
|
| 11 |
+
text=re.sub(r'\[[^\]]*\]',' ',text).lower().replace('ё','е')
|
| 12 |
+
return re.findall(r'[a-zа-я0-9]+',text)
|
| 13 |
+
|
| 14 |
+
def distance(a,b):
|
| 15 |
+
prev=list(range(len(b)+1))
|
| 16 |
+
for i,x in enumerate(a,1):
|
| 17 |
+
row=[i]
|
| 18 |
+
for j,y in enumerate(b,1):row.append(min(row[-1]+1,prev[j]+1,prev[j-1]+(x!=y)))
|
| 19 |
+
prev=row
|
| 20 |
+
return prev[-1]
|
| 21 |
+
|
| 22 |
+
@torch.inference_mode()
|
| 23 |
+
def main():
|
| 24 |
+
p=argparse.ArgumentParser();p.add_argument('runs',nargs='+');p.add_argument('--output',required=True);a=p.parse_args()
|
| 25 |
+
model=AutoModelForSpeechSeq2Seq.from_pretrained('models/whisper-large-v3-turbo',torch_dtype=torch.float16,attn_implementation='sdpa').cuda().eval()
|
| 26 |
+
processor=AutoProcessor.from_pretrained('models/whisper-large-v3-turbo')
|
| 27 |
+
reference=words(json.loads(Path('prompts/default-ru.json').read_text())['lyrics']);report=[]
|
| 28 |
+
for run in a.runs:
|
| 29 |
+
signal,sr=sf.read(Path(run)/'audio.wav',dtype='float32');signal=signal.mean(axis=1)
|
| 30 |
+
from math import gcd
|
| 31 |
+
g=gcd(sr,16000);signal=resample_poly(signal,16000//g,sr//g)
|
| 32 |
+
inputs=processor(signal,sampling_rate=16000,return_tensors='pt',return_attention_mask=True,truncation=False,padding='longest')
|
| 33 |
+
inputs={k:v.cuda().to(torch.float16) if k=='input_features' else v.cuda() for k,v in inputs.items()}
|
| 34 |
+
start=time.perf_counter()
|
| 35 |
+
ids=model.generate(**inputs,language='ru',task='transcribe',return_timestamps=True)
|
| 36 |
+
text=processor.batch_decode(ids,skip_special_tokens=True)[0]
|
| 37 |
+
hypothesis=words(text)
|
| 38 |
+
item=dict(run=run,transcript=text,reference_words=len(reference),transcript_words=len(hypothesis),word_error_rate=distance(reference,hypothesis)/len(reference),seconds=time.perf_counter()-start)
|
| 39 |
+
report.append(item);print(json.dumps(item,ensure_ascii=False),flush=True)
|
| 40 |
+
Path(a.output).write_text(json.dumps(dict(results=report,interpretation='Whisper ASR is imperfect on singing; word error rate is a lyric intelligibility diagnostic, not a music-quality score.'),ensure_ascii=False,indent=2))
|
| 41 |
+
|
| 42 |
+
if __name__=='__main__':main()
|
evaluation/harness/benchmark.py
ADDED
|
@@ -0,0 +1,121 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Original/packed YuE2 measurements with replayable semantic and noise inputs."""
|
| 3 |
+
import argparse
|
| 4 |
+
import hashlib
|
| 5 |
+
import json
|
| 6 |
+
import os
|
| 7 |
+
from pathlib import Path
|
| 8 |
+
import threading
|
| 9 |
+
import time
|
| 10 |
+
|
| 11 |
+
import numpy as np
|
| 12 |
+
import torch
|
| 13 |
+
import pynvml
|
| 14 |
+
|
| 15 |
+
from yue2 import YuE2Pipeline
|
| 16 |
+
from yue2.pipeline import SemanticResult, SymbolicPlan
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
class Monitor:
|
| 20 |
+
def __init__(self):
|
| 21 |
+
pynvml.nvmlInit()
|
| 22 |
+
self.handle=pynvml.nvmlDeviceGetHandleByIndex(0)
|
| 23 |
+
self.stop=threading.Event()
|
| 24 |
+
self.peak=0
|
| 25 |
+
self.thread=threading.Thread(target=self.run,daemon=True)
|
| 26 |
+
def run(self):
|
| 27 |
+
while not self.stop.is_set():
|
| 28 |
+
self.peak=max(self.peak,pynvml.nvmlDeviceGetMemoryInfo(self.handle).used)
|
| 29 |
+
self.stop.wait(.05)
|
| 30 |
+
def __enter__(self):
|
| 31 |
+
torch.cuda.reset_peak_memory_stats()
|
| 32 |
+
self.thread.start()
|
| 33 |
+
return self
|
| 34 |
+
def __exit__(self,*args):
|
| 35 |
+
self.stop.set(); self.thread.join()
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def digest(array):
|
| 39 |
+
return hashlib.sha256(np.ascontiguousarray(array).tobytes()).hexdigest()
|
| 40 |
+
|
| 41 |
+
|
| 42 |
+
def main():
|
| 43 |
+
p=argparse.ArgumentParser()
|
| 44 |
+
p.add_argument('--model',default='models/YuE2-3B')
|
| 45 |
+
p.add_argument('--vae',default='models/YuE2-Vae')
|
| 46 |
+
p.add_argument('--prompt',default='prompts/default-ru.json')
|
| 47 |
+
p.add_argument('--output',required=True)
|
| 48 |
+
p.add_argument('--replay')
|
| 49 |
+
p.add_argument('--frames',type=int)
|
| 50 |
+
p.add_argument('--repeat',type=int,default=1)
|
| 51 |
+
p.add_argument('--activation-bits',type=int)
|
| 52 |
+
p.add_argument('--gemv',action='store_true')
|
| 53 |
+
p.add_argument('--fuse-packed',action='store_true')
|
| 54 |
+
p.add_argument('--nar-graph',action='store_true')
|
| 55 |
+
p.add_argument('--trim-cache',action='store_true')
|
| 56 |
+
p.add_argument('--offload-ar',action='store_true')
|
| 57 |
+
a=p.parse_args()
|
| 58 |
+
if a.nar_graph:
|
| 59 |
+
from nar_graph import enable as enable_nar
|
| 60 |
+
enable_nar()
|
| 61 |
+
if a.gemv:
|
| 62 |
+
from enable_gemv import enable
|
| 63 |
+
enable()
|
| 64 |
+
out=Path(a.output); out.mkdir(parents=True,exist_ok=True)
|
| 65 |
+
if (Path(a.model)/'orbitquant_manifest.json').exists():
|
| 66 |
+
from yue2_orbit import OrbitPipeline
|
| 67 |
+
pipe=OrbitPipeline.from_pretrained(a.model,vae=a.vae,device='cuda',memory_budget_gib=30,progress=False)
|
| 68 |
+
pipe.fused=a.fuse_packed
|
| 69 |
+
pipe.trim_cache=a.trim_cache
|
| 70 |
+
if a.activation_bits is not None:
|
| 71 |
+
pipe.activation_bits=a.activation_bits
|
| 72 |
+
else:
|
| 73 |
+
pipe=YuE2Pipeline.from_pretrained(a.model,vae=a.vae,device='cuda',memory_budget_gib=30,progress=False)
|
| 74 |
+
pipe.offload_ar=a.offload_ar
|
| 75 |
+
prompt=json.loads(Path(a.prompt).read_text())
|
| 76 |
+
request={k:prompt[k] for k in ('style','lyrics','seed','cot') if k in prompt}
|
| 77 |
+
print('loading model',flush=True)
|
| 78 |
+
t=time.perf_counter(); pipe._load_model(); torch.cuda.synchronize(); load=time.perf_counter()-t
|
| 79 |
+
for index in range(a.repeat):
|
| 80 |
+
dest=out/f'run-{index:02d}'; dest.mkdir(exist_ok=True)
|
| 81 |
+
print(f'generation {index} start replay={a.replay} frames={a.frames}',flush=True)
|
| 82 |
+
with Monitor() as monitor:
|
| 83 |
+
torch.cuda.synchronize(); start=time.perf_counter()
|
| 84 |
+
if a.replay:
|
| 85 |
+
original=Path(a.replay)
|
| 86 |
+
plan=SymbolicPlan.load(original)
|
| 87 |
+
tokens=np.load(original/'semantic.npy').tolist()
|
| 88 |
+
if a.frames: tokens=tokens[:a.frames]
|
| 89 |
+
semantic=SemanticResult(plan,tokens,{},False)
|
| 90 |
+
nar_start=time.perf_counter(); latents=pipe.synthesize(semantic); torch.cuda.synchronize()
|
| 91 |
+
nar_seconds=time.perf_counter()-nar_start
|
| 92 |
+
vae_start=time.perf_counter(); audio=pipe.decode(latents); torch.cuda.synchronize()
|
| 93 |
+
timing=dict(nar_seconds=nar_seconds,vae_seconds=time.perf_counter()-vae_start,
|
| 94 |
+
e2e_seconds=time.perf_counter()-start)
|
| 95 |
+
from yue2.pipeline import SongResult
|
| 96 |
+
song=SongResult(audio,48000,semantic,latents,pipe.effective_config(plan.request),pipe.weights,timing,'controlled-replay')
|
| 97 |
+
else:
|
| 98 |
+
song=pipe(**request)
|
| 99 |
+
torch.cuda.synchronize()
|
| 100 |
+
elapsed=time.perf_counter()-start
|
| 101 |
+
song.save_artifacts(dest)
|
| 102 |
+
song.save(dest/'audio.wav')
|
| 103 |
+
noise=torch.randn((len(song.semantic.tokens),64),generator=torch.Generator(device='cpu').manual_seed(song.semantic.plan.request.seed),dtype=torch.float32).numpy()
|
| 104 |
+
np.save(dest/'initial_noise.npy',noise)
|
| 105 |
+
metrics=dict(run=index,wall_seconds=elapsed,load_seconds=load,audio_seconds=len(song.audio)/48000,
|
| 106 |
+
rtf=elapsed/(len(song.audio)/48000),nvml_peak_bytes=monitor.peak,
|
| 107 |
+
torch_peak_allocated_bytes=torch.cuda.max_memory_allocated(),torch_peak_reserved_bytes=torch.cuda.max_memory_reserved(),
|
| 108 |
+
gpu=torch.cuda.get_device_name(),torch=torch.__version__,cuda=torch.version.cuda,
|
| 109 |
+
timing=song.timing,truncated=song.truncated,finite=bool(np.isfinite(song.audio).all()),
|
| 110 |
+
audio_peak=float(np.abs(song.audio).max()),audio_rms=float(np.sqrt(np.mean(song.audio**2))),
|
| 111 |
+
clipped_fraction=float(np.mean(np.abs(song.audio)>=1)),semantic_sha256=digest(np.asarray(song.semantic.tokens,dtype=np.int32)),
|
| 112 |
+
noise_sha256=digest(noise),latent_sha256=digest(song.latents),waveform_sha256=digest(song.audio),
|
| 113 |
+
source_or_quantized_model=str(a.model),replay=a.replay,frames=a.frames,gemv=a.gemv,fuse_packed=a.fuse_packed,nar_graph=a.nar_graph,trim_cache=a.trim_cache,offload_ar=a.offload_ar,
|
| 114 |
+
measurement='first process request' if index==0 else 'warm same-process request',pid=os.getpid())
|
| 115 |
+
if hasattr(pipe,'runtime_report'): metrics['packed_runtime']=pipe.runtime_report()
|
| 116 |
+
(dest/'metrics.json').write_text(json.dumps(metrics,indent=2))
|
| 117 |
+
print(json.dumps(metrics),flush=True)
|
| 118 |
+
del song
|
| 119 |
+
pipe.close()
|
| 120 |
+
|
| 121 |
+
if __name__=='__main__': main()
|
evaluation/harness/clean_validate.py
ADDED
|
@@ -0,0 +1,59 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Run the downloaded release with the original checkpoint path unavailable."""
|
| 2 |
+
from pathlib import Path
|
| 3 |
+
import json,os,sys,time,importlib.util,hashlib
|
| 4 |
+
from huggingface_hub import snapshot_download
|
| 5 |
+
import numpy as np
|
| 6 |
+
import torch
|
| 7 |
+
|
| 8 |
+
root=Path('/workspace/yue2')
|
| 9 |
+
info=json.loads((root/'reports/candidate-hub-final-runtime.json').read_text())
|
| 10 |
+
folder=Path(snapshot_download(info['repo'],revision=info['revision'],cache_dir=str(root/'clean-hf-cache')))
|
| 11 |
+
manifest=json.loads((folder/'weights_manifest.json').read_text())
|
| 12 |
+
def sha(path):
|
| 13 |
+
h=hashlib.sha256()
|
| 14 |
+
with path.open('rb') as f:
|
| 15 |
+
for block in iter(lambda:f.read(8*1024*1024),b''):h.update(block)
|
| 16 |
+
return h.hexdigest()
|
| 17 |
+
assert sha(folder/'model.safetensors')==manifest['files']['model.safetensors']['sha256']
|
| 18 |
+
original=root/'models/YuE2-3B';hidden=root/'models/YuE2-3B-clean-test-unavailable'
|
| 19 |
+
assert original.exists() and not hidden.exists()
|
| 20 |
+
original.rename(hidden)
|
| 21 |
+
try:
|
| 22 |
+
spec=importlib.util.spec_from_file_location('downloaded_release',folder/'run.py')
|
| 23 |
+
release=importlib.util.module_from_spec(spec);spec.loader.exec_module(release)
|
| 24 |
+
# VAE is also resolved by the released entrypoint at its pinned Hub revision.
|
| 25 |
+
pipe=release.load_pipeline(progress=False)
|
| 26 |
+
from benchmark import Monitor,digest
|
| 27 |
+
request=json.loads((folder/'prompts/default-ru.json').read_text())
|
| 28 |
+
start=time.perf_counter();pipe._load_model();torch.cuda.synchronize();load_seconds=time.perf_counter()-start
|
| 29 |
+
expected=json.loads((root/'benchmark/ru-final-w4a4-full/run-01/metrics.json').read_text())
|
| 30 |
+
results=[]
|
| 31 |
+
for i in range(2):
|
| 32 |
+
out=root/f'benchmark/clean-release/run-{i:02d}';out.mkdir(parents=True,exist_ok=True)
|
| 33 |
+
with Monitor() as monitor:
|
| 34 |
+
torch.cuda.synchronize();start=time.perf_counter()
|
| 35 |
+
song=pipe(**{k:request[k] for k in ['style','lyrics','cot','seed']})
|
| 36 |
+
torch.cuda.synchronize();elapsed=time.perf_counter()-start
|
| 37 |
+
song.save_artifacts(out);song.save(out/'audio.wav')
|
| 38 |
+
semantic=digest(np.asarray(song.semantic.tokens,dtype=np.int32));wave=digest(song.audio)
|
| 39 |
+
assert semantic==expected['semantic_sha256'],'Published runtime changed semantic generation'
|
| 40 |
+
assert wave==expected['waveform_sha256'],'Published runtime changed waveform'
|
| 41 |
+
runtime=pipe.runtime_report()
|
| 42 |
+
assert runtime['modules']==224 and runtime['logical_projections']==392
|
| 43 |
+
assert runtime['decoded_weight_caches']==0 and runtime['int8_weight_caches']==0
|
| 44 |
+
assert runtime['effective_modes']=={'native_packed_matmul':224}
|
| 45 |
+
report=dict(run=i,repo=info['repo'],revision=info['revision'],model_sha256=sha(folder/'model.safetensors'),
|
| 46 |
+
source_checkpoint_path_unavailable=not original.exists(),module_file=sys.modules['yue2'].__file__,
|
| 47 |
+
wall_seconds=elapsed,load_seconds=load_seconds,audio_seconds=len(song.audio)/48000,
|
| 48 |
+
rtf=elapsed/(len(song.audio)/48000),nvml_peak_bytes=monitor.peak,
|
| 49 |
+
torch_peak_allocated_bytes=torch.cuda.max_memory_allocated(),torch_peak_reserved_bytes=torch.cuda.max_memory_reserved(),
|
| 50 |
+
semantic_sha256=semantic,waveform_sha256=wave,packed_runtime=runtime,config=song.config,timing=song.timing,
|
| 51 |
+
finite=bool(np.isfinite(song.audio).all()),truncated=song.truncated,
|
| 52 |
+
measurement='first process request' if i==0 else 'warm same-process request')
|
| 53 |
+
(out/'metrics.json').write_text(json.dumps(report,indent=2))
|
| 54 |
+
results.append(report);print(json.dumps(report),flush=True)
|
| 55 |
+
del song
|
| 56 |
+
pipe.close()
|
| 57 |
+
(root/'reports/clean-load-proof.json').write_text(json.dumps(dict(status='passed',runs=results),indent=2))
|
| 58 |
+
finally:
|
| 59 |
+
hidden.rename(original)
|
evaluation/harness/satellite.py
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""A one-shot <60 second remote job/liveness snapshot, never a daemon."""
|
| 3 |
+
import argparse
|
| 4 |
+
import json
|
| 5 |
+
from pathlib import Path
|
| 6 |
+
import subprocess
|
| 7 |
+
import time
|
| 8 |
+
|
| 9 |
+
p=argparse.ArgumentParser()
|
| 10 |
+
p.add_argument("job")
|
| 11 |
+
a=p.parse_args()
|
| 12 |
+
job=Path(a.job)
|
| 13 |
+
start=time.monotonic()
|
| 14 |
+
for offset in (0,15,30,45):
|
| 15 |
+
time.sleep(max(0,start+offset-time.monotonic()))
|
| 16 |
+
value=json.loads((job/'status.json').read_text()) if (job/'status.json').exists() else {"status":"missing"}
|
| 17 |
+
alive=False
|
| 18 |
+
if value.get('pid'):
|
| 19 |
+
result=subprocess.run(['ps','-p',str(value['pid']),'-o','pid=,stat=,etime=,pcpu=,pmem=,comm='],capture_output=True,text=True)
|
| 20 |
+
alive=result.returncode==0
|
| 21 |
+
value['process']=result.stdout.strip()
|
| 22 |
+
log=job/'output.log'
|
| 23 |
+
value['log_age_seconds']=round(time.time()-log.stat().st_mtime,1) if log.exists() else None
|
| 24 |
+
value['log_tail']=subprocess.run(['tail','-n','8' if value['status']!='running' else '1',str(log)],capture_output=True,text=True).stdout.strip()
|
| 25 |
+
value['gpu']=subprocess.run(['nvidia-smi','--query-gpu=name,utilization.gpu,memory.used,power.draw','--format=csv,noheader'],capture_output=True,text=True,timeout=5).stdout.strip()
|
| 26 |
+
value['alive']=alive
|
| 27 |
+
print(json.dumps(value),flush=True)
|
| 28 |
+
if value['status']!='running' or not alive:
|
| 29 |
+
break
|
evaluation/harness/supervise.py
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Detached child runner with unbuffered minute heartbeats and atomic status."""
|
| 3 |
+
import argparse
|
| 4 |
+
import json
|
| 5 |
+
import os
|
| 6 |
+
from pathlib import Path
|
| 7 |
+
import subprocess
|
| 8 |
+
import time
|
| 9 |
+
|
| 10 |
+
p = argparse.ArgumentParser()
|
| 11 |
+
p.add_argument("--job", required=True)
|
| 12 |
+
p.add_argument("command", nargs=argparse.REMAINDER)
|
| 13 |
+
a = p.parse_args()
|
| 14 |
+
job = Path(a.job)
|
| 15 |
+
job.mkdir(parents=True, exist_ok=True)
|
| 16 |
+
command = a.command[1:] if a.command[:1] == ["--"] else a.command
|
| 17 |
+
started = time.time()
|
| 18 |
+
with (job / "output.log").open("a", buffering=1) as log:
|
| 19 |
+
child = subprocess.Popen(command, stdout=log, stderr=subprocess.STDOUT,
|
| 20 |
+
env={**os.environ, "PYTHONUNBUFFERED": "1"})
|
| 21 |
+
def status():
|
| 22 |
+
code = child.poll()
|
| 23 |
+
value = dict(status="running" if code is None else "complete" if code == 0 else "failed",
|
| 24 |
+
pid=child.pid, supervisor_pid=os.getpid(), started=started,
|
| 25 |
+
updated=time.time(), elapsed_seconds=time.time()-started, exit_code=code)
|
| 26 |
+
temporary = job / "status.tmp"
|
| 27 |
+
temporary.write_text(json.dumps(value, indent=2))
|
| 28 |
+
temporary.replace(job / "status.json")
|
| 29 |
+
print(f"heartbeat status={value['status']} pid={child.pid} elapsed={value['elapsed_seconds']:.0f}s exit={code}", flush=True)
|
| 30 |
+
return code
|
| 31 |
+
while status() is None:
|
| 32 |
+
try:
|
| 33 |
+
child.wait(timeout=60)
|
| 34 |
+
except subprocess.TimeoutExpired:
|
| 35 |
+
pass
|
| 36 |
+
raise SystemExit(child.returncode)
|
evaluation/harness/test_fused.py
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import torch
|
| 2 |
+
from orbitquant import OrbitQuantConfig
|
| 3 |
+
from orbitquant.layers import OrbitQuantLinear
|
| 4 |
+
from test_packed_graph import small_model
|
| 5 |
+
from fuse_packed import fuse_model
|
| 6 |
+
from yue2.cuda_graph import GraphAR
|
| 7 |
+
from enable_gemv import enable
|
| 8 |
+
|
| 9 |
+
def test_fused_graph_and_storage():
|
| 10 |
+
enable()
|
| 11 |
+
model=small_model()
|
| 12 |
+
before=sum(m.packed_weight_indices.numel() for m in model.modules() if isinstance(m,OrbitQuantLinear))
|
| 13 |
+
g=GraphAR(model,[[1,2,3]],8,capture=True)
|
| 14 |
+
expected=[g.prefill().clone()]+[g.step(t).clone() for t in (4,5,6,7)]
|
| 15 |
+
g.close()
|
| 16 |
+
fuse_model(model,{'config':OrbitQuantConfig().to_dict()})
|
| 17 |
+
after=sum(m.packed_weight_indices.numel() for m in model.modules() if isinstance(m,OrbitQuantLinear))
|
| 18 |
+
assert before==after
|
| 19 |
+
g=GraphAR(model,[[1,2,3]],8,capture=True)
|
| 20 |
+
actual=[g.prefill().clone()]+[g.step(t).clone() for t in (4,5,6,7)]
|
| 21 |
+
for a,b in zip(actual,expected):torch.testing.assert_close(a,b,rtol=0,atol=0)
|
| 22 |
+
g.close()
|
| 23 |
+
# Both device transfers retain one owner and correct compatibility views.
|
| 24 |
+
model.cpu().cuda()
|
| 25 |
+
for layer in model.model.layers:
|
| 26 |
+
assert layer.self_attn.q_proj._group is layer.self_attn.packed_qkv
|
| 27 |
+
assert layer.mlp.up_proj._group is layer.mlp.packed_gate_up
|
evaluation/harness/test_gemv.py
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import json
|
| 2 |
+
from pathlib import Path
|
| 3 |
+
import pytest
|
| 4 |
+
import torch
|
| 5 |
+
from kernels import get_local_kernel
|
| 6 |
+
from orbitquant.kernels.native_packed_matmul import matmul_packed_w4a4_int8_with_native_kernel as original
|
| 7 |
+
|
| 8 |
+
kernel=get_local_kernel(Path('/workspace/yue2/orbitquant-gemv/build'), 'orbitquant_gemv')
|
| 9 |
+
|
| 10 |
+
@pytest.mark.parametrize('rows,n,k',[(1,2048,2048),(2,6144,2048),(8,2048,6144),(1,129,128),(2,1024,2048)])
|
| 11 |
+
@pytest.mark.parametrize('dtype',[torch.bfloat16,torch.float16])
|
| 12 |
+
def test_exact_integer_gemv(rows,n,k,dtype):
|
| 13 |
+
torch.manual_seed(123)
|
| 14 |
+
x=torch.randint(0,256,(rows,k//2),device='cuda',dtype=torch.uint8)
|
| 15 |
+
w=torch.randint(0,256,(n*k//2,),device='cuda',dtype=torch.uint8)
|
| 16 |
+
xn=torch.rand(rows,device='cuda');wn=torch.rand(n,device='cuda',dtype=torch.bfloat16)
|
| 17 |
+
ac=torch.arange(-8,8,device='cuda',dtype=torch.int8);wc=ac.flip(0).contiguous()
|
| 18 |
+
bias=None
|
| 19 |
+
kw=dict(activation_scale=.03125,weight_scale=.0625,bias=bias,output_dtype=dtype)
|
| 20 |
+
ref=original(x,w,xn,wn,ac,wc,out_features=n,in_features=k,**kw)
|
| 21 |
+
out=kernel.gemv(x,w,xn,wn,ac,wc,**kw)
|
| 22 |
+
torch.testing.assert_close(out,ref,rtol=0,atol=0)
|
| 23 |
+
stream=torch.cuda.Stream();stream.wait_stream(torch.cuda.current_stream())
|
| 24 |
+
with torch.cuda.stream(stream):
|
| 25 |
+
for _ in range(3):kernel.gemv(x,w,xn,wn,ac,wc,**kw)
|
| 26 |
+
torch.cuda.current_stream().wait_stream(stream)
|
| 27 |
+
graph=torch.cuda.CUDAGraph()
|
| 28 |
+
with torch.cuda.graph(graph): captured=kernel.gemv(x,w,xn,wn,ac,wc,**kw)
|
| 29 |
+
graph.replay();torch.cuda.synchronize()
|
| 30 |
+
torch.testing.assert_close(captured,ref,rtol=0,atol=0)
|
| 31 |
+
|
| 32 |
+
def test_rejects_bad_weight_shape():
|
| 33 |
+
x=torch.zeros(1,64,device='cuda',dtype=torch.uint8)
|
| 34 |
+
with pytest.raises(RuntimeError,match='packed weight size'):
|
| 35 |
+
kernel.gemv(x,x.flatten(),torch.ones(1,device='cuda'),torch.ones(2,device='cuda',dtype=torch.bfloat16),torch.zeros(16,device='cuda',dtype=torch.int8),torch.zeros(16,device='cuda',dtype=torch.int8),activation_scale=1,weight_scale=1)
|
evaluation/harness/test_packed_graph.py
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The packed graph must match its eager path across KV updates and reload."""
|
| 2 |
+
import torch
|
| 3 |
+
from yue2.modeling_yue2 import YuE2Config, YuE2ForCausalLM
|
| 4 |
+
from yue2.cuda_graph import GraphAR
|
| 5 |
+
from orbitquant import OrbitQuantConfig
|
| 6 |
+
from orbitquant.layers import OrbitQuantLinear
|
| 7 |
+
|
| 8 |
+
|
| 9 |
+
def small_model():
|
| 10 |
+
torch.manual_seed(42)
|
| 11 |
+
config=YuE2Config(hidden_size=128,intermediate_size=256,num_hidden_layers=2,
|
| 12 |
+
num_attention_heads=4,num_key_value_heads=2,head_dim=32,
|
| 13 |
+
vocab_size=256,max_position_embeddings=128,max_latent_frames=128)
|
| 14 |
+
model=YuE2ForCausalLM(config).to(device='cuda',dtype=torch.bfloat16).eval()
|
| 15 |
+
for name,module in list(model.named_modules()):
|
| 16 |
+
if isinstance(module,torch.nn.Linear) and name.startswith('model.layers.'):
|
| 17 |
+
parent,leaf=name.rsplit('.',1)
|
| 18 |
+
model.get_submodule(parent)._modules[leaf]=OrbitQuantLinear.from_linear(module,config=OrbitQuantConfig(),module_name=name)
|
| 19 |
+
return model
|
| 20 |
+
|
| 21 |
+
|
| 22 |
+
def test_packed_graph_matches_eager_fixed_tokens():
|
| 23 |
+
model=small_model()
|
| 24 |
+
graph=GraphAR(model,[[1,2,3]],8,capture=True)
|
| 25 |
+
eager=GraphAR(model,[[1,2,3]],8,capture=False)
|
| 26 |
+
try:
|
| 27 |
+
torch.testing.assert_close(graph.prefill(),eager.prefill(),rtol=0,atol=0)
|
| 28 |
+
for token in (4,5,6,7):
|
| 29 |
+
torch.testing.assert_close(graph.step(token).clone(),eager.step(token),rtol=0,atol=0)
|
| 30 |
+
finally:
|
| 31 |
+
graph.close(); eager.close()
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def test_packed_graph_rejects_bf16_fusion():
|
| 35 |
+
import pytest
|
| 36 |
+
with pytest.raises(ValueError,match='packed'):
|
| 37 |
+
GraphAR(small_model(),[[1,2]],4,capture=False,fuse_projections=True)
|
| 38 |
+
|
| 39 |
+
|
| 40 |
+
def test_packed_reload_preserves_logits(tmp_path):
|
| 41 |
+
import json
|
| 42 |
+
from safetensors.torch import save_file
|
| 43 |
+
from yue2_orbit import load_packed
|
| 44 |
+
model=small_model()
|
| 45 |
+
modules=[dict(name=name,in_features=m.in_features,out_features=m.out_features,bias=m.bias is not None)
|
| 46 |
+
for name,m in model.named_modules() if isinstance(m,OrbitQuantLinear)]
|
| 47 |
+
(tmp_path/'config.json').write_text(json.dumps(model.config.to_dict()))
|
| 48 |
+
(tmp_path/'orbitquant_manifest.json').write_text(json.dumps(dict(config=OrbitQuantConfig().to_dict(),modules=modules)))
|
| 49 |
+
save_file({k:v.detach().cpu().contiguous() for k,v in model.state_dict().items()},str(tmp_path/'model.safetensors'))
|
| 50 |
+
restored,_=load_packed(tmp_path)
|
| 51 |
+
ids=torch.tensor([[1,2,3]],device='cuda')
|
| 52 |
+
with torch.inference_mode():
|
| 53 |
+
torch.testing.assert_close(model(ids).logits,restored(ids).logits,rtol=0,atol=0)
|
| 54 |
+
assert all(m._dequantized_weight_cache is None for m in restored.modules() if isinstance(m,OrbitQuantLinear))
|
evaluation/harness/vae_bench.py
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Measure the existing VAE chunk-size knob independently of transformer changes."""
|
| 2 |
+
import json,time
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
import numpy as np
|
| 5 |
+
import torch
|
| 6 |
+
import soundfile as sf
|
| 7 |
+
from yue2.modeling_vae import YuE2VAE
|
| 8 |
+
from benchmark import Monitor
|
| 9 |
+
|
| 10 |
+
@torch.inference_mode()
|
| 11 |
+
def main():
|
| 12 |
+
# Match the pinned pipeline's numerical settings, including convolution TF32.
|
| 13 |
+
torch.backends.cudnn.benchmark=False
|
| 14 |
+
torch.backends.cudnn.deterministic=True
|
| 15 |
+
torch.backends.cuda.matmul.allow_tf32=False
|
| 16 |
+
torch.backends.cudnn.allow_tf32=False
|
| 17 |
+
torch.backends.cuda.matmul.allow_fp16_reduced_precision_reduction=False
|
| 18 |
+
torch.set_float32_matmul_precision('highest')
|
| 19 |
+
model=YuE2VAE.from_pretrained('models/YuE2-Vae',decoder_only=True,device='cuda',local_files_only=True)
|
| 20 |
+
root=Path('benchmark/ru-final-w4a4-full/run-01')
|
| 21 |
+
z=torch.from_numpy(np.load(root/'latent.npy')).T.unsqueeze(0)
|
| 22 |
+
reference,sr=sf.read(root/'audio.wav',dtype='float32')
|
| 23 |
+
results=[]
|
| 24 |
+
for frames in (1024,512,256):
|
| 25 |
+
for repeat in range(3):
|
| 26 |
+
torch.cuda.empty_cache()
|
| 27 |
+
with Monitor() as monitor:
|
| 28 |
+
torch.cuda.synchronize();start=time.perf_counter()
|
| 29 |
+
audio=model.decode_tiled(z,core_frames=frames,halo_frames=16,output_device='cpu')[0].float().clamp(-1,1).T.contiguous().numpy()
|
| 30 |
+
torch.cuda.synchronize();elapsed=time.perf_counter()-start
|
| 31 |
+
diff=audio-reference
|
| 32 |
+
row=dict(pipeline_numerics=True,core_frames=frames,repeat=repeat,seconds=elapsed,nvml_peak_bytes=monitor.peak,torch_peak_allocated_bytes=torch.cuda.max_memory_allocated(),max_abs_difference=float(np.abs(diff).max()),relative_l2=float(np.linalg.norm(diff)/(np.linalg.norm(reference)+1e-12)),waveform_correlation=float(np.corrcoef(audio.ravel(),reference.ravel())[0,1]))
|
| 33 |
+
results.append(row);print(json.dumps(row),flush=True)
|
| 34 |
+
Path('reports/vae-chunk-benchmark.json').write_text(json.dumps(results,indent=2))
|
| 35 |
+
|
| 36 |
+
if __name__=='__main__': main()
|
evaluation/ru-ar-quality.json
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"candidate": "models/YuE2-OrbitQuant-W4A4",
|
| 3 |
+
"activation_bits": 4,
|
| 4 |
+
"positions": 128,
|
| 5 |
+
"teacher_forced_kl_mean": 0.031312376260757446,
|
| 6 |
+
"teacher_forced_kl_max": 0.12084035575389862,
|
| 7 |
+
"top1_agreement": 0.828125,
|
| 8 |
+
"source_nll": 3.8141632080078125,
|
| 9 |
+
"candidate_nll": 3.8379945755004883,
|
| 10 |
+
"interpretation": "128 positions from one held baseline prefix; numerical diagnostic, not a general quality benchmark"
|
| 11 |
+
}
|
evaluation/ru-asr.json
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"results": [
|
| 3 |
+
{
|
| 4 |
+
"run": "benchmark/ru-original-full/run-01",
|
| 5 |
+
"transcript": " Ночной трамвай Стираю с улиц тишину И этот город слышит нас одних Пускай часы торопятся вперед Мы не считаем пройденных шагов Пока над крышами рассвет встает Мне хватит самых обыкновенных слов Останьтесь со мной до утра Пока не погасли огни, нам эта короткая ночь Подарит другие дни. Останься со мной до утра и просто за руку держи. Когда просыпается город, мы снова учимся жить. На остановке мокрая сканья И теплый ветер трогает руках Я столько раз искал тебя во снах Но ты стоишь и смотришь на меня Останься со мной до утра, пока не потасли огни Нам эта короткая ночь подарит другие дни Останься со мной до утра и просто за руку держи Когда просыпается город Пока просыпается город Останьте со мной до утра",
|
| 6 |
+
"reference_words": 137,
|
| 7 |
+
"transcript_words": 124,
|
| 8 |
+
"word_error_rate": 0.1386861313868613,
|
| 9 |
+
"seconds": 1.5337588579859585
|
| 10 |
+
},
|
| 11 |
+
{
|
| 12 |
+
"run": "benchmark/ru-final-w4a4-full/run-01",
|
| 13 |
+
"transcript": " Ночной трамвай уходят за мосты В стекле дрожат последние дни Я собираю с улиц тишину И этот город слышит нас одни Пускай часы торопятся вперед Мы не считаем пройденных шагов Пока над крышами рассвет встает Мне хватит самых обыкновенных слов Останьте со мной до утра, пока не погасли огни Нам эта короткая ночь подарит другие дни Останьте со мной до утра и просто за руку держи Когда просыпается город, мы снова учимся жить На остановке мокрая сканья И теплый ветер трогает рукав Я столько раз искал тебя во снах Но ты стоишь и смотришь на меня Останься со мной до утра, пока не погасли огни Нам эта короткая ночь подарит другие дни Останься со мной до утра и просто за руку держи Когда просыпается город, мы снова учимся жить Пока просыпается город Останьте за мной до утра",
|
| 14 |
+
"reference_words": 137,
|
| 15 |
+
"transcript_words": 137,
|
| 16 |
+
"word_error_rate": 0.058394160583941604,
|
| 17 |
+
"seconds": 0.9663123651407659
|
| 18 |
+
}
|
| 19 |
+
],
|
| 20 |
+
"interpretation": "Whisper ASR is imperfect on singing; word error rate is a lyric intelligibility diagnostic, not a music-quality score."
|
| 21 |
+
}
|
evaluation/ru-final-w4a4-full/run-00/metrics.json
ADDED
|
@@ -0,0 +1,82 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"run": 0,
|
| 3 |
+
"wall_seconds": 45.04905349877663,
|
| 4 |
+
"load_seconds": 1.029712200164795,
|
| 5 |
+
"audio_seconds": 163.83866666666665,
|
| 6 |
+
"rtf": 0.2749598395501467,
|
| 7 |
+
"nvml_peak_bytes": 9756803072,
|
| 8 |
+
"torch_peak_allocated_bytes": 5495350272,
|
| 9 |
+
"torch_peak_reserved_bytes": 8495562752,
|
| 10 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 11 |
+
"torch": "2.10.0+cu128",
|
| 12 |
+
"cuda": "12.8",
|
| 13 |
+
"timing": {
|
| 14 |
+
"abc": {
|
| 15 |
+
"seconds": 10.528724248055369,
|
| 16 |
+
"prefill_seconds": 0.8370891781523824,
|
| 17 |
+
"ttft_seconds": 0.9959756580647081,
|
| 18 |
+
"output_tokens": 1772,
|
| 19 |
+
"content_tokens": 1771,
|
| 20 |
+
"output_tps": 168.30149201858757,
|
| 21 |
+
"prefix_tokens": 423,
|
| 22 |
+
"cfg_branches": 1,
|
| 23 |
+
"execution": "cuda_graph",
|
| 24 |
+
"attention": "flash"
|
| 25 |
+
},
|
| 26 |
+
"semantic": {
|
| 27 |
+
"seconds": 23.20693418988958,
|
| 28 |
+
"prefill_seconds": 0.3349543798249215,
|
| 29 |
+
"ttft_seconds": 0.33607721398584545,
|
| 30 |
+
"output_tokens": 4097,
|
| 31 |
+
"content_tokens": 4096,
|
| 32 |
+
"output_tps": 176.54206137167893,
|
| 33 |
+
"prefix_tokens": 2196,
|
| 34 |
+
"cfg_branches": 1,
|
| 35 |
+
"execution": "cuda_graph",
|
| 36 |
+
"attention": "flash"
|
| 37 |
+
},
|
| 38 |
+
"nar_seconds": 7.658700210042298,
|
| 39 |
+
"vae_seconds": 3.6495362292043865,
|
| 40 |
+
"load": {
|
| 41 |
+
"resolve_and_integrity_seconds": 2.4894183888100088,
|
| 42 |
+
"mot_load_seconds": 1.0296437400393188
|
| 43 |
+
},
|
| 44 |
+
"e2e_seconds": 45.04794106609188
|
| 45 |
+
},
|
| 46 |
+
"truncated": {
|
| 47 |
+
"abc": false,
|
| 48 |
+
"semantic": false
|
| 49 |
+
},
|
| 50 |
+
"finite": true,
|
| 51 |
+
"audio_peak": 0.9706120491027832,
|
| 52 |
+
"audio_rms": 0.14440932869911194,
|
| 53 |
+
"clipped_fraction": 0.0,
|
| 54 |
+
"semantic_sha256": "45d407d6c04fb19e54f29a17b622a89060240e5b8d8af4b0f3eee1fc338e5251",
|
| 55 |
+
"noise_sha256": "0f1e184cfce6b99c4e2b2f5dd27a470c894fc6a1dcbe6bd0cab6ec635d343f29",
|
| 56 |
+
"latent_sha256": "e371f6a61d2d2d13a9ff1003aeae8bf6355a62569c7707988c49f0053bd1e76f",
|
| 57 |
+
"waveform_sha256": "eceb3ad260035470803e5827bc88192819a5ca08f01536dd3d580abb4b190d54",
|
| 58 |
+
"source_or_quantized_model": "models/YuE2-OrbitQuant-W4A4",
|
| 59 |
+
"replay": null,
|
| 60 |
+
"frames": null,
|
| 61 |
+
"gemv": true,
|
| 62 |
+
"fuse_packed": true,
|
| 63 |
+
"nar_graph": false,
|
| 64 |
+
"trim_cache": true,
|
| 65 |
+
"measurement": "first process request",
|
| 66 |
+
"pid": 10795,
|
| 67 |
+
"packed_runtime": {
|
| 68 |
+
"modules": 224,
|
| 69 |
+
"logical_projections": 392,
|
| 70 |
+
"fused": true,
|
| 71 |
+
"effective_modes": {
|
| 72 |
+
"native_packed_matmul": 224
|
| 73 |
+
},
|
| 74 |
+
"activation_backends": {
|
| 75 |
+
"native_cuda_int8_surrogate": 168,
|
| 76 |
+
"triton_cuda_packed_w4": 56
|
| 77 |
+
},
|
| 78 |
+
"decoded_weight_caches": 0,
|
| 79 |
+
"int8_weight_caches": 0,
|
| 80 |
+
"packed_weight_bytes": 1409286144
|
| 81 |
+
}
|
| 82 |
+
}
|
evaluation/ru-final-w4a4-full/run-01/metrics.json
ADDED
|
@@ -0,0 +1,82 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"run": 1,
|
| 3 |
+
"wall_seconds": 42.934636591002345,
|
| 4 |
+
"load_seconds": 1.029712200164795,
|
| 5 |
+
"audio_seconds": 163.83866666666665,
|
| 6 |
+
"rtf": 0.2620543578907035,
|
| 7 |
+
"nvml_peak_bytes": 9857466368,
|
| 8 |
+
"torch_peak_allocated_bytes": 5782598144,
|
| 9 |
+
"torch_peak_reserved_bytes": 8596226048,
|
| 10 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 11 |
+
"torch": "2.10.0+cu128",
|
| 12 |
+
"cuda": "12.8",
|
| 13 |
+
"timing": {
|
| 14 |
+
"abc": {
|
| 15 |
+
"seconds": 9.711376916151494,
|
| 16 |
+
"prefill_seconds": 0.15963919600471854,
|
| 17 |
+
"ttft_seconds": 0.16055710101500154,
|
| 18 |
+
"output_tokens": 1772,
|
| 19 |
+
"content_tokens": 1771,
|
| 20 |
+
"output_tps": 182.46640155145198,
|
| 21 |
+
"prefix_tokens": 423,
|
| 22 |
+
"cfg_branches": 1,
|
| 23 |
+
"execution": "cuda_graph",
|
| 24 |
+
"attention": "flash"
|
| 25 |
+
},
|
| 26 |
+
"semantic": {
|
| 27 |
+
"seconds": 23.016886444063857,
|
| 28 |
+
"prefill_seconds": 0.15774337714537978,
|
| 29 |
+
"ttft_seconds": 0.1586433621123433,
|
| 30 |
+
"output_tokens": 4097,
|
| 31 |
+
"content_tokens": 4096,
|
| 32 |
+
"output_tps": 177.9997485740141,
|
| 33 |
+
"prefix_tokens": 2196,
|
| 34 |
+
"cfg_branches": 1,
|
| 35 |
+
"execution": "cuda_graph",
|
| 36 |
+
"attention": "flash"
|
| 37 |
+
},
|
| 38 |
+
"nar_seconds": 7.616418123943731,
|
| 39 |
+
"vae_seconds": 2.2683653261046857,
|
| 40 |
+
"load": {
|
| 41 |
+
"resolve_and_integrity_seconds": 2.4894183888100088,
|
| 42 |
+
"mot_load_seconds": 1.0296437400393188
|
| 43 |
+
},
|
| 44 |
+
"e2e_seconds": 42.93380210502073
|
| 45 |
+
},
|
| 46 |
+
"truncated": {
|
| 47 |
+
"abc": false,
|
| 48 |
+
"semantic": false
|
| 49 |
+
},
|
| 50 |
+
"finite": true,
|
| 51 |
+
"audio_peak": 0.9706120491027832,
|
| 52 |
+
"audio_rms": 0.14440932869911194,
|
| 53 |
+
"clipped_fraction": 0.0,
|
| 54 |
+
"semantic_sha256": "45d407d6c04fb19e54f29a17b622a89060240e5b8d8af4b0f3eee1fc338e5251",
|
| 55 |
+
"noise_sha256": "0f1e184cfce6b99c4e2b2f5dd27a470c894fc6a1dcbe6bd0cab6ec635d343f29",
|
| 56 |
+
"latent_sha256": "e371f6a61d2d2d13a9ff1003aeae8bf6355a62569c7707988c49f0053bd1e76f",
|
| 57 |
+
"waveform_sha256": "eceb3ad260035470803e5827bc88192819a5ca08f01536dd3d580abb4b190d54",
|
| 58 |
+
"source_or_quantized_model": "models/YuE2-OrbitQuant-W4A4",
|
| 59 |
+
"replay": null,
|
| 60 |
+
"frames": null,
|
| 61 |
+
"gemv": true,
|
| 62 |
+
"fuse_packed": true,
|
| 63 |
+
"nar_graph": false,
|
| 64 |
+
"trim_cache": true,
|
| 65 |
+
"measurement": "warm same-process request",
|
| 66 |
+
"pid": 10795,
|
| 67 |
+
"packed_runtime": {
|
| 68 |
+
"modules": 224,
|
| 69 |
+
"logical_projections": 392,
|
| 70 |
+
"fused": true,
|
| 71 |
+
"effective_modes": {
|
| 72 |
+
"native_packed_matmul": 224
|
| 73 |
+
},
|
| 74 |
+
"activation_backends": {
|
| 75 |
+
"native_cuda_int8_surrogate": 168,
|
| 76 |
+
"triton_cuda_packed_w4": 56
|
| 77 |
+
},
|
| 78 |
+
"decoded_weight_caches": 0,
|
| 79 |
+
"int8_weight_caches": 0,
|
| 80 |
+
"packed_weight_bytes": 1409286144
|
| 81 |
+
}
|
| 82 |
+
}
|
evaluation/ru-offload-w4a4-full/run-00/metrics.json
ADDED
|
@@ -0,0 +1,83 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"run": 0,
|
| 3 |
+
"wall_seconds": 46.23842817801051,
|
| 4 |
+
"load_seconds": 1.14848616393283,
|
| 5 |
+
"audio_seconds": 163.83866666666665,
|
| 6 |
+
"rtf": 0.2822192655661902,
|
| 7 |
+
"nvml_peak_bytes": 9756803072,
|
| 8 |
+
"torch_peak_allocated_bytes": 5495350272,
|
| 9 |
+
"torch_peak_reserved_bytes": 8495562752,
|
| 10 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 11 |
+
"torch": "2.10.0+cu128",
|
| 12 |
+
"cuda": "12.8",
|
| 13 |
+
"timing": {
|
| 14 |
+
"abc": {
|
| 15 |
+
"seconds": 10.56912293494679,
|
| 16 |
+
"prefill_seconds": 0.8523490249644965,
|
| 17 |
+
"ttft_seconds": 1.038547317031771,
|
| 18 |
+
"output_tokens": 1772,
|
| 19 |
+
"content_tokens": 1771,
|
| 20 |
+
"output_tps": 167.65818799787866,
|
| 21 |
+
"prefix_tokens": 423,
|
| 22 |
+
"cfg_branches": 1,
|
| 23 |
+
"execution": "cuda_graph",
|
| 24 |
+
"attention": "flash"
|
| 25 |
+
},
|
| 26 |
+
"semantic": {
|
| 27 |
+
"seconds": 23.106563664972782,
|
| 28 |
+
"prefill_seconds": 0.28357866406440735,
|
| 29 |
+
"ttft_seconds": 0.2846065699122846,
|
| 30 |
+
"output_tokens": 4097,
|
| 31 |
+
"content_tokens": 4096,
|
| 32 |
+
"output_tps": 177.30892656317556,
|
| 33 |
+
"prefix_tokens": 2196,
|
| 34 |
+
"cfg_branches": 1,
|
| 35 |
+
"execution": "cuda_graph",
|
| 36 |
+
"attention": "flash"
|
| 37 |
+
},
|
| 38 |
+
"nar_seconds": 9.069145092042163,
|
| 39 |
+
"vae_seconds": 3.4875931590795517,
|
| 40 |
+
"load": {
|
| 41 |
+
"resolve_and_integrity_seconds": 2.487326374044642,
|
| 42 |
+
"mot_load_seconds": 1.1484210339840502
|
| 43 |
+
},
|
| 44 |
+
"e2e_seconds": 46.23729805299081
|
| 45 |
+
},
|
| 46 |
+
"truncated": {
|
| 47 |
+
"abc": false,
|
| 48 |
+
"semantic": false
|
| 49 |
+
},
|
| 50 |
+
"finite": true,
|
| 51 |
+
"audio_peak": 0.9706120491027832,
|
| 52 |
+
"audio_rms": 0.14440932869911194,
|
| 53 |
+
"clipped_fraction": 0.0,
|
| 54 |
+
"semantic_sha256": "45d407d6c04fb19e54f29a17b622a89060240e5b8d8af4b0f3eee1fc338e5251",
|
| 55 |
+
"noise_sha256": "0f1e184cfce6b99c4e2b2f5dd27a470c894fc6a1dcbe6bd0cab6ec635d343f29",
|
| 56 |
+
"latent_sha256": "e371f6a61d2d2d13a9ff1003aeae8bf6355a62569c7707988c49f0053bd1e76f",
|
| 57 |
+
"waveform_sha256": "eceb3ad260035470803e5827bc88192819a5ca08f01536dd3d580abb4b190d54",
|
| 58 |
+
"source_or_quantized_model": "models/YuE2-OrbitQuant-W4A4",
|
| 59 |
+
"replay": null,
|
| 60 |
+
"frames": null,
|
| 61 |
+
"gemv": true,
|
| 62 |
+
"fuse_packed": true,
|
| 63 |
+
"nar_graph": false,
|
| 64 |
+
"trim_cache": true,
|
| 65 |
+
"offload_ar": true,
|
| 66 |
+
"measurement": "first process request",
|
| 67 |
+
"pid": 11579,
|
| 68 |
+
"packed_runtime": {
|
| 69 |
+
"modules": 224,
|
| 70 |
+
"logical_projections": 392,
|
| 71 |
+
"fused": true,
|
| 72 |
+
"effective_modes": {
|
| 73 |
+
"native_packed_matmul": 224
|
| 74 |
+
},
|
| 75 |
+
"activation_backends": {
|
| 76 |
+
"native_cuda_int8_surrogate": 168,
|
| 77 |
+
"triton_cuda_packed_w4": 56
|
| 78 |
+
},
|
| 79 |
+
"decoded_weight_caches": 0,
|
| 80 |
+
"int8_weight_caches": 0,
|
| 81 |
+
"packed_weight_bytes": 1409286144
|
| 82 |
+
}
|
| 83 |
+
}
|
evaluation/ru-offload-w4a4-full/run-01/metrics.json
ADDED
|
@@ -0,0 +1,83 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"run": 1,
|
| 3 |
+
"wall_seconds": 44.32611982291564,
|
| 4 |
+
"load_seconds": 1.14848616393283,
|
| 5 |
+
"audio_seconds": 163.83866666666665,
|
| 6 |
+
"rtf": 0.27054736665489415,
|
| 7 |
+
"nvml_peak_bytes": 9857466368,
|
| 8 |
+
"torch_peak_allocated_bytes": 5782598144,
|
| 9 |
+
"torch_peak_reserved_bytes": 8596226048,
|
| 10 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 11 |
+
"torch": "2.10.0+cu128",
|
| 12 |
+
"cuda": "12.8",
|
| 13 |
+
"timing": {
|
| 14 |
+
"abc": {
|
| 15 |
+
"seconds": 9.711860368028283,
|
| 16 |
+
"prefill_seconds": 0.15409997710958123,
|
| 17 |
+
"ttft_seconds": 0.15508739417418838,
|
| 18 |
+
"output_tokens": 1772,
|
| 19 |
+
"content_tokens": 1771,
|
| 20 |
+
"output_tps": 182.4573184591362,
|
| 21 |
+
"prefix_tokens": 423,
|
| 22 |
+
"cfg_branches": 1,
|
| 23 |
+
"execution": "cuda_graph",
|
| 24 |
+
"attention": "flash"
|
| 25 |
+
},
|
| 26 |
+
"semantic": {
|
| 27 |
+
"seconds": 23.071978967869654,
|
| 28 |
+
"prefill_seconds": 0.15794157283380628,
|
| 29 |
+
"ttft_seconds": 0.15890546003356576,
|
| 30 |
+
"output_tokens": 4097,
|
| 31 |
+
"content_tokens": 4096,
|
| 32 |
+
"output_tps": 177.57471111193092,
|
| 33 |
+
"prefix_tokens": 2196,
|
| 34 |
+
"cfg_branches": 1,
|
| 35 |
+
"execution": "cuda_graph",
|
| 36 |
+
"attention": "flash"
|
| 37 |
+
},
|
| 38 |
+
"nar_seconds": 8.86763684405014,
|
| 39 |
+
"vae_seconds": 2.319226206978783,
|
| 40 |
+
"load": {
|
| 41 |
+
"resolve_and_integrity_seconds": 2.487326374044642,
|
| 42 |
+
"mot_load_seconds": 1.1484210339840502
|
| 43 |
+
},
|
| 44 |
+
"e2e_seconds": 44.32531815697439
|
| 45 |
+
},
|
| 46 |
+
"truncated": {
|
| 47 |
+
"abc": false,
|
| 48 |
+
"semantic": false
|
| 49 |
+
},
|
| 50 |
+
"finite": true,
|
| 51 |
+
"audio_peak": 0.9706120491027832,
|
| 52 |
+
"audio_rms": 0.14440932869911194,
|
| 53 |
+
"clipped_fraction": 0.0,
|
| 54 |
+
"semantic_sha256": "45d407d6c04fb19e54f29a17b622a89060240e5b8d8af4b0f3eee1fc338e5251",
|
| 55 |
+
"noise_sha256": "0f1e184cfce6b99c4e2b2f5dd27a470c894fc6a1dcbe6bd0cab6ec635d343f29",
|
| 56 |
+
"latent_sha256": "e371f6a61d2d2d13a9ff1003aeae8bf6355a62569c7707988c49f0053bd1e76f",
|
| 57 |
+
"waveform_sha256": "eceb3ad260035470803e5827bc88192819a5ca08f01536dd3d580abb4b190d54",
|
| 58 |
+
"source_or_quantized_model": "models/YuE2-OrbitQuant-W4A4",
|
| 59 |
+
"replay": null,
|
| 60 |
+
"frames": null,
|
| 61 |
+
"gemv": true,
|
| 62 |
+
"fuse_packed": true,
|
| 63 |
+
"nar_graph": false,
|
| 64 |
+
"trim_cache": true,
|
| 65 |
+
"offload_ar": true,
|
| 66 |
+
"measurement": "warm same-process request",
|
| 67 |
+
"pid": 11579,
|
| 68 |
+
"packed_runtime": {
|
| 69 |
+
"modules": 224,
|
| 70 |
+
"logical_projections": 392,
|
| 71 |
+
"fused": true,
|
| 72 |
+
"effective_modes": {
|
| 73 |
+
"native_packed_matmul": 224
|
| 74 |
+
},
|
| 75 |
+
"activation_backends": {
|
| 76 |
+
"native_cuda_int8_surrogate": 168,
|
| 77 |
+
"triton_cuda_packed_w4": 56
|
| 78 |
+
},
|
| 79 |
+
"decoded_weight_caches": 0,
|
| 80 |
+
"int8_weight_caches": 0,
|
| 81 |
+
"packed_weight_bytes": 1409286144
|
| 82 |
+
}
|
| 83 |
+
}
|
evaluation/ru-original-full/run-00/metrics.json
ADDED
|
@@ -0,0 +1,64 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"run": 0,
|
| 3 |
+
"wall_seconds": 51.78475882997736,
|
| 4 |
+
"load_seconds": 3.197218642104417,
|
| 5 |
+
"audio_seconds": 164.63866666666667,
|
| 6 |
+
"rtf": 0.3145358248972135,
|
| 7 |
+
"nvml_peak_bytes": 10583080960,
|
| 8 |
+
"torch_peak_allocated_bytes": 8692496896,
|
| 9 |
+
"torch_peak_reserved_bytes": 9332326400,
|
| 10 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 11 |
+
"torch": "2.10.0+cu128",
|
| 12 |
+
"cuda": "12.8",
|
| 13 |
+
"timing": {
|
| 14 |
+
"abc": {
|
| 15 |
+
"seconds": 10.894125336082652,
|
| 16 |
+
"prefill_seconds": 0.48003934998996556,
|
| 17 |
+
"ttft_seconds": 0.6661999989300966,
|
| 18 |
+
"output_tokens": 1669,
|
| 19 |
+
"content_tokens": 1668,
|
| 20 |
+
"output_tps": 153.20183571526127,
|
| 21 |
+
"prefix_tokens": 423,
|
| 22 |
+
"cfg_branches": 1,
|
| 23 |
+
"execution": "cuda_graph",
|
| 24 |
+
"attention": "flash"
|
| 25 |
+
},
|
| 26 |
+
"semantic": {
|
| 27 |
+
"seconds": 26.13979333708994,
|
| 28 |
+
"prefill_seconds": 0.11473918193951249,
|
| 29 |
+
"ttft_seconds": 0.11574821802787483,
|
| 30 |
+
"output_tokens": 4117,
|
| 31 |
+
"content_tokens": 4116,
|
| 32 |
+
"output_tps": 157.49933241279146,
|
| 33 |
+
"prefix_tokens": 2093,
|
| 34 |
+
"cfg_branches": 1,
|
| 35 |
+
"execution": "cuda_graph",
|
| 36 |
+
"attention": "flash"
|
| 37 |
+
},
|
| 38 |
+
"nar_seconds": 8.817161875078455,
|
| 39 |
+
"vae_seconds": 5.922186557902023,
|
| 40 |
+
"load": {
|
| 41 |
+
"resolve_and_integrity_seconds": 4.985906707821414,
|
| 42 |
+
"mot_load_seconds": 0.2366083210799843
|
| 43 |
+
},
|
| 44 |
+
"e2e_seconds": 51.78419973095879
|
| 45 |
+
},
|
| 46 |
+
"truncated": {
|
| 47 |
+
"abc": false,
|
| 48 |
+
"semantic": false
|
| 49 |
+
},
|
| 50 |
+
"finite": true,
|
| 51 |
+
"audio_peak": 1.0,
|
| 52 |
+
"audio_rms": 0.15364815294742584,
|
| 53 |
+
"clipped_fraction": 1.3286672227666243e-06,
|
| 54 |
+
"semantic_sha256": "f37633f5006e9b80ae72a0321d663e1623e939d5963dd9493a1b592126db122e",
|
| 55 |
+
"noise_sha256": "c8aa36bd1c94b4d8d0f311f823e2e948ddc876c9951d98bc8a47e8c5b5bffb1c",
|
| 56 |
+
"latent_sha256": "752ebcaea4ffae13bfeb169655d6c84381b1db186871feffc0c0105670f16b73",
|
| 57 |
+
"waveform_sha256": "a8a6ea78c6b5dc4478b4e54d538625282ee87f3a10b4127323cc737f881e0338",
|
| 58 |
+
"source_or_quantized_model": "models/YuE2-3B",
|
| 59 |
+
"replay": null,
|
| 60 |
+
"frames": null,
|
| 61 |
+
"gemv": false,
|
| 62 |
+
"measurement": "first process request",
|
| 63 |
+
"pid": 5444
|
| 64 |
+
}
|
evaluation/ru-original-full/run-01/metrics.json
ADDED
|
@@ -0,0 +1,64 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"run": 1,
|
| 3 |
+
"wall_seconds": 50.16780533106066,
|
| 4 |
+
"load_seconds": 3.197218642104417,
|
| 5 |
+
"audio_seconds": 164.63866666666667,
|
| 6 |
+
"rtf": 0.3047145992297921,
|
| 7 |
+
"nvml_peak_bytes": 10908139520,
|
| 8 |
+
"torch_peak_allocated_bytes": 8977385472,
|
| 9 |
+
"torch_peak_reserved_bytes": 9646899200,
|
| 10 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 11 |
+
"torch": "2.10.0+cu128",
|
| 12 |
+
"cuda": "12.8",
|
| 13 |
+
"timing": {
|
| 14 |
+
"abc": {
|
| 15 |
+
"seconds": 10.320648127002642,
|
| 16 |
+
"prefill_seconds": 0.08523392281495035,
|
| 17 |
+
"ttft_seconds": 0.0860717708710581,
|
| 18 |
+
"output_tokens": 1669,
|
| 19 |
+
"content_tokens": 1668,
|
| 20 |
+
"output_tps": 161.714650035716,
|
| 21 |
+
"prefix_tokens": 423,
|
| 22 |
+
"cfg_branches": 1,
|
| 23 |
+
"execution": "cuda_graph",
|
| 24 |
+
"attention": "flash"
|
| 25 |
+
},
|
| 26 |
+
"semantic": {
|
| 27 |
+
"seconds": 26.154879739973694,
|
| 28 |
+
"prefill_seconds": 0.11253222916275263,
|
| 29 |
+
"ttft_seconds": 0.11347067612223327,
|
| 30 |
+
"output_tokens": 4117,
|
| 31 |
+
"content_tokens": 4116,
|
| 32 |
+
"output_tps": 157.40848518251076,
|
| 33 |
+
"prefix_tokens": 2093,
|
| 34 |
+
"cfg_branches": 1,
|
| 35 |
+
"execution": "cuda_graph",
|
| 36 |
+
"attention": "flash"
|
| 37 |
+
},
|
| 38 |
+
"nar_seconds": 8.794583374867216,
|
| 39 |
+
"vae_seconds": 4.3119754288345575,
|
| 40 |
+
"load": {
|
| 41 |
+
"resolve_and_integrity_seconds": 4.985906707821414,
|
| 42 |
+
"mot_load_seconds": 0.2366083210799843
|
| 43 |
+
},
|
| 44 |
+
"e2e_seconds": 50.16745935194194
|
| 45 |
+
},
|
| 46 |
+
"truncated": {
|
| 47 |
+
"abc": false,
|
| 48 |
+
"semantic": false
|
| 49 |
+
},
|
| 50 |
+
"finite": true,
|
| 51 |
+
"audio_peak": 1.0,
|
| 52 |
+
"audio_rms": 0.15364815294742584,
|
| 53 |
+
"clipped_fraction": 1.3286672227666243e-06,
|
| 54 |
+
"semantic_sha256": "f37633f5006e9b80ae72a0321d663e1623e939d5963dd9493a1b592126db122e",
|
| 55 |
+
"noise_sha256": "c8aa36bd1c94b4d8d0f311f823e2e948ddc876c9951d98bc8a47e8c5b5bffb1c",
|
| 56 |
+
"latent_sha256": "752ebcaea4ffae13bfeb169655d6c84381b1db186871feffc0c0105670f16b73",
|
| 57 |
+
"waveform_sha256": "a8a6ea78c6b5dc4478b4e54d538625282ee87f3a10b4127323cc737f881e0338",
|
| 58 |
+
"source_or_quantized_model": "models/YuE2-3B",
|
| 59 |
+
"replay": null,
|
| 60 |
+
"frames": null,
|
| 61 |
+
"gemv": false,
|
| 62 |
+
"measurement": "warm same-process request",
|
| 63 |
+
"pid": 5444
|
| 64 |
+
}
|
evaluation/ru-original-replay1500/run-00/metrics.json
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"run": 0,
|
| 3 |
+
"wall_seconds": 9.195442344993353,
|
| 4 |
+
"load_seconds": 3.375524572096765,
|
| 5 |
+
"audio_seconds": 59.998666666666665,
|
| 6 |
+
"rtf": 0.15326077821162726,
|
| 7 |
+
"nvml_peak_bytes": 9165406208,
|
| 8 |
+
"torch_peak_allocated_bytes": 7884326912,
|
| 9 |
+
"torch_peak_reserved_bytes": 8009023488,
|
| 10 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 11 |
+
"torch": "2.10.0+cu128",
|
| 12 |
+
"cuda": "12.8",
|
| 13 |
+
"timing": {
|
| 14 |
+
"nar_seconds": 3.627831295831129,
|
| 15 |
+
"vae_seconds": 5.565822640201077,
|
| 16 |
+
"e2e_seconds": 9.195178817026317
|
| 17 |
+
},
|
| 18 |
+
"truncated": {
|
| 19 |
+
"abc": false,
|
| 20 |
+
"semantic": false
|
| 21 |
+
},
|
| 22 |
+
"finite": true,
|
| 23 |
+
"audio_peak": 0.8817682862281799,
|
| 24 |
+
"audio_rms": 0.13392212986946106,
|
| 25 |
+
"clipped_fraction": 0.0,
|
| 26 |
+
"semantic_sha256": "105e53ffbbe671c28587ce52c917f57698087ed47bf6b63792a377e357236c6b",
|
| 27 |
+
"noise_sha256": "304490172c3b67d1dd3984d61d39e3cfc0d311ad6df79200674f815bc1081eae",
|
| 28 |
+
"latent_sha256": "b9461dadf0c0a16941fa0178ea5f3845f32d1f6bda70de4fe17541b2e81faabf",
|
| 29 |
+
"waveform_sha256": "9a80ae8267c22cd066aaf7cd1f7e5a12c7fcaedfb3a2e0b5f80b95ee2a47e482",
|
| 30 |
+
"source_or_quantized_model": "models/YuE2-3B",
|
| 31 |
+
"replay": "benchmark/ru-original-full/run-00",
|
| 32 |
+
"frames": 1500,
|
| 33 |
+
"gemv": false,
|
| 34 |
+
"measurement": "first process request",
|
| 35 |
+
"pid": 6449
|
| 36 |
+
}
|
evaluation/ru-original-replay1500/run-01/metrics.json
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"run": 1,
|
| 3 |
+
"wall_seconds": 7.551422053948045,
|
| 4 |
+
"load_seconds": 3.375524572096765,
|
| 5 |
+
"audio_seconds": 59.998666666666665,
|
| 6 |
+
"rtf": 0.12585983111760338,
|
| 7 |
+
"nvml_peak_bytes": 9469493248,
|
| 8 |
+
"torch_peak_allocated_bytes": 8150668800,
|
| 9 |
+
"torch_peak_reserved_bytes": 8304721920,
|
| 10 |
+
"gpu": "NVIDIA GeForce RTX 5090",
|
| 11 |
+
"torch": "2.10.0+cu128",
|
| 12 |
+
"cuda": "12.8",
|
| 13 |
+
"timing": {
|
| 14 |
+
"nar_seconds": 4.007696468848735,
|
| 15 |
+
"vae_seconds": 3.542188974097371,
|
| 16 |
+
"e2e_seconds": 7.551166985882446
|
| 17 |
+
},
|
| 18 |
+
"truncated": {
|
| 19 |
+
"abc": false,
|
| 20 |
+
"semantic": false
|
| 21 |
+
},
|
| 22 |
+
"finite": true,
|
| 23 |
+
"audio_peak": 0.8817682862281799,
|
| 24 |
+
"audio_rms": 0.13392212986946106,
|
| 25 |
+
"clipped_fraction": 0.0,
|
| 26 |
+
"semantic_sha256": "105e53ffbbe671c28587ce52c917f57698087ed47bf6b63792a377e357236c6b",
|
| 27 |
+
"noise_sha256": "304490172c3b67d1dd3984d61d39e3cfc0d311ad6df79200674f815bc1081eae",
|
| 28 |
+
"latent_sha256": "b9461dadf0c0a16941fa0178ea5f3845f32d1f6bda70de4fe17541b2e81faabf",
|
| 29 |
+
"waveform_sha256": "9a80ae8267c22cd066aaf7cd1f7e5a12c7fcaedfb3a2e0b5f80b95ee2a47e482",
|
| 30 |
+
"source_or_quantized_model": "models/YuE2-3B",
|
| 31 |
+
"replay": "benchmark/ru-original-full/run-00",
|
| 32 |
+
"frames": 1500,
|
| 33 |
+
"gemv": false,
|
| 34 |
+
"measurement": "warm same-process request",
|
| 35 |
+
"pid": 6449
|
| 36 |
+
}
|
evaluation/vae-chunk-benchmark.json
ADDED
|
@@ -0,0 +1,101 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[
|
| 2 |
+
{
|
| 3 |
+
"pipeline_numerics": true,
|
| 4 |
+
"core_frames": 1024,
|
| 5 |
+
"repeat": 0,
|
| 6 |
+
"seconds": 1.1320553519763052,
|
| 7 |
+
"nvml_peak_bytes": 8204910592,
|
| 8 |
+
"torch_peak_allocated_bytes": 2871368192,
|
| 9 |
+
"max_abs_difference": 0.0,
|
| 10 |
+
"relative_l2": 0.0,
|
| 11 |
+
"waveform_correlation": 1.0
|
| 12 |
+
},
|
| 13 |
+
{
|
| 14 |
+
"pipeline_numerics": true,
|
| 15 |
+
"core_frames": 1024,
|
| 16 |
+
"repeat": 1,
|
| 17 |
+
"seconds": 0.9049297058954835,
|
| 18 |
+
"nvml_peak_bytes": 8286699520,
|
| 19 |
+
"torch_peak_allocated_bytes": 2870057472,
|
| 20 |
+
"max_abs_difference": 0.0,
|
| 21 |
+
"relative_l2": 0.0,
|
| 22 |
+
"waveform_correlation": 1.0
|
| 23 |
+
},
|
| 24 |
+
{
|
| 25 |
+
"pipeline_numerics": true,
|
| 26 |
+
"core_frames": 1024,
|
| 27 |
+
"repeat": 2,
|
| 28 |
+
"seconds": 0.9072453379631042,
|
| 29 |
+
"nvml_peak_bytes": 8286699520,
|
| 30 |
+
"torch_peak_allocated_bytes": 2870581760,
|
| 31 |
+
"max_abs_difference": 0.0,
|
| 32 |
+
"relative_l2": 0.0,
|
| 33 |
+
"waveform_correlation": 1.0
|
| 34 |
+
},
|
| 35 |
+
{
|
| 36 |
+
"pipeline_numerics": true,
|
| 37 |
+
"core_frames": 512,
|
| 38 |
+
"repeat": 0,
|
| 39 |
+
"seconds": 0.9228434218093753,
|
| 40 |
+
"nvml_peak_bytes": 5040308224,
|
| 41 |
+
"torch_peak_allocated_bytes": 1737859584,
|
| 42 |
+
"max_abs_difference": 0.0,
|
| 43 |
+
"relative_l2": 0.0,
|
| 44 |
+
"waveform_correlation": 1.0
|
| 45 |
+
},
|
| 46 |
+
{
|
| 47 |
+
"pipeline_numerics": true,
|
| 48 |
+
"core_frames": 512,
|
| 49 |
+
"repeat": 1,
|
| 50 |
+
"seconds": 0.9203053209930658,
|
| 51 |
+
"nvml_peak_bytes": 5010948096,
|
| 52 |
+
"torch_peak_allocated_bytes": 1737193984,
|
| 53 |
+
"max_abs_difference": 0.0,
|
| 54 |
+
"relative_l2": 0.0,
|
| 55 |
+
"waveform_correlation": 1.0
|
| 56 |
+
},
|
| 57 |
+
{
|
| 58 |
+
"pipeline_numerics": true,
|
| 59 |
+
"core_frames": 512,
|
| 60 |
+
"repeat": 2,
|
| 61 |
+
"seconds": 0.9190280558541417,
|
| 62 |
+
"nvml_peak_bytes": 5021433856,
|
| 63 |
+
"torch_peak_allocated_bytes": 1738250752,
|
| 64 |
+
"max_abs_difference": 0.0,
|
| 65 |
+
"relative_l2": 0.0,
|
| 66 |
+
"waveform_correlation": 1.0
|
| 67 |
+
},
|
| 68 |
+
{
|
| 69 |
+
"pipeline_numerics": true,
|
| 70 |
+
"core_frames": 256,
|
| 71 |
+
"repeat": 0,
|
| 72 |
+
"seconds": 1.0967565891332924,
|
| 73 |
+
"nvml_peak_bytes": 3282894848,
|
| 74 |
+
"torch_peak_allocated_bytes": 1172341248,
|
| 75 |
+
"max_abs_difference": 3.874301910400391e-06,
|
| 76 |
+
"relative_l2": 1.3670459111381206e-06,
|
| 77 |
+
"waveform_correlation": 0.9999999999990744
|
| 78 |
+
},
|
| 79 |
+
{
|
| 80 |
+
"pipeline_numerics": true,
|
| 81 |
+
"core_frames": 256,
|
| 82 |
+
"repeat": 1,
|
| 83 |
+
"seconds": 1.0739446009974927,
|
| 84 |
+
"nvml_peak_bytes": 3282894848,
|
| 85 |
+
"torch_peak_allocated_bytes": 1172406784,
|
| 86 |
+
"max_abs_difference": 3.874301910400391e-06,
|
| 87 |
+
"relative_l2": 1.3670459111381206e-06,
|
| 88 |
+
"waveform_correlation": 0.9999999999990744
|
| 89 |
+
},
|
| 90 |
+
{
|
| 91 |
+
"pipeline_numerics": true,
|
| 92 |
+
"core_frames": 256,
|
| 93 |
+
"repeat": 2,
|
| 94 |
+
"seconds": 1.1024859179742634,
|
| 95 |
+
"nvml_peak_bytes": 3316449280,
|
| 96 |
+
"torch_peak_allocated_bytes": 1172470272,
|
| 97 |
+
"max_abs_difference": 3.874301910400391e-06,
|
| 98 |
+
"relative_l2": 1.3670459111381206e-06,
|
| 99 |
+
"waveform_correlation": 0.9999999999990744
|
| 100 |
+
}
|
| 101 |
+
]
|
run.py
CHANGED
|
@@ -6,7 +6,7 @@ ROOT=Path(__file__).absolute().parent
|
|
| 6 |
sys.path.insert(0,str(ROOT/'runtime/kernels/base'))
|
| 7 |
sys.path.insert(0,str(ROOT/'runtime'))
|
| 8 |
|
| 9 |
-
def load_pipeline(*,vae=None,low_memory=False,progress=True):
|
| 10 |
import torch
|
| 11 |
if not torch.cuda.is_available():raise RuntimeError('This release requires an NVIDIA CUDA GPU.')
|
| 12 |
if torch.cuda.get_device_capability()!=(12,0):
|
|
@@ -19,15 +19,15 @@ def load_pipeline(*,vae=None,low_memory=False,progress=True):
|
|
| 19 |
from enable_gemv import enable
|
| 20 |
from yue2_orbit import OrbitPipeline
|
| 21 |
enable()
|
| 22 |
-
pipe=OrbitPipeline.from_pretrained(str(ROOT),vae=vae,device='cuda',memory_budget_gib=30,vae_core_frames=512,progress=progress)
|
| 23 |
pipe.fused=not low_memory
|
| 24 |
pipe.trim_cache=True
|
| 25 |
return pipe
|
| 26 |
|
| 27 |
def main():
|
| 28 |
-
p=argparse.ArgumentParser();p.add_argument('--prompt',default=str(ROOT/'prompts/default-ru.json'));p.add_argument('--output',default='song');p.add_argument('--vae');p.add_argument('--low-memory',action='store_true');a=p.parse_args()
|
| 29 |
request=json.loads(Path(a.prompt).read_text());out=Path(a.output);out.mkdir(parents=True,exist_ok=True)
|
| 30 |
-
with load_pipeline(vae=a.vae,low_memory=a.low_memory) as pipe:
|
| 31 |
start=time.perf_counter()
|
| 32 |
song=pipe(**{k:request[k] for k in ('style','lyrics','seed','cot') if k in request})
|
| 33 |
song.save_artifacts(out);song.save(out/'audio.wav')
|
|
|
|
| 6 |
sys.path.insert(0,str(ROOT/'runtime/kernels/base'))
|
| 7 |
sys.path.insert(0,str(ROOT/'runtime'))
|
| 8 |
|
| 9 |
+
def load_pipeline(*,vae=None,low_memory=False,offload_ar=False,progress=True):
|
| 10 |
import torch
|
| 11 |
if not torch.cuda.is_available():raise RuntimeError('This release requires an NVIDIA CUDA GPU.')
|
| 12 |
if torch.cuda.get_device_capability()!=(12,0):
|
|
|
|
| 19 |
from enable_gemv import enable
|
| 20 |
from yue2_orbit import OrbitPipeline
|
| 21 |
enable()
|
| 22 |
+
pipe=OrbitPipeline.from_pretrained(str(ROOT),vae=vae,device='cuda',memory_budget_gib=30,vae_core_frames=512,offload_ar=offload_ar,progress=progress)
|
| 23 |
pipe.fused=not low_memory
|
| 24 |
pipe.trim_cache=True
|
| 25 |
return pipe
|
| 26 |
|
| 27 |
def main():
|
| 28 |
+
p=argparse.ArgumentParser();p.add_argument('--prompt',default=str(ROOT/'prompts/default-ru.json'));p.add_argument('--output',default='song');p.add_argument('--vae');p.add_argument('--low-memory',action='store_true');p.add_argument('--offload-ar',action='store_true');a=p.parse_args()
|
| 29 |
request=json.loads(Path(a.prompt).read_text());out=Path(a.output);out.mkdir(parents=True,exist_ok=True)
|
| 30 |
+
with load_pipeline(vae=a.vae,low_memory=a.low_memory,offload_ar=a.offload_ar) as pipe:
|
| 31 |
start=time.perf_counter()
|
| 32 |
song=pipe(**{k:request[k] for k in ('style','lyrics','seed','cot') if k in request})
|
| 33 |
song.save_artifacts(out);song.save(out/'audio.wav')
|