WaveCut commited on
Commit
bb3819f
·
verified ·
1 Parent(s): f75fe51

Document packed runtime and expose optional AR offload

Browse files
README.md CHANGED
@@ -1,9 +1,70 @@
1
  ---
 
2
  license: cc-by-nc-4.0
3
  base_model: m-a-p/YuE2-3B
4
- pipeline_tag: text-to-audio
5
- inference: false
 
 
 
 
 
 
 
 
6
  ---
7
- # YuE2-3B OrbitQuant W4A4
8
 
9
- Private release candidate. Publication and final model card follow independent clean-load verification.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ library_name: transformers
3
  license: cc-by-nc-4.0
4
  base_model: m-a-p/YuE2-3B
5
+ tags:
6
+ - music-generation
7
+ - orbitquant
8
+ - quantization
9
+ - 4-bit
10
+ - custom-code
11
+ language:
12
+ - en
13
+ - zh
14
+ - ru
15
  ---
16
+ # YuE2-3B · OrbitQuant W4A4
17
 
18
+ Packed 4-bit weights and activations for all 392 transformer projections in YuE2-3B, with a directly loadable runtime and a dedicated CUDA GEMV kernel. The default example is **«До утра»**, an original Russian synth-pop song. Russian is an experiment: the upstream model advertises Chinese and English, and this example does not establish general Russian-language reliability.
19
+
20
+ This is a **partial-model W4A4 release**. Embeddings, output heads, normalization, auxiliary projections and the separate FP32 audio VAE retain their original precision. No fine-tuning or distillation was performed.
21
+
22
+ ## Architecture
23
+
24
+ YuE2 has 28 mixture-of-transformers layers with fixed AR/NAR token-type routing, hidden width 2048, MLP width 6144, 16 query heads and 8 KV heads. Each branch has Q/K/V/O attention and gate/up/down SwiGLU projections. RMSNorm and rotary position embeddings remain unchanged. This is not learned sparse MoE routing.
25
+
26
+ The autoregressive branch plans an ABC score and generates semantic audio tokens. The non-autoregressive transformer uses flow matching with 32 midpoint integration steps to produce 64-dimensional latents at 25 Hz. The separate Oobleck-style VAE decodes these to 48 kHz stereo using convolutions, transposed convolutions and SnakeBeta activations; it is not a transformer. The upstream semantic encoder is not included in the released generation checkpoint.
27
+
28
+ ## Storage and execution
29
+
30
+ 2,818,572,288 transformer weights occupy 1,409,286,144 packed bytes. The complete model checkpoint is **3,035,926,840 bytes**, versus **7,261,441,640 bytes** upstream: 2.39× smaller (58.2% reduction). The separate 530,512,720-byte VAE is unchanged and downloaded at a pinned revision.
31
+
32
+ Compatible Q/K/V and gate/up projections are grouped: 224 physical packed modules represent 392 logical projections without duplicated packed storage. A dedicated bias-free CUDA W4A4 GEMV handles 1–8 rows; larger inputs retain OrbitQuant's packed path. No persistent decoded BF16 or INT8 weight caches are used. Original AR CUDA graphs and prefix caching were already present upstream and are not credited as new optimizations.
33
+
34
+ The default VAE core is 512 frames with the original 16-frame halo. On the recorded Russian song this was waveform-identical to 1024 frames, while lowering the decoder's isolated memory peak. This is a tested example, not a universal bitwise-equivalence guarantee.
35
+
36
+ ## Supported runtime
37
+
38
+ **Validated only on NVIDIA RTX 5090 (SM120), Linux x86_64, glibc ≥ 2.34, Python 3.12, PyTorch 2.10.0 / CUDA 12.8.** Packaged kernels use Python ABI3. Other GPUs and Torch/CUDA combinations require rebuilding and validation; the runner rejects unsupported GPU capability. This is not a generic `AutoModel.from_pretrained` integration.
39
+
40
+ ```bash
41
+ hf download WaveCut/YuE2-3B-OrbitQuant-W4A4 --local-dir YuE2-W4A4
42
+ cd YuE2-W4A4
43
+ python3.12 -m venv .venv
44
+ source .venv/bin/activate
45
+ pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cu128
46
+ pip install -r requirements.txt
47
+ python run.py --output song
48
+ ```
49
+
50
+ The output includes `audio.wav` and generation artifacts. `--prompt` accepts a JSON file containing `style`, `lyrics`, `seed`, and `cot`. `--low-memory` disables grouped projections; it is an alternative profile, not a guaranteed speedup. `--offload-ar` enables upstream AR offload: our same-song test was slower (44.33 vs 42.93 seconds) with unchanged overall peak before the VAE chunk-size change, so it is not the default. Acoustic CUDA graph capture was tested but not selected as a default due to its small benefit and additional memory.
51
+
52
+ ## Measurement protocol
53
+
54
+ Measurements use one RTX 5090 rented at $0.69/hour. Reported generation time excludes checkpoint loading unless explicitly named otherwise. First-request and warm same-process results are separate. NVML peak is device memory sampled during generation; PyTorch allocated/reserved values are different metrics. Full songs may have different semantic sequences and durations even with the same prompt and seed. The controlled 60-second acoustic comparison replays identical 1500 semantic tokens and initial noise.
55
+
56
+ See `evaluation/` for exact metrics, hashes, source pins and the clean-load proof. The published runtime was downloaded into a separate HF cache and executed with the original checkpoint directory unavailable. All packed module counts and absence of decoded caches were asserted. Model checkpoint SHA256: `e169774b8e79f423e6e8586221558acd28a35bacc67aae912d9a8a905231134a`.
57
+
58
+ ## Quality and limitations
59
+
60
+ For the controlled Russian acoustic pair, latent cosine similarity was 0.97822, waveform correlation 0.90434, and spectral convergence 0.16642. On 128 teacher-forced AR positions, mean KL was 0.03131 nats and top-1 agreement 82.81%. These are diagnostics, not perceptual scores.
61
+
62
+ Whisper-large-v3-turbo measured WER 13.87% for the original Russian song and 5.84% for the quantized song against the supplied lyrics. This is one example with different generated music and an imperfect automatic recognizer; it does not show a general quality improvement. Listen to both samples. MP3 previews are lossy encodes without loudness normalization or other postprocessing.
63
+
64
+ Earlier unoptimized packed execution was substantially slower than BF16, illustrating why packing alone is insufficient. The included kernel/runtime changes are required for the reported profile. Low-memory hardware feasibility is not established solely by the measured peak; loading and graph capture can have separate requirements.
65
+
66
+ ## Provenance and license
67
+
68
+ Source pins are in `source-lock.json`; runtime changes are in `runtime/OPTIMIZATIONS.md`; standalone kernel sources and tests are in `kernel-source/`. OrbitQuant 0.9.2 is included as a wheel built from its pinned source revision. Original licenses and third-party notices are retained.
69
+
70
+ Model weights are **CC BY-NC 4.0**, inherited from YuE2. This release does not grant commercial rights. Runtime and kernel components retain their respective code licenses. Credit M-A-P / YuE2 and OrbitQuant when using this derivative.
evaluation/ar-fused-fixed.json ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "models/YuE2-OrbitQuant-W4A4",
3
+ "gemv": true,
4
+ "fuse_packed": true,
5
+ "replay": "benchmark/ru-original-full/run-00",
6
+ "gpu": "NVIDIA GeForce RTX 5090",
7
+ "results": [
8
+ {
9
+ "repeat": 0,
10
+ "prefill_seconds": 1.1080292710103095,
11
+ "wall_seconds": 2.194309923099354,
12
+ "cuda_ms": 2194.27587890625,
13
+ "tokens": 512,
14
+ "tps": 233.33075907382553
15
+ },
16
+ {
17
+ "repeat": 1,
18
+ "prefill_seconds": 0.14538272097706795,
19
+ "wall_seconds": 2.1817037297878414,
20
+ "cuda_ms": 2181.6689453125,
21
+ "tokens": 512,
22
+ "tps": 234.67897726415362
23
+ },
24
+ {
25
+ "repeat": 2,
26
+ "prefill_seconds": 0.14446253678761423,
27
+ "wall_seconds": 2.199943969026208,
28
+ "cuda_ms": 2199.907958984375,
29
+ "tokens": 512,
30
+ "tps": 232.7332001217439
31
+ }
32
+ ]
33
+ }
evaluation/ar-gemv-fixed.json ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "models/YuE2-OrbitQuant-W4A4",
3
+ "gemv": true,
4
+ "fuse_packed": false,
5
+ "replay": "benchmark/ru-original-full/run-00",
6
+ "gpu": "NVIDIA GeForce RTX 5090",
7
+ "results": [
8
+ {
9
+ "repeat": 0,
10
+ "prefill_seconds": 1.698816369054839,
11
+ "wall_seconds": 2.863719446817413,
12
+ "cuda_ms": 2863.684326171875,
13
+ "tokens": 512,
14
+ "tps": 178.78846357278812
15
+ },
16
+ {
17
+ "repeat": 1,
18
+ "prefill_seconds": 0.17642259295098484,
19
+ "wall_seconds": 2.8788871839642525,
20
+ "cuda_ms": 2878.85400390625,
21
+ "tokens": 512,
22
+ "tps": 177.84649667826565
23
+ },
24
+ {
25
+ "repeat": 2,
26
+ "prefill_seconds": 0.17390872887335718,
27
+ "wall_seconds": 2.814904622035101,
28
+ "cuda_ms": 2814.87451171875,
29
+ "tokens": 512,
30
+ "tps": 181.88893363990346
31
+ }
32
+ ]
33
+ }
evaluation/ar-original-fixed.json ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "models/YuE2-3B",
3
+ "gemv": false,
4
+ "fuse_packed": false,
5
+ "replay": "benchmark/ru-original-full/run-00",
6
+ "gpu": "NVIDIA GeForce RTX 5090",
7
+ "results": [
8
+ {
9
+ "repeat": 0,
10
+ "prefill_seconds": 0.4288525169249624,
11
+ "wall_seconds": 2.5710276649333537,
12
+ "cuda_ms": 2570.99365234375,
13
+ "tokens": 512,
14
+ "tps": 199.14215898305866
15
+ },
16
+ {
17
+ "repeat": 1,
18
+ "prefill_seconds": 0.09946558880619705,
19
+ "wall_seconds": 2.5770173692144454,
20
+ "cuda_ms": 2576.986328125,
21
+ "tokens": 512,
22
+ "tps": 198.6792972823747
23
+ },
24
+ {
25
+ "repeat": 2,
26
+ "prefill_seconds": 0.0979860769584775,
27
+ "wall_seconds": 2.5706858979538083,
28
+ "cuda_ms": 2570.656005859375,
29
+ "tokens": 512,
30
+ "tps": 199.16863449071596
31
+ }
32
+ ]
33
+ }
evaluation/asr-source.json ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ {
2
+ "repo": "openai/whisper-large-v3-turbo",
3
+ "revision": "41f01f3fe87f28c78e2fbf8b568835947dd65ed9"
4
+ }
evaluation/harness/ar_bench.py ADDED
@@ -0,0 +1,37 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Identical prefix and forced token trace for fair AR runtime measurements."""
2
+ import argparse,json,time,gc
3
+ from pathlib import Path
4
+ import numpy as np
5
+ import torch
6
+ from yue2.modeling_yue2 import YuE2ForCausalLM
7
+ from yue2.cuda_graph import GraphAR
8
+ from yue2.protocol import CODEC_OFFSET
9
+ from yue2_orbit import load_packed
10
+
11
+ @torch.inference_mode()
12
+ def main():
13
+ p=argparse.ArgumentParser();p.add_argument('--model',default='models/YuE2-3B');p.add_argument('--replay',default='benchmark/ru-original-full/run-00');p.add_argument('--output',required=True);p.add_argument('--gemv',action='store_true');p.add_argument('--fuse-packed',action='store_true');p.add_argument('--tokens',type=int,default=512);a=p.parse_args()
14
+ if a.gemv:
15
+ from enable_gemv import enable
16
+ enable()
17
+ if (Path(a.model)/'orbitquant_manifest.json').exists():
18
+ model,manifest=load_packed(a.model)
19
+ if a.fuse_packed:
20
+ from fuse_packed import fuse_model
21
+ fuse_model(model,manifest)
22
+ else:model=YuE2ForCausalLM.from_pretrained(a.model,torch_dtype=torch.bfloat16).cuda().eval()
23
+ prefix=np.load(Path(a.replay)/'prefix.npy').tolist();tokens=np.load(Path(a.replay)/'semantic.npy').tolist()[:a.tokens]
24
+ results=[]
25
+ for repeat in range(3):
26
+ g=GraphAR(model,[prefix],len(tokens)+1,capture=True)
27
+ t=time.perf_counter();g.prefill();torch.cuda.synchronize();prefill=time.perf_counter()-t
28
+ start,end=torch.cuda.Event(enable_timing=True),torch.cuda.Event(enable_timing=True)
29
+ torch.cuda.synchronize();t=time.perf_counter();start.record()
30
+ for token in tokens:g.step(int(token)+CODEC_OFFSET)
31
+ end.record();end.synchronize();wall=time.perf_counter()-t
32
+ results.append(dict(repeat=repeat,prefill_seconds=prefill,wall_seconds=wall,cuda_ms=start.elapsed_time(end),tokens=len(tokens),tps=len(tokens)/wall))
33
+ g.close()
34
+ value=dict(model=a.model,gemv=a.gemv,fuse_packed=a.fuse_packed,replay=a.replay,gpu=torch.cuda.get_device_name(),results=results)
35
+ Path(a.output).write_text(json.dumps(value,indent=2));print(json.dumps(value),flush=True)
36
+
37
+ if __name__=='__main__':main()
evaluation/harness/ar_quality.py ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Teacher-forced next-token KL on a saved baseline sequence, never sampled paths."""
3
+ import argparse
4
+ import gc
5
+ import json
6
+ from pathlib import Path
7
+ import numpy as np
8
+ import torch
9
+ from yue2.modeling_yue2 import YuE2ForCausalLM
10
+ from yue2.protocol import CODEC_OFFSET,CODEC_SIZE,MUSIC_END
11
+ from yue2_orbit import load_packed
12
+
13
+
14
+ @torch.inference_mode()
15
+ def logits(model,prefix,tokens):
16
+ # Position 0 predicts first semantic token; every later input is from baseline.
17
+ ids=torch.tensor([prefix+[t+CODEC_OFFSET for t in tokens[:127]]],device='cuda')
18
+ result=model(ids,logits_to_keep=128,use_cache=False).logits[0].float()
19
+ return torch.cat([result[:,MUSIC_END:MUSIC_END+1],result[:,CODEC_OFFSET:CODEC_OFFSET+CODEC_SIZE]],dim=1).cpu()
20
+
21
+
22
+ @torch.inference_mode()
23
+ def main():
24
+ p=argparse.ArgumentParser(); p.add_argument('--replay',default='benchmark/original-full/run-00'); p.add_argument('--candidate',default='models/YuE2-OrbitQuant-W4A4'); p.add_argument('--output',default='reports/ar-quality.json'); p.add_argument('--activation-bits',type=int); a=p.parse_args()
25
+ root=Path(a.replay); prefix=np.load(root/'prefix.npy').tolist(); tokens=np.load(root/'semantic.npy').tolist()
26
+ if len(tokens)<128: raise ValueError('At least 128 baseline semantic tokens required')
27
+ model=YuE2ForCausalLM.from_pretrained('models/YuE2-3B',torch_dtype=torch.bfloat16,local_files_only=True).cuda().eval()
28
+ ref=logits(model,prefix,tokens); del model; gc.collect(); torch.cuda.empty_cache()
29
+ model,manifest=load_packed(a.candidate,activation_bits=a.activation_bits)
30
+ quant=logits(model,prefix,tokens)
31
+ lr=ref.log_softmax(-1); lq=quant.log_softmax(-1)
32
+ kl=(lr.exp()*(lr-lq)).sum(-1)
33
+ targets=torch.tensor(tokens[:128])+1
34
+ result=dict(candidate=a.candidate,activation_bits=a.activation_bits or manifest['config']['activation_bits'],positions=128,
35
+ teacher_forced_kl_mean=float(kl.mean()),teacher_forced_kl_max=float(kl.max()),
36
+ top1_agreement=float((ref.argmax(-1)==quant.argmax(-1)).float().mean()),
37
+ source_nll=float(-lr.gather(1,targets[:,None]).mean()),candidate_nll=float(-lq.gather(1,targets[:,None]).mean()),
38
+ interpretation='128 positions from one held baseline prefix; numerical diagnostic, not a general quality benchmark')
39
+ Path(a.output).write_text(json.dumps(result,indent=2)); print(json.dumps(result),flush=True)
40
+
41
+ if __name__=='__main__': main()
evaluation/harness/asr_evaluate.py ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Whisper lyric transcription diagnostic; no perceptual-quality claims."""
2
+ import argparse,json,re,time
3
+ from pathlib import Path
4
+ import numpy as np
5
+ import soundfile as sf
6
+ from scipy.signal import resample_poly
7
+ import torch
8
+ from transformers import AutoModelForSpeechSeq2Seq,AutoProcessor
9
+
10
+ def words(text):
11
+ text=re.sub(r'\[[^\]]*\]',' ',text).lower().replace('ё','е')
12
+ return re.findall(r'[a-zа-я0-9]+',text)
13
+
14
+ def distance(a,b):
15
+ prev=list(range(len(b)+1))
16
+ for i,x in enumerate(a,1):
17
+ row=[i]
18
+ for j,y in enumerate(b,1):row.append(min(row[-1]+1,prev[j]+1,prev[j-1]+(x!=y)))
19
+ prev=row
20
+ return prev[-1]
21
+
22
+ @torch.inference_mode()
23
+ def main():
24
+ p=argparse.ArgumentParser();p.add_argument('runs',nargs='+');p.add_argument('--output',required=True);a=p.parse_args()
25
+ model=AutoModelForSpeechSeq2Seq.from_pretrained('models/whisper-large-v3-turbo',torch_dtype=torch.float16,attn_implementation='sdpa').cuda().eval()
26
+ processor=AutoProcessor.from_pretrained('models/whisper-large-v3-turbo')
27
+ reference=words(json.loads(Path('prompts/default-ru.json').read_text())['lyrics']);report=[]
28
+ for run in a.runs:
29
+ signal,sr=sf.read(Path(run)/'audio.wav',dtype='float32');signal=signal.mean(axis=1)
30
+ from math import gcd
31
+ g=gcd(sr,16000);signal=resample_poly(signal,16000//g,sr//g)
32
+ inputs=processor(signal,sampling_rate=16000,return_tensors='pt',return_attention_mask=True,truncation=False,padding='longest')
33
+ inputs={k:v.cuda().to(torch.float16) if k=='input_features' else v.cuda() for k,v in inputs.items()}
34
+ start=time.perf_counter()
35
+ ids=model.generate(**inputs,language='ru',task='transcribe',return_timestamps=True)
36
+ text=processor.batch_decode(ids,skip_special_tokens=True)[0]
37
+ hypothesis=words(text)
38
+ item=dict(run=run,transcript=text,reference_words=len(reference),transcript_words=len(hypothesis),word_error_rate=distance(reference,hypothesis)/len(reference),seconds=time.perf_counter()-start)
39
+ report.append(item);print(json.dumps(item,ensure_ascii=False),flush=True)
40
+ Path(a.output).write_text(json.dumps(dict(results=report,interpretation='Whisper ASR is imperfect on singing; word error rate is a lyric intelligibility diagnostic, not a music-quality score.'),ensure_ascii=False,indent=2))
41
+
42
+ if __name__=='__main__':main()
evaluation/harness/benchmark.py ADDED
@@ -0,0 +1,121 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Original/packed YuE2 measurements with replayable semantic and noise inputs."""
3
+ import argparse
4
+ import hashlib
5
+ import json
6
+ import os
7
+ from pathlib import Path
8
+ import threading
9
+ import time
10
+
11
+ import numpy as np
12
+ import torch
13
+ import pynvml
14
+
15
+ from yue2 import YuE2Pipeline
16
+ from yue2.pipeline import SemanticResult, SymbolicPlan
17
+
18
+
19
+ class Monitor:
20
+ def __init__(self):
21
+ pynvml.nvmlInit()
22
+ self.handle=pynvml.nvmlDeviceGetHandleByIndex(0)
23
+ self.stop=threading.Event()
24
+ self.peak=0
25
+ self.thread=threading.Thread(target=self.run,daemon=True)
26
+ def run(self):
27
+ while not self.stop.is_set():
28
+ self.peak=max(self.peak,pynvml.nvmlDeviceGetMemoryInfo(self.handle).used)
29
+ self.stop.wait(.05)
30
+ def __enter__(self):
31
+ torch.cuda.reset_peak_memory_stats()
32
+ self.thread.start()
33
+ return self
34
+ def __exit__(self,*args):
35
+ self.stop.set(); self.thread.join()
36
+
37
+
38
+ def digest(array):
39
+ return hashlib.sha256(np.ascontiguousarray(array).tobytes()).hexdigest()
40
+
41
+
42
+ def main():
43
+ p=argparse.ArgumentParser()
44
+ p.add_argument('--model',default='models/YuE2-3B')
45
+ p.add_argument('--vae',default='models/YuE2-Vae')
46
+ p.add_argument('--prompt',default='prompts/default-ru.json')
47
+ p.add_argument('--output',required=True)
48
+ p.add_argument('--replay')
49
+ p.add_argument('--frames',type=int)
50
+ p.add_argument('--repeat',type=int,default=1)
51
+ p.add_argument('--activation-bits',type=int)
52
+ p.add_argument('--gemv',action='store_true')
53
+ p.add_argument('--fuse-packed',action='store_true')
54
+ p.add_argument('--nar-graph',action='store_true')
55
+ p.add_argument('--trim-cache',action='store_true')
56
+ p.add_argument('--offload-ar',action='store_true')
57
+ a=p.parse_args()
58
+ if a.nar_graph:
59
+ from nar_graph import enable as enable_nar
60
+ enable_nar()
61
+ if a.gemv:
62
+ from enable_gemv import enable
63
+ enable()
64
+ out=Path(a.output); out.mkdir(parents=True,exist_ok=True)
65
+ if (Path(a.model)/'orbitquant_manifest.json').exists():
66
+ from yue2_orbit import OrbitPipeline
67
+ pipe=OrbitPipeline.from_pretrained(a.model,vae=a.vae,device='cuda',memory_budget_gib=30,progress=False)
68
+ pipe.fused=a.fuse_packed
69
+ pipe.trim_cache=a.trim_cache
70
+ if a.activation_bits is not None:
71
+ pipe.activation_bits=a.activation_bits
72
+ else:
73
+ pipe=YuE2Pipeline.from_pretrained(a.model,vae=a.vae,device='cuda',memory_budget_gib=30,progress=False)
74
+ pipe.offload_ar=a.offload_ar
75
+ prompt=json.loads(Path(a.prompt).read_text())
76
+ request={k:prompt[k] for k in ('style','lyrics','seed','cot') if k in prompt}
77
+ print('loading model',flush=True)
78
+ t=time.perf_counter(); pipe._load_model(); torch.cuda.synchronize(); load=time.perf_counter()-t
79
+ for index in range(a.repeat):
80
+ dest=out/f'run-{index:02d}'; dest.mkdir(exist_ok=True)
81
+ print(f'generation {index} start replay={a.replay} frames={a.frames}',flush=True)
82
+ with Monitor() as monitor:
83
+ torch.cuda.synchronize(); start=time.perf_counter()
84
+ if a.replay:
85
+ original=Path(a.replay)
86
+ plan=SymbolicPlan.load(original)
87
+ tokens=np.load(original/'semantic.npy').tolist()
88
+ if a.frames: tokens=tokens[:a.frames]
89
+ semantic=SemanticResult(plan,tokens,{},False)
90
+ nar_start=time.perf_counter(); latents=pipe.synthesize(semantic); torch.cuda.synchronize()
91
+ nar_seconds=time.perf_counter()-nar_start
92
+ vae_start=time.perf_counter(); audio=pipe.decode(latents); torch.cuda.synchronize()
93
+ timing=dict(nar_seconds=nar_seconds,vae_seconds=time.perf_counter()-vae_start,
94
+ e2e_seconds=time.perf_counter()-start)
95
+ from yue2.pipeline import SongResult
96
+ song=SongResult(audio,48000,semantic,latents,pipe.effective_config(plan.request),pipe.weights,timing,'controlled-replay')
97
+ else:
98
+ song=pipe(**request)
99
+ torch.cuda.synchronize()
100
+ elapsed=time.perf_counter()-start
101
+ song.save_artifacts(dest)
102
+ song.save(dest/'audio.wav')
103
+ noise=torch.randn((len(song.semantic.tokens),64),generator=torch.Generator(device='cpu').manual_seed(song.semantic.plan.request.seed),dtype=torch.float32).numpy()
104
+ np.save(dest/'initial_noise.npy',noise)
105
+ metrics=dict(run=index,wall_seconds=elapsed,load_seconds=load,audio_seconds=len(song.audio)/48000,
106
+ rtf=elapsed/(len(song.audio)/48000),nvml_peak_bytes=monitor.peak,
107
+ torch_peak_allocated_bytes=torch.cuda.max_memory_allocated(),torch_peak_reserved_bytes=torch.cuda.max_memory_reserved(),
108
+ gpu=torch.cuda.get_device_name(),torch=torch.__version__,cuda=torch.version.cuda,
109
+ timing=song.timing,truncated=song.truncated,finite=bool(np.isfinite(song.audio).all()),
110
+ audio_peak=float(np.abs(song.audio).max()),audio_rms=float(np.sqrt(np.mean(song.audio**2))),
111
+ clipped_fraction=float(np.mean(np.abs(song.audio)>=1)),semantic_sha256=digest(np.asarray(song.semantic.tokens,dtype=np.int32)),
112
+ noise_sha256=digest(noise),latent_sha256=digest(song.latents),waveform_sha256=digest(song.audio),
113
+ source_or_quantized_model=str(a.model),replay=a.replay,frames=a.frames,gemv=a.gemv,fuse_packed=a.fuse_packed,nar_graph=a.nar_graph,trim_cache=a.trim_cache,offload_ar=a.offload_ar,
114
+ measurement='first process request' if index==0 else 'warm same-process request',pid=os.getpid())
115
+ if hasattr(pipe,'runtime_report'): metrics['packed_runtime']=pipe.runtime_report()
116
+ (dest/'metrics.json').write_text(json.dumps(metrics,indent=2))
117
+ print(json.dumps(metrics),flush=True)
118
+ del song
119
+ pipe.close()
120
+
121
+ if __name__=='__main__': main()
evaluation/harness/clean_validate.py ADDED
@@ -0,0 +1,59 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Run the downloaded release with the original checkpoint path unavailable."""
2
+ from pathlib import Path
3
+ import json,os,sys,time,importlib.util,hashlib
4
+ from huggingface_hub import snapshot_download
5
+ import numpy as np
6
+ import torch
7
+
8
+ root=Path('/workspace/yue2')
9
+ info=json.loads((root/'reports/candidate-hub-final-runtime.json').read_text())
10
+ folder=Path(snapshot_download(info['repo'],revision=info['revision'],cache_dir=str(root/'clean-hf-cache')))
11
+ manifest=json.loads((folder/'weights_manifest.json').read_text())
12
+ def sha(path):
13
+ h=hashlib.sha256()
14
+ with path.open('rb') as f:
15
+ for block in iter(lambda:f.read(8*1024*1024),b''):h.update(block)
16
+ return h.hexdigest()
17
+ assert sha(folder/'model.safetensors')==manifest['files']['model.safetensors']['sha256']
18
+ original=root/'models/YuE2-3B';hidden=root/'models/YuE2-3B-clean-test-unavailable'
19
+ assert original.exists() and not hidden.exists()
20
+ original.rename(hidden)
21
+ try:
22
+ spec=importlib.util.spec_from_file_location('downloaded_release',folder/'run.py')
23
+ release=importlib.util.module_from_spec(spec);spec.loader.exec_module(release)
24
+ # VAE is also resolved by the released entrypoint at its pinned Hub revision.
25
+ pipe=release.load_pipeline(progress=False)
26
+ from benchmark import Monitor,digest
27
+ request=json.loads((folder/'prompts/default-ru.json').read_text())
28
+ start=time.perf_counter();pipe._load_model();torch.cuda.synchronize();load_seconds=time.perf_counter()-start
29
+ expected=json.loads((root/'benchmark/ru-final-w4a4-full/run-01/metrics.json').read_text())
30
+ results=[]
31
+ for i in range(2):
32
+ out=root/f'benchmark/clean-release/run-{i:02d}';out.mkdir(parents=True,exist_ok=True)
33
+ with Monitor() as monitor:
34
+ torch.cuda.synchronize();start=time.perf_counter()
35
+ song=pipe(**{k:request[k] for k in ['style','lyrics','cot','seed']})
36
+ torch.cuda.synchronize();elapsed=time.perf_counter()-start
37
+ song.save_artifacts(out);song.save(out/'audio.wav')
38
+ semantic=digest(np.asarray(song.semantic.tokens,dtype=np.int32));wave=digest(song.audio)
39
+ assert semantic==expected['semantic_sha256'],'Published runtime changed semantic generation'
40
+ assert wave==expected['waveform_sha256'],'Published runtime changed waveform'
41
+ runtime=pipe.runtime_report()
42
+ assert runtime['modules']==224 and runtime['logical_projections']==392
43
+ assert runtime['decoded_weight_caches']==0 and runtime['int8_weight_caches']==0
44
+ assert runtime['effective_modes']=={'native_packed_matmul':224}
45
+ report=dict(run=i,repo=info['repo'],revision=info['revision'],model_sha256=sha(folder/'model.safetensors'),
46
+ source_checkpoint_path_unavailable=not original.exists(),module_file=sys.modules['yue2'].__file__,
47
+ wall_seconds=elapsed,load_seconds=load_seconds,audio_seconds=len(song.audio)/48000,
48
+ rtf=elapsed/(len(song.audio)/48000),nvml_peak_bytes=monitor.peak,
49
+ torch_peak_allocated_bytes=torch.cuda.max_memory_allocated(),torch_peak_reserved_bytes=torch.cuda.max_memory_reserved(),
50
+ semantic_sha256=semantic,waveform_sha256=wave,packed_runtime=runtime,config=song.config,timing=song.timing,
51
+ finite=bool(np.isfinite(song.audio).all()),truncated=song.truncated,
52
+ measurement='first process request' if i==0 else 'warm same-process request')
53
+ (out/'metrics.json').write_text(json.dumps(report,indent=2))
54
+ results.append(report);print(json.dumps(report),flush=True)
55
+ del song
56
+ pipe.close()
57
+ (root/'reports/clean-load-proof.json').write_text(json.dumps(dict(status='passed',runs=results),indent=2))
58
+ finally:
59
+ hidden.rename(original)
evaluation/harness/satellite.py ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """A one-shot <60 second remote job/liveness snapshot, never a daemon."""
3
+ import argparse
4
+ import json
5
+ from pathlib import Path
6
+ import subprocess
7
+ import time
8
+
9
+ p=argparse.ArgumentParser()
10
+ p.add_argument("job")
11
+ a=p.parse_args()
12
+ job=Path(a.job)
13
+ start=time.monotonic()
14
+ for offset in (0,15,30,45):
15
+ time.sleep(max(0,start+offset-time.monotonic()))
16
+ value=json.loads((job/'status.json').read_text()) if (job/'status.json').exists() else {"status":"missing"}
17
+ alive=False
18
+ if value.get('pid'):
19
+ result=subprocess.run(['ps','-p',str(value['pid']),'-o','pid=,stat=,etime=,pcpu=,pmem=,comm='],capture_output=True,text=True)
20
+ alive=result.returncode==0
21
+ value['process']=result.stdout.strip()
22
+ log=job/'output.log'
23
+ value['log_age_seconds']=round(time.time()-log.stat().st_mtime,1) if log.exists() else None
24
+ value['log_tail']=subprocess.run(['tail','-n','8' if value['status']!='running' else '1',str(log)],capture_output=True,text=True).stdout.strip()
25
+ value['gpu']=subprocess.run(['nvidia-smi','--query-gpu=name,utilization.gpu,memory.used,power.draw','--format=csv,noheader'],capture_output=True,text=True,timeout=5).stdout.strip()
26
+ value['alive']=alive
27
+ print(json.dumps(value),flush=True)
28
+ if value['status']!='running' or not alive:
29
+ break
evaluation/harness/supervise.py ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Detached child runner with unbuffered minute heartbeats and atomic status."""
3
+ import argparse
4
+ import json
5
+ import os
6
+ from pathlib import Path
7
+ import subprocess
8
+ import time
9
+
10
+ p = argparse.ArgumentParser()
11
+ p.add_argument("--job", required=True)
12
+ p.add_argument("command", nargs=argparse.REMAINDER)
13
+ a = p.parse_args()
14
+ job = Path(a.job)
15
+ job.mkdir(parents=True, exist_ok=True)
16
+ command = a.command[1:] if a.command[:1] == ["--"] else a.command
17
+ started = time.time()
18
+ with (job / "output.log").open("a", buffering=1) as log:
19
+ child = subprocess.Popen(command, stdout=log, stderr=subprocess.STDOUT,
20
+ env={**os.environ, "PYTHONUNBUFFERED": "1"})
21
+ def status():
22
+ code = child.poll()
23
+ value = dict(status="running" if code is None else "complete" if code == 0 else "failed",
24
+ pid=child.pid, supervisor_pid=os.getpid(), started=started,
25
+ updated=time.time(), elapsed_seconds=time.time()-started, exit_code=code)
26
+ temporary = job / "status.tmp"
27
+ temporary.write_text(json.dumps(value, indent=2))
28
+ temporary.replace(job / "status.json")
29
+ print(f"heartbeat status={value['status']} pid={child.pid} elapsed={value['elapsed_seconds']:.0f}s exit={code}", flush=True)
30
+ return code
31
+ while status() is None:
32
+ try:
33
+ child.wait(timeout=60)
34
+ except subprocess.TimeoutExpired:
35
+ pass
36
+ raise SystemExit(child.returncode)
evaluation/harness/test_fused.py ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import torch
2
+ from orbitquant import OrbitQuantConfig
3
+ from orbitquant.layers import OrbitQuantLinear
4
+ from test_packed_graph import small_model
5
+ from fuse_packed import fuse_model
6
+ from yue2.cuda_graph import GraphAR
7
+ from enable_gemv import enable
8
+
9
+ def test_fused_graph_and_storage():
10
+ enable()
11
+ model=small_model()
12
+ before=sum(m.packed_weight_indices.numel() for m in model.modules() if isinstance(m,OrbitQuantLinear))
13
+ g=GraphAR(model,[[1,2,3]],8,capture=True)
14
+ expected=[g.prefill().clone()]+[g.step(t).clone() for t in (4,5,6,7)]
15
+ g.close()
16
+ fuse_model(model,{'config':OrbitQuantConfig().to_dict()})
17
+ after=sum(m.packed_weight_indices.numel() for m in model.modules() if isinstance(m,OrbitQuantLinear))
18
+ assert before==after
19
+ g=GraphAR(model,[[1,2,3]],8,capture=True)
20
+ actual=[g.prefill().clone()]+[g.step(t).clone() for t in (4,5,6,7)]
21
+ for a,b in zip(actual,expected):torch.testing.assert_close(a,b,rtol=0,atol=0)
22
+ g.close()
23
+ # Both device transfers retain one owner and correct compatibility views.
24
+ model.cpu().cuda()
25
+ for layer in model.model.layers:
26
+ assert layer.self_attn.q_proj._group is layer.self_attn.packed_qkv
27
+ assert layer.mlp.up_proj._group is layer.mlp.packed_gate_up
evaluation/harness/test_gemv.py ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import json
2
+ from pathlib import Path
3
+ import pytest
4
+ import torch
5
+ from kernels import get_local_kernel
6
+ from orbitquant.kernels.native_packed_matmul import matmul_packed_w4a4_int8_with_native_kernel as original
7
+
8
+ kernel=get_local_kernel(Path('/workspace/yue2/orbitquant-gemv/build'), 'orbitquant_gemv')
9
+
10
+ @pytest.mark.parametrize('rows,n,k',[(1,2048,2048),(2,6144,2048),(8,2048,6144),(1,129,128),(2,1024,2048)])
11
+ @pytest.mark.parametrize('dtype',[torch.bfloat16,torch.float16])
12
+ def test_exact_integer_gemv(rows,n,k,dtype):
13
+ torch.manual_seed(123)
14
+ x=torch.randint(0,256,(rows,k//2),device='cuda',dtype=torch.uint8)
15
+ w=torch.randint(0,256,(n*k//2,),device='cuda',dtype=torch.uint8)
16
+ xn=torch.rand(rows,device='cuda');wn=torch.rand(n,device='cuda',dtype=torch.bfloat16)
17
+ ac=torch.arange(-8,8,device='cuda',dtype=torch.int8);wc=ac.flip(0).contiguous()
18
+ bias=None
19
+ kw=dict(activation_scale=.03125,weight_scale=.0625,bias=bias,output_dtype=dtype)
20
+ ref=original(x,w,xn,wn,ac,wc,out_features=n,in_features=k,**kw)
21
+ out=kernel.gemv(x,w,xn,wn,ac,wc,**kw)
22
+ torch.testing.assert_close(out,ref,rtol=0,atol=0)
23
+ stream=torch.cuda.Stream();stream.wait_stream(torch.cuda.current_stream())
24
+ with torch.cuda.stream(stream):
25
+ for _ in range(3):kernel.gemv(x,w,xn,wn,ac,wc,**kw)
26
+ torch.cuda.current_stream().wait_stream(stream)
27
+ graph=torch.cuda.CUDAGraph()
28
+ with torch.cuda.graph(graph): captured=kernel.gemv(x,w,xn,wn,ac,wc,**kw)
29
+ graph.replay();torch.cuda.synchronize()
30
+ torch.testing.assert_close(captured,ref,rtol=0,atol=0)
31
+
32
+ def test_rejects_bad_weight_shape():
33
+ x=torch.zeros(1,64,device='cuda',dtype=torch.uint8)
34
+ with pytest.raises(RuntimeError,match='packed weight size'):
35
+ kernel.gemv(x,x.flatten(),torch.ones(1,device='cuda'),torch.ones(2,device='cuda',dtype=torch.bfloat16),torch.zeros(16,device='cuda',dtype=torch.int8),torch.zeros(16,device='cuda',dtype=torch.int8),activation_scale=1,weight_scale=1)
evaluation/harness/test_packed_graph.py ADDED
@@ -0,0 +1,54 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The packed graph must match its eager path across KV updates and reload."""
2
+ import torch
3
+ from yue2.modeling_yue2 import YuE2Config, YuE2ForCausalLM
4
+ from yue2.cuda_graph import GraphAR
5
+ from orbitquant import OrbitQuantConfig
6
+ from orbitquant.layers import OrbitQuantLinear
7
+
8
+
9
+ def small_model():
10
+ torch.manual_seed(42)
11
+ config=YuE2Config(hidden_size=128,intermediate_size=256,num_hidden_layers=2,
12
+ num_attention_heads=4,num_key_value_heads=2,head_dim=32,
13
+ vocab_size=256,max_position_embeddings=128,max_latent_frames=128)
14
+ model=YuE2ForCausalLM(config).to(device='cuda',dtype=torch.bfloat16).eval()
15
+ for name,module in list(model.named_modules()):
16
+ if isinstance(module,torch.nn.Linear) and name.startswith('model.layers.'):
17
+ parent,leaf=name.rsplit('.',1)
18
+ model.get_submodule(parent)._modules[leaf]=OrbitQuantLinear.from_linear(module,config=OrbitQuantConfig(),module_name=name)
19
+ return model
20
+
21
+
22
+ def test_packed_graph_matches_eager_fixed_tokens():
23
+ model=small_model()
24
+ graph=GraphAR(model,[[1,2,3]],8,capture=True)
25
+ eager=GraphAR(model,[[1,2,3]],8,capture=False)
26
+ try:
27
+ torch.testing.assert_close(graph.prefill(),eager.prefill(),rtol=0,atol=0)
28
+ for token in (4,5,6,7):
29
+ torch.testing.assert_close(graph.step(token).clone(),eager.step(token),rtol=0,atol=0)
30
+ finally:
31
+ graph.close(); eager.close()
32
+
33
+
34
+ def test_packed_graph_rejects_bf16_fusion():
35
+ import pytest
36
+ with pytest.raises(ValueError,match='packed'):
37
+ GraphAR(small_model(),[[1,2]],4,capture=False,fuse_projections=True)
38
+
39
+
40
+ def test_packed_reload_preserves_logits(tmp_path):
41
+ import json
42
+ from safetensors.torch import save_file
43
+ from yue2_orbit import load_packed
44
+ model=small_model()
45
+ modules=[dict(name=name,in_features=m.in_features,out_features=m.out_features,bias=m.bias is not None)
46
+ for name,m in model.named_modules() if isinstance(m,OrbitQuantLinear)]
47
+ (tmp_path/'config.json').write_text(json.dumps(model.config.to_dict()))
48
+ (tmp_path/'orbitquant_manifest.json').write_text(json.dumps(dict(config=OrbitQuantConfig().to_dict(),modules=modules)))
49
+ save_file({k:v.detach().cpu().contiguous() for k,v in model.state_dict().items()},str(tmp_path/'model.safetensors'))
50
+ restored,_=load_packed(tmp_path)
51
+ ids=torch.tensor([[1,2,3]],device='cuda')
52
+ with torch.inference_mode():
53
+ torch.testing.assert_close(model(ids).logits,restored(ids).logits,rtol=0,atol=0)
54
+ assert all(m._dequantized_weight_cache is None for m in restored.modules() if isinstance(m,OrbitQuantLinear))
evaluation/harness/vae_bench.py ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Measure the existing VAE chunk-size knob independently of transformer changes."""
2
+ import json,time
3
+ from pathlib import Path
4
+ import numpy as np
5
+ import torch
6
+ import soundfile as sf
7
+ from yue2.modeling_vae import YuE2VAE
8
+ from benchmark import Monitor
9
+
10
+ @torch.inference_mode()
11
+ def main():
12
+ # Match the pinned pipeline's numerical settings, including convolution TF32.
13
+ torch.backends.cudnn.benchmark=False
14
+ torch.backends.cudnn.deterministic=True
15
+ torch.backends.cuda.matmul.allow_tf32=False
16
+ torch.backends.cudnn.allow_tf32=False
17
+ torch.backends.cuda.matmul.allow_fp16_reduced_precision_reduction=False
18
+ torch.set_float32_matmul_precision('highest')
19
+ model=YuE2VAE.from_pretrained('models/YuE2-Vae',decoder_only=True,device='cuda',local_files_only=True)
20
+ root=Path('benchmark/ru-final-w4a4-full/run-01')
21
+ z=torch.from_numpy(np.load(root/'latent.npy')).T.unsqueeze(0)
22
+ reference,sr=sf.read(root/'audio.wav',dtype='float32')
23
+ results=[]
24
+ for frames in (1024,512,256):
25
+ for repeat in range(3):
26
+ torch.cuda.empty_cache()
27
+ with Monitor() as monitor:
28
+ torch.cuda.synchronize();start=time.perf_counter()
29
+ audio=model.decode_tiled(z,core_frames=frames,halo_frames=16,output_device='cpu')[0].float().clamp(-1,1).T.contiguous().numpy()
30
+ torch.cuda.synchronize();elapsed=time.perf_counter()-start
31
+ diff=audio-reference
32
+ row=dict(pipeline_numerics=True,core_frames=frames,repeat=repeat,seconds=elapsed,nvml_peak_bytes=monitor.peak,torch_peak_allocated_bytes=torch.cuda.max_memory_allocated(),max_abs_difference=float(np.abs(diff).max()),relative_l2=float(np.linalg.norm(diff)/(np.linalg.norm(reference)+1e-12)),waveform_correlation=float(np.corrcoef(audio.ravel(),reference.ravel())[0,1]))
33
+ results.append(row);print(json.dumps(row),flush=True)
34
+ Path('reports/vae-chunk-benchmark.json').write_text(json.dumps(results,indent=2))
35
+
36
+ if __name__=='__main__': main()
evaluation/ru-ar-quality.json ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "candidate": "models/YuE2-OrbitQuant-W4A4",
3
+ "activation_bits": 4,
4
+ "positions": 128,
5
+ "teacher_forced_kl_mean": 0.031312376260757446,
6
+ "teacher_forced_kl_max": 0.12084035575389862,
7
+ "top1_agreement": 0.828125,
8
+ "source_nll": 3.8141632080078125,
9
+ "candidate_nll": 3.8379945755004883,
10
+ "interpretation": "128 positions from one held baseline prefix; numerical diagnostic, not a general quality benchmark"
11
+ }
evaluation/ru-asr.json ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results": [
3
+ {
4
+ "run": "benchmark/ru-original-full/run-01",
5
+ "transcript": " Ночной трамвай Стираю с улиц тишину И этот город слышит нас одних Пускай часы торопятся вперед Мы не считаем пройденных шагов Пока над крышами рассвет встает Мне хватит самых обыкновенных слов Останьтесь со мной до утра Пока не погасли огни, нам эта короткая ночь Подарит другие дни. Останься со мной до утра и просто за руку держи. Когда просыпается город, мы снова учимся жить. На остановке мокрая сканья И теплый ветер трогает руках Я столько раз искал тебя во снах Но ты стоишь и смотришь на меня Останься со мной до утра, пока не потасли огни Нам эта короткая ночь подарит другие дни Останься со мной до утра и просто за руку держи Когда просыпается город Пока просыпается город Останьте со мной до утра",
6
+ "reference_words": 137,
7
+ "transcript_words": 124,
8
+ "word_error_rate": 0.1386861313868613,
9
+ "seconds": 1.5337588579859585
10
+ },
11
+ {
12
+ "run": "benchmark/ru-final-w4a4-full/run-01",
13
+ "transcript": " Ночной трамвай уходят за мосты В стекле дрожат последние дни Я собираю с улиц тишину И этот город слышит нас одни Пускай часы торопятся вперед Мы не считаем пройденных шагов Пока над крышами рассвет встает Мне хватит самых обыкновенных слов Останьте со мной до утра, пока не погасли огни Нам эта короткая ночь подарит другие дни Останьте со мной до утра и просто за руку держи Когда просыпается город, мы снова учимся жить На остановке мокрая сканья И теплый ветер трогает рукав Я столько раз искал тебя во снах Но ты стоишь и смотришь на меня Останься со мной до утра, пока не погасли огни Нам эта короткая ночь подарит другие дни Останься со мной до утра и просто за руку держи Когда просыпается город, мы снова учимся жить Пока просыпается город Останьте за мной до утра",
14
+ "reference_words": 137,
15
+ "transcript_words": 137,
16
+ "word_error_rate": 0.058394160583941604,
17
+ "seconds": 0.9663123651407659
18
+ }
19
+ ],
20
+ "interpretation": "Whisper ASR is imperfect on singing; word error rate is a lyric intelligibility diagnostic, not a music-quality score."
21
+ }
evaluation/ru-final-w4a4-full/run-00/metrics.json ADDED
@@ -0,0 +1,82 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "run": 0,
3
+ "wall_seconds": 45.04905349877663,
4
+ "load_seconds": 1.029712200164795,
5
+ "audio_seconds": 163.83866666666665,
6
+ "rtf": 0.2749598395501467,
7
+ "nvml_peak_bytes": 9756803072,
8
+ "torch_peak_allocated_bytes": 5495350272,
9
+ "torch_peak_reserved_bytes": 8495562752,
10
+ "gpu": "NVIDIA GeForce RTX 5090",
11
+ "torch": "2.10.0+cu128",
12
+ "cuda": "12.8",
13
+ "timing": {
14
+ "abc": {
15
+ "seconds": 10.528724248055369,
16
+ "prefill_seconds": 0.8370891781523824,
17
+ "ttft_seconds": 0.9959756580647081,
18
+ "output_tokens": 1772,
19
+ "content_tokens": 1771,
20
+ "output_tps": 168.30149201858757,
21
+ "prefix_tokens": 423,
22
+ "cfg_branches": 1,
23
+ "execution": "cuda_graph",
24
+ "attention": "flash"
25
+ },
26
+ "semantic": {
27
+ "seconds": 23.20693418988958,
28
+ "prefill_seconds": 0.3349543798249215,
29
+ "ttft_seconds": 0.33607721398584545,
30
+ "output_tokens": 4097,
31
+ "content_tokens": 4096,
32
+ "output_tps": 176.54206137167893,
33
+ "prefix_tokens": 2196,
34
+ "cfg_branches": 1,
35
+ "execution": "cuda_graph",
36
+ "attention": "flash"
37
+ },
38
+ "nar_seconds": 7.658700210042298,
39
+ "vae_seconds": 3.6495362292043865,
40
+ "load": {
41
+ "resolve_and_integrity_seconds": 2.4894183888100088,
42
+ "mot_load_seconds": 1.0296437400393188
43
+ },
44
+ "e2e_seconds": 45.04794106609188
45
+ },
46
+ "truncated": {
47
+ "abc": false,
48
+ "semantic": false
49
+ },
50
+ "finite": true,
51
+ "audio_peak": 0.9706120491027832,
52
+ "audio_rms": 0.14440932869911194,
53
+ "clipped_fraction": 0.0,
54
+ "semantic_sha256": "45d407d6c04fb19e54f29a17b622a89060240e5b8d8af4b0f3eee1fc338e5251",
55
+ "noise_sha256": "0f1e184cfce6b99c4e2b2f5dd27a470c894fc6a1dcbe6bd0cab6ec635d343f29",
56
+ "latent_sha256": "e371f6a61d2d2d13a9ff1003aeae8bf6355a62569c7707988c49f0053bd1e76f",
57
+ "waveform_sha256": "eceb3ad260035470803e5827bc88192819a5ca08f01536dd3d580abb4b190d54",
58
+ "source_or_quantized_model": "models/YuE2-OrbitQuant-W4A4",
59
+ "replay": null,
60
+ "frames": null,
61
+ "gemv": true,
62
+ "fuse_packed": true,
63
+ "nar_graph": false,
64
+ "trim_cache": true,
65
+ "measurement": "first process request",
66
+ "pid": 10795,
67
+ "packed_runtime": {
68
+ "modules": 224,
69
+ "logical_projections": 392,
70
+ "fused": true,
71
+ "effective_modes": {
72
+ "native_packed_matmul": 224
73
+ },
74
+ "activation_backends": {
75
+ "native_cuda_int8_surrogate": 168,
76
+ "triton_cuda_packed_w4": 56
77
+ },
78
+ "decoded_weight_caches": 0,
79
+ "int8_weight_caches": 0,
80
+ "packed_weight_bytes": 1409286144
81
+ }
82
+ }
evaluation/ru-final-w4a4-full/run-01/metrics.json ADDED
@@ -0,0 +1,82 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "run": 1,
3
+ "wall_seconds": 42.934636591002345,
4
+ "load_seconds": 1.029712200164795,
5
+ "audio_seconds": 163.83866666666665,
6
+ "rtf": 0.2620543578907035,
7
+ "nvml_peak_bytes": 9857466368,
8
+ "torch_peak_allocated_bytes": 5782598144,
9
+ "torch_peak_reserved_bytes": 8596226048,
10
+ "gpu": "NVIDIA GeForce RTX 5090",
11
+ "torch": "2.10.0+cu128",
12
+ "cuda": "12.8",
13
+ "timing": {
14
+ "abc": {
15
+ "seconds": 9.711376916151494,
16
+ "prefill_seconds": 0.15963919600471854,
17
+ "ttft_seconds": 0.16055710101500154,
18
+ "output_tokens": 1772,
19
+ "content_tokens": 1771,
20
+ "output_tps": 182.46640155145198,
21
+ "prefix_tokens": 423,
22
+ "cfg_branches": 1,
23
+ "execution": "cuda_graph",
24
+ "attention": "flash"
25
+ },
26
+ "semantic": {
27
+ "seconds": 23.016886444063857,
28
+ "prefill_seconds": 0.15774337714537978,
29
+ "ttft_seconds": 0.1586433621123433,
30
+ "output_tokens": 4097,
31
+ "content_tokens": 4096,
32
+ "output_tps": 177.9997485740141,
33
+ "prefix_tokens": 2196,
34
+ "cfg_branches": 1,
35
+ "execution": "cuda_graph",
36
+ "attention": "flash"
37
+ },
38
+ "nar_seconds": 7.616418123943731,
39
+ "vae_seconds": 2.2683653261046857,
40
+ "load": {
41
+ "resolve_and_integrity_seconds": 2.4894183888100088,
42
+ "mot_load_seconds": 1.0296437400393188
43
+ },
44
+ "e2e_seconds": 42.93380210502073
45
+ },
46
+ "truncated": {
47
+ "abc": false,
48
+ "semantic": false
49
+ },
50
+ "finite": true,
51
+ "audio_peak": 0.9706120491027832,
52
+ "audio_rms": 0.14440932869911194,
53
+ "clipped_fraction": 0.0,
54
+ "semantic_sha256": "45d407d6c04fb19e54f29a17b622a89060240e5b8d8af4b0f3eee1fc338e5251",
55
+ "noise_sha256": "0f1e184cfce6b99c4e2b2f5dd27a470c894fc6a1dcbe6bd0cab6ec635d343f29",
56
+ "latent_sha256": "e371f6a61d2d2d13a9ff1003aeae8bf6355a62569c7707988c49f0053bd1e76f",
57
+ "waveform_sha256": "eceb3ad260035470803e5827bc88192819a5ca08f01536dd3d580abb4b190d54",
58
+ "source_or_quantized_model": "models/YuE2-OrbitQuant-W4A4",
59
+ "replay": null,
60
+ "frames": null,
61
+ "gemv": true,
62
+ "fuse_packed": true,
63
+ "nar_graph": false,
64
+ "trim_cache": true,
65
+ "measurement": "warm same-process request",
66
+ "pid": 10795,
67
+ "packed_runtime": {
68
+ "modules": 224,
69
+ "logical_projections": 392,
70
+ "fused": true,
71
+ "effective_modes": {
72
+ "native_packed_matmul": 224
73
+ },
74
+ "activation_backends": {
75
+ "native_cuda_int8_surrogate": 168,
76
+ "triton_cuda_packed_w4": 56
77
+ },
78
+ "decoded_weight_caches": 0,
79
+ "int8_weight_caches": 0,
80
+ "packed_weight_bytes": 1409286144
81
+ }
82
+ }
evaluation/ru-offload-w4a4-full/run-00/metrics.json ADDED
@@ -0,0 +1,83 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "run": 0,
3
+ "wall_seconds": 46.23842817801051,
4
+ "load_seconds": 1.14848616393283,
5
+ "audio_seconds": 163.83866666666665,
6
+ "rtf": 0.2822192655661902,
7
+ "nvml_peak_bytes": 9756803072,
8
+ "torch_peak_allocated_bytes": 5495350272,
9
+ "torch_peak_reserved_bytes": 8495562752,
10
+ "gpu": "NVIDIA GeForce RTX 5090",
11
+ "torch": "2.10.0+cu128",
12
+ "cuda": "12.8",
13
+ "timing": {
14
+ "abc": {
15
+ "seconds": 10.56912293494679,
16
+ "prefill_seconds": 0.8523490249644965,
17
+ "ttft_seconds": 1.038547317031771,
18
+ "output_tokens": 1772,
19
+ "content_tokens": 1771,
20
+ "output_tps": 167.65818799787866,
21
+ "prefix_tokens": 423,
22
+ "cfg_branches": 1,
23
+ "execution": "cuda_graph",
24
+ "attention": "flash"
25
+ },
26
+ "semantic": {
27
+ "seconds": 23.106563664972782,
28
+ "prefill_seconds": 0.28357866406440735,
29
+ "ttft_seconds": 0.2846065699122846,
30
+ "output_tokens": 4097,
31
+ "content_tokens": 4096,
32
+ "output_tps": 177.30892656317556,
33
+ "prefix_tokens": 2196,
34
+ "cfg_branches": 1,
35
+ "execution": "cuda_graph",
36
+ "attention": "flash"
37
+ },
38
+ "nar_seconds": 9.069145092042163,
39
+ "vae_seconds": 3.4875931590795517,
40
+ "load": {
41
+ "resolve_and_integrity_seconds": 2.487326374044642,
42
+ "mot_load_seconds": 1.1484210339840502
43
+ },
44
+ "e2e_seconds": 46.23729805299081
45
+ },
46
+ "truncated": {
47
+ "abc": false,
48
+ "semantic": false
49
+ },
50
+ "finite": true,
51
+ "audio_peak": 0.9706120491027832,
52
+ "audio_rms": 0.14440932869911194,
53
+ "clipped_fraction": 0.0,
54
+ "semantic_sha256": "45d407d6c04fb19e54f29a17b622a89060240e5b8d8af4b0f3eee1fc338e5251",
55
+ "noise_sha256": "0f1e184cfce6b99c4e2b2f5dd27a470c894fc6a1dcbe6bd0cab6ec635d343f29",
56
+ "latent_sha256": "e371f6a61d2d2d13a9ff1003aeae8bf6355a62569c7707988c49f0053bd1e76f",
57
+ "waveform_sha256": "eceb3ad260035470803e5827bc88192819a5ca08f01536dd3d580abb4b190d54",
58
+ "source_or_quantized_model": "models/YuE2-OrbitQuant-W4A4",
59
+ "replay": null,
60
+ "frames": null,
61
+ "gemv": true,
62
+ "fuse_packed": true,
63
+ "nar_graph": false,
64
+ "trim_cache": true,
65
+ "offload_ar": true,
66
+ "measurement": "first process request",
67
+ "pid": 11579,
68
+ "packed_runtime": {
69
+ "modules": 224,
70
+ "logical_projections": 392,
71
+ "fused": true,
72
+ "effective_modes": {
73
+ "native_packed_matmul": 224
74
+ },
75
+ "activation_backends": {
76
+ "native_cuda_int8_surrogate": 168,
77
+ "triton_cuda_packed_w4": 56
78
+ },
79
+ "decoded_weight_caches": 0,
80
+ "int8_weight_caches": 0,
81
+ "packed_weight_bytes": 1409286144
82
+ }
83
+ }
evaluation/ru-offload-w4a4-full/run-01/metrics.json ADDED
@@ -0,0 +1,83 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "run": 1,
3
+ "wall_seconds": 44.32611982291564,
4
+ "load_seconds": 1.14848616393283,
5
+ "audio_seconds": 163.83866666666665,
6
+ "rtf": 0.27054736665489415,
7
+ "nvml_peak_bytes": 9857466368,
8
+ "torch_peak_allocated_bytes": 5782598144,
9
+ "torch_peak_reserved_bytes": 8596226048,
10
+ "gpu": "NVIDIA GeForce RTX 5090",
11
+ "torch": "2.10.0+cu128",
12
+ "cuda": "12.8",
13
+ "timing": {
14
+ "abc": {
15
+ "seconds": 9.711860368028283,
16
+ "prefill_seconds": 0.15409997710958123,
17
+ "ttft_seconds": 0.15508739417418838,
18
+ "output_tokens": 1772,
19
+ "content_tokens": 1771,
20
+ "output_tps": 182.4573184591362,
21
+ "prefix_tokens": 423,
22
+ "cfg_branches": 1,
23
+ "execution": "cuda_graph",
24
+ "attention": "flash"
25
+ },
26
+ "semantic": {
27
+ "seconds": 23.071978967869654,
28
+ "prefill_seconds": 0.15794157283380628,
29
+ "ttft_seconds": 0.15890546003356576,
30
+ "output_tokens": 4097,
31
+ "content_tokens": 4096,
32
+ "output_tps": 177.57471111193092,
33
+ "prefix_tokens": 2196,
34
+ "cfg_branches": 1,
35
+ "execution": "cuda_graph",
36
+ "attention": "flash"
37
+ },
38
+ "nar_seconds": 8.86763684405014,
39
+ "vae_seconds": 2.319226206978783,
40
+ "load": {
41
+ "resolve_and_integrity_seconds": 2.487326374044642,
42
+ "mot_load_seconds": 1.1484210339840502
43
+ },
44
+ "e2e_seconds": 44.32531815697439
45
+ },
46
+ "truncated": {
47
+ "abc": false,
48
+ "semantic": false
49
+ },
50
+ "finite": true,
51
+ "audio_peak": 0.9706120491027832,
52
+ "audio_rms": 0.14440932869911194,
53
+ "clipped_fraction": 0.0,
54
+ "semantic_sha256": "45d407d6c04fb19e54f29a17b622a89060240e5b8d8af4b0f3eee1fc338e5251",
55
+ "noise_sha256": "0f1e184cfce6b99c4e2b2f5dd27a470c894fc6a1dcbe6bd0cab6ec635d343f29",
56
+ "latent_sha256": "e371f6a61d2d2d13a9ff1003aeae8bf6355a62569c7707988c49f0053bd1e76f",
57
+ "waveform_sha256": "eceb3ad260035470803e5827bc88192819a5ca08f01536dd3d580abb4b190d54",
58
+ "source_or_quantized_model": "models/YuE2-OrbitQuant-W4A4",
59
+ "replay": null,
60
+ "frames": null,
61
+ "gemv": true,
62
+ "fuse_packed": true,
63
+ "nar_graph": false,
64
+ "trim_cache": true,
65
+ "offload_ar": true,
66
+ "measurement": "warm same-process request",
67
+ "pid": 11579,
68
+ "packed_runtime": {
69
+ "modules": 224,
70
+ "logical_projections": 392,
71
+ "fused": true,
72
+ "effective_modes": {
73
+ "native_packed_matmul": 224
74
+ },
75
+ "activation_backends": {
76
+ "native_cuda_int8_surrogate": 168,
77
+ "triton_cuda_packed_w4": 56
78
+ },
79
+ "decoded_weight_caches": 0,
80
+ "int8_weight_caches": 0,
81
+ "packed_weight_bytes": 1409286144
82
+ }
83
+ }
evaluation/ru-original-full/run-00/metrics.json ADDED
@@ -0,0 +1,64 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "run": 0,
3
+ "wall_seconds": 51.78475882997736,
4
+ "load_seconds": 3.197218642104417,
5
+ "audio_seconds": 164.63866666666667,
6
+ "rtf": 0.3145358248972135,
7
+ "nvml_peak_bytes": 10583080960,
8
+ "torch_peak_allocated_bytes": 8692496896,
9
+ "torch_peak_reserved_bytes": 9332326400,
10
+ "gpu": "NVIDIA GeForce RTX 5090",
11
+ "torch": "2.10.0+cu128",
12
+ "cuda": "12.8",
13
+ "timing": {
14
+ "abc": {
15
+ "seconds": 10.894125336082652,
16
+ "prefill_seconds": 0.48003934998996556,
17
+ "ttft_seconds": 0.6661999989300966,
18
+ "output_tokens": 1669,
19
+ "content_tokens": 1668,
20
+ "output_tps": 153.20183571526127,
21
+ "prefix_tokens": 423,
22
+ "cfg_branches": 1,
23
+ "execution": "cuda_graph",
24
+ "attention": "flash"
25
+ },
26
+ "semantic": {
27
+ "seconds": 26.13979333708994,
28
+ "prefill_seconds": 0.11473918193951249,
29
+ "ttft_seconds": 0.11574821802787483,
30
+ "output_tokens": 4117,
31
+ "content_tokens": 4116,
32
+ "output_tps": 157.49933241279146,
33
+ "prefix_tokens": 2093,
34
+ "cfg_branches": 1,
35
+ "execution": "cuda_graph",
36
+ "attention": "flash"
37
+ },
38
+ "nar_seconds": 8.817161875078455,
39
+ "vae_seconds": 5.922186557902023,
40
+ "load": {
41
+ "resolve_and_integrity_seconds": 4.985906707821414,
42
+ "mot_load_seconds": 0.2366083210799843
43
+ },
44
+ "e2e_seconds": 51.78419973095879
45
+ },
46
+ "truncated": {
47
+ "abc": false,
48
+ "semantic": false
49
+ },
50
+ "finite": true,
51
+ "audio_peak": 1.0,
52
+ "audio_rms": 0.15364815294742584,
53
+ "clipped_fraction": 1.3286672227666243e-06,
54
+ "semantic_sha256": "f37633f5006e9b80ae72a0321d663e1623e939d5963dd9493a1b592126db122e",
55
+ "noise_sha256": "c8aa36bd1c94b4d8d0f311f823e2e948ddc876c9951d98bc8a47e8c5b5bffb1c",
56
+ "latent_sha256": "752ebcaea4ffae13bfeb169655d6c84381b1db186871feffc0c0105670f16b73",
57
+ "waveform_sha256": "a8a6ea78c6b5dc4478b4e54d538625282ee87f3a10b4127323cc737f881e0338",
58
+ "source_or_quantized_model": "models/YuE2-3B",
59
+ "replay": null,
60
+ "frames": null,
61
+ "gemv": false,
62
+ "measurement": "first process request",
63
+ "pid": 5444
64
+ }
evaluation/ru-original-full/run-01/metrics.json ADDED
@@ -0,0 +1,64 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "run": 1,
3
+ "wall_seconds": 50.16780533106066,
4
+ "load_seconds": 3.197218642104417,
5
+ "audio_seconds": 164.63866666666667,
6
+ "rtf": 0.3047145992297921,
7
+ "nvml_peak_bytes": 10908139520,
8
+ "torch_peak_allocated_bytes": 8977385472,
9
+ "torch_peak_reserved_bytes": 9646899200,
10
+ "gpu": "NVIDIA GeForce RTX 5090",
11
+ "torch": "2.10.0+cu128",
12
+ "cuda": "12.8",
13
+ "timing": {
14
+ "abc": {
15
+ "seconds": 10.320648127002642,
16
+ "prefill_seconds": 0.08523392281495035,
17
+ "ttft_seconds": 0.0860717708710581,
18
+ "output_tokens": 1669,
19
+ "content_tokens": 1668,
20
+ "output_tps": 161.714650035716,
21
+ "prefix_tokens": 423,
22
+ "cfg_branches": 1,
23
+ "execution": "cuda_graph",
24
+ "attention": "flash"
25
+ },
26
+ "semantic": {
27
+ "seconds": 26.154879739973694,
28
+ "prefill_seconds": 0.11253222916275263,
29
+ "ttft_seconds": 0.11347067612223327,
30
+ "output_tokens": 4117,
31
+ "content_tokens": 4116,
32
+ "output_tps": 157.40848518251076,
33
+ "prefix_tokens": 2093,
34
+ "cfg_branches": 1,
35
+ "execution": "cuda_graph",
36
+ "attention": "flash"
37
+ },
38
+ "nar_seconds": 8.794583374867216,
39
+ "vae_seconds": 4.3119754288345575,
40
+ "load": {
41
+ "resolve_and_integrity_seconds": 4.985906707821414,
42
+ "mot_load_seconds": 0.2366083210799843
43
+ },
44
+ "e2e_seconds": 50.16745935194194
45
+ },
46
+ "truncated": {
47
+ "abc": false,
48
+ "semantic": false
49
+ },
50
+ "finite": true,
51
+ "audio_peak": 1.0,
52
+ "audio_rms": 0.15364815294742584,
53
+ "clipped_fraction": 1.3286672227666243e-06,
54
+ "semantic_sha256": "f37633f5006e9b80ae72a0321d663e1623e939d5963dd9493a1b592126db122e",
55
+ "noise_sha256": "c8aa36bd1c94b4d8d0f311f823e2e948ddc876c9951d98bc8a47e8c5b5bffb1c",
56
+ "latent_sha256": "752ebcaea4ffae13bfeb169655d6c84381b1db186871feffc0c0105670f16b73",
57
+ "waveform_sha256": "a8a6ea78c6b5dc4478b4e54d538625282ee87f3a10b4127323cc737f881e0338",
58
+ "source_or_quantized_model": "models/YuE2-3B",
59
+ "replay": null,
60
+ "frames": null,
61
+ "gemv": false,
62
+ "measurement": "warm same-process request",
63
+ "pid": 5444
64
+ }
evaluation/ru-original-replay1500/run-00/metrics.json ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "run": 0,
3
+ "wall_seconds": 9.195442344993353,
4
+ "load_seconds": 3.375524572096765,
5
+ "audio_seconds": 59.998666666666665,
6
+ "rtf": 0.15326077821162726,
7
+ "nvml_peak_bytes": 9165406208,
8
+ "torch_peak_allocated_bytes": 7884326912,
9
+ "torch_peak_reserved_bytes": 8009023488,
10
+ "gpu": "NVIDIA GeForce RTX 5090",
11
+ "torch": "2.10.0+cu128",
12
+ "cuda": "12.8",
13
+ "timing": {
14
+ "nar_seconds": 3.627831295831129,
15
+ "vae_seconds": 5.565822640201077,
16
+ "e2e_seconds": 9.195178817026317
17
+ },
18
+ "truncated": {
19
+ "abc": false,
20
+ "semantic": false
21
+ },
22
+ "finite": true,
23
+ "audio_peak": 0.8817682862281799,
24
+ "audio_rms": 0.13392212986946106,
25
+ "clipped_fraction": 0.0,
26
+ "semantic_sha256": "105e53ffbbe671c28587ce52c917f57698087ed47bf6b63792a377e357236c6b",
27
+ "noise_sha256": "304490172c3b67d1dd3984d61d39e3cfc0d311ad6df79200674f815bc1081eae",
28
+ "latent_sha256": "b9461dadf0c0a16941fa0178ea5f3845f32d1f6bda70de4fe17541b2e81faabf",
29
+ "waveform_sha256": "9a80ae8267c22cd066aaf7cd1f7e5a12c7fcaedfb3a2e0b5f80b95ee2a47e482",
30
+ "source_or_quantized_model": "models/YuE2-3B",
31
+ "replay": "benchmark/ru-original-full/run-00",
32
+ "frames": 1500,
33
+ "gemv": false,
34
+ "measurement": "first process request",
35
+ "pid": 6449
36
+ }
evaluation/ru-original-replay1500/run-01/metrics.json ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "run": 1,
3
+ "wall_seconds": 7.551422053948045,
4
+ "load_seconds": 3.375524572096765,
5
+ "audio_seconds": 59.998666666666665,
6
+ "rtf": 0.12585983111760338,
7
+ "nvml_peak_bytes": 9469493248,
8
+ "torch_peak_allocated_bytes": 8150668800,
9
+ "torch_peak_reserved_bytes": 8304721920,
10
+ "gpu": "NVIDIA GeForce RTX 5090",
11
+ "torch": "2.10.0+cu128",
12
+ "cuda": "12.8",
13
+ "timing": {
14
+ "nar_seconds": 4.007696468848735,
15
+ "vae_seconds": 3.542188974097371,
16
+ "e2e_seconds": 7.551166985882446
17
+ },
18
+ "truncated": {
19
+ "abc": false,
20
+ "semantic": false
21
+ },
22
+ "finite": true,
23
+ "audio_peak": 0.8817682862281799,
24
+ "audio_rms": 0.13392212986946106,
25
+ "clipped_fraction": 0.0,
26
+ "semantic_sha256": "105e53ffbbe671c28587ce52c917f57698087ed47bf6b63792a377e357236c6b",
27
+ "noise_sha256": "304490172c3b67d1dd3984d61d39e3cfc0d311ad6df79200674f815bc1081eae",
28
+ "latent_sha256": "b9461dadf0c0a16941fa0178ea5f3845f32d1f6bda70de4fe17541b2e81faabf",
29
+ "waveform_sha256": "9a80ae8267c22cd066aaf7cd1f7e5a12c7fcaedfb3a2e0b5f80b95ee2a47e482",
30
+ "source_or_quantized_model": "models/YuE2-3B",
31
+ "replay": "benchmark/ru-original-full/run-00",
32
+ "frames": 1500,
33
+ "gemv": false,
34
+ "measurement": "warm same-process request",
35
+ "pid": 6449
36
+ }
evaluation/vae-chunk-benchmark.json ADDED
@@ -0,0 +1,101 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "pipeline_numerics": true,
4
+ "core_frames": 1024,
5
+ "repeat": 0,
6
+ "seconds": 1.1320553519763052,
7
+ "nvml_peak_bytes": 8204910592,
8
+ "torch_peak_allocated_bytes": 2871368192,
9
+ "max_abs_difference": 0.0,
10
+ "relative_l2": 0.0,
11
+ "waveform_correlation": 1.0
12
+ },
13
+ {
14
+ "pipeline_numerics": true,
15
+ "core_frames": 1024,
16
+ "repeat": 1,
17
+ "seconds": 0.9049297058954835,
18
+ "nvml_peak_bytes": 8286699520,
19
+ "torch_peak_allocated_bytes": 2870057472,
20
+ "max_abs_difference": 0.0,
21
+ "relative_l2": 0.0,
22
+ "waveform_correlation": 1.0
23
+ },
24
+ {
25
+ "pipeline_numerics": true,
26
+ "core_frames": 1024,
27
+ "repeat": 2,
28
+ "seconds": 0.9072453379631042,
29
+ "nvml_peak_bytes": 8286699520,
30
+ "torch_peak_allocated_bytes": 2870581760,
31
+ "max_abs_difference": 0.0,
32
+ "relative_l2": 0.0,
33
+ "waveform_correlation": 1.0
34
+ },
35
+ {
36
+ "pipeline_numerics": true,
37
+ "core_frames": 512,
38
+ "repeat": 0,
39
+ "seconds": 0.9228434218093753,
40
+ "nvml_peak_bytes": 5040308224,
41
+ "torch_peak_allocated_bytes": 1737859584,
42
+ "max_abs_difference": 0.0,
43
+ "relative_l2": 0.0,
44
+ "waveform_correlation": 1.0
45
+ },
46
+ {
47
+ "pipeline_numerics": true,
48
+ "core_frames": 512,
49
+ "repeat": 1,
50
+ "seconds": 0.9203053209930658,
51
+ "nvml_peak_bytes": 5010948096,
52
+ "torch_peak_allocated_bytes": 1737193984,
53
+ "max_abs_difference": 0.0,
54
+ "relative_l2": 0.0,
55
+ "waveform_correlation": 1.0
56
+ },
57
+ {
58
+ "pipeline_numerics": true,
59
+ "core_frames": 512,
60
+ "repeat": 2,
61
+ "seconds": 0.9190280558541417,
62
+ "nvml_peak_bytes": 5021433856,
63
+ "torch_peak_allocated_bytes": 1738250752,
64
+ "max_abs_difference": 0.0,
65
+ "relative_l2": 0.0,
66
+ "waveform_correlation": 1.0
67
+ },
68
+ {
69
+ "pipeline_numerics": true,
70
+ "core_frames": 256,
71
+ "repeat": 0,
72
+ "seconds": 1.0967565891332924,
73
+ "nvml_peak_bytes": 3282894848,
74
+ "torch_peak_allocated_bytes": 1172341248,
75
+ "max_abs_difference": 3.874301910400391e-06,
76
+ "relative_l2": 1.3670459111381206e-06,
77
+ "waveform_correlation": 0.9999999999990744
78
+ },
79
+ {
80
+ "pipeline_numerics": true,
81
+ "core_frames": 256,
82
+ "repeat": 1,
83
+ "seconds": 1.0739446009974927,
84
+ "nvml_peak_bytes": 3282894848,
85
+ "torch_peak_allocated_bytes": 1172406784,
86
+ "max_abs_difference": 3.874301910400391e-06,
87
+ "relative_l2": 1.3670459111381206e-06,
88
+ "waveform_correlation": 0.9999999999990744
89
+ },
90
+ {
91
+ "pipeline_numerics": true,
92
+ "core_frames": 256,
93
+ "repeat": 2,
94
+ "seconds": 1.1024859179742634,
95
+ "nvml_peak_bytes": 3316449280,
96
+ "torch_peak_allocated_bytes": 1172470272,
97
+ "max_abs_difference": 3.874301910400391e-06,
98
+ "relative_l2": 1.3670459111381206e-06,
99
+ "waveform_correlation": 0.9999999999990744
100
+ }
101
+ ]
run.py CHANGED
@@ -6,7 +6,7 @@ ROOT=Path(__file__).absolute().parent
6
  sys.path.insert(0,str(ROOT/'runtime/kernels/base'))
7
  sys.path.insert(0,str(ROOT/'runtime'))
8
 
9
- def load_pipeline(*,vae=None,low_memory=False,progress=True):
10
  import torch
11
  if not torch.cuda.is_available():raise RuntimeError('This release requires an NVIDIA CUDA GPU.')
12
  if torch.cuda.get_device_capability()!=(12,0):
@@ -19,15 +19,15 @@ def load_pipeline(*,vae=None,low_memory=False,progress=True):
19
  from enable_gemv import enable
20
  from yue2_orbit import OrbitPipeline
21
  enable()
22
- pipe=OrbitPipeline.from_pretrained(str(ROOT),vae=vae,device='cuda',memory_budget_gib=30,vae_core_frames=512,progress=progress)
23
  pipe.fused=not low_memory
24
  pipe.trim_cache=True
25
  return pipe
26
 
27
  def main():
28
- p=argparse.ArgumentParser();p.add_argument('--prompt',default=str(ROOT/'prompts/default-ru.json'));p.add_argument('--output',default='song');p.add_argument('--vae');p.add_argument('--low-memory',action='store_true');a=p.parse_args()
29
  request=json.loads(Path(a.prompt).read_text());out=Path(a.output);out.mkdir(parents=True,exist_ok=True)
30
- with load_pipeline(vae=a.vae,low_memory=a.low_memory) as pipe:
31
  start=time.perf_counter()
32
  song=pipe(**{k:request[k] for k in ('style','lyrics','seed','cot') if k in request})
33
  song.save_artifacts(out);song.save(out/'audio.wav')
 
6
  sys.path.insert(0,str(ROOT/'runtime/kernels/base'))
7
  sys.path.insert(0,str(ROOT/'runtime'))
8
 
9
+ def load_pipeline(*,vae=None,low_memory=False,offload_ar=False,progress=True):
10
  import torch
11
  if not torch.cuda.is_available():raise RuntimeError('This release requires an NVIDIA CUDA GPU.')
12
  if torch.cuda.get_device_capability()!=(12,0):
 
19
  from enable_gemv import enable
20
  from yue2_orbit import OrbitPipeline
21
  enable()
22
+ pipe=OrbitPipeline.from_pretrained(str(ROOT),vae=vae,device='cuda',memory_budget_gib=30,vae_core_frames=512,offload_ar=offload_ar,progress=progress)
23
  pipe.fused=not low_memory
24
  pipe.trim_cache=True
25
  return pipe
26
 
27
  def main():
28
+ p=argparse.ArgumentParser();p.add_argument('--prompt',default=str(ROOT/'prompts/default-ru.json'));p.add_argument('--output',default='song');p.add_argument('--vae');p.add_argument('--low-memory',action='store_true');p.add_argument('--offload-ar',action='store_true');a=p.parse_args()
29
  request=json.loads(Path(a.prompt).read_text());out=Path(a.output);out.mkdir(parents=True,exist_ok=True)
30
+ with load_pipeline(vae=a.vae,low_memory=a.low_memory,offload_ar=a.offload_ar) as pipe:
31
  start=time.perf_counter()
32
  song=pipe(**{k:request[k] for k in ('style','lyrics','seed','cot') if k in request})
33
  song.save_artifacts(out);song.save(out/'audio.wav')