thoddnn/parakeet-tdt-0.6b-v3-mlx-4bit

This is a stability mirror of animaslabs/parakeet-tdt-0.6b-v3-mlx-4bit at source revision 65247a0a9e735426eba06056a9535f7e67dcbbb9. The model weights and configuration are unchanged. The original conversion and its attribution are preserved below.

This model was converted to MLX format, 4-bit quantized from nvidia/parakeet-tdt-0.6b-v3 using the scripts in this github repo. Please refer to original model card for more details on the model.

Usage

Quantized models require calling mlx.nn.quantize() before loading weights.

import json
import mlx.nn as nn
from huggingface_hub import hf_hub_download
from parakeet_mlx.utils import from_config

# Download and load config
config_path = hf_hub_download("thoddnn/parakeet-tdt-0.6b-v3-mlx-4bit", "config.json")
with open(config_path) as f:
    config = json.load(f)

# Build model and apply quantization structure
model = from_config(config)
nn.quantize(
    model,
    bits=config["quantization"]["bits"],
    group_size=config["quantization"]["group_size"],
)

# Load quantized weights
weights_path = hf_hub_download("thoddnn/parakeet-tdt-0.6b-v3-mlx-4bit", "model.safetensors")
model.load_weights(weights_path)

# Transcribe
result = model.transcribe("audio.wav")
print(result.text)
Downloads last month
40
Safetensors
Model size
0.6B params
Tensor type
F32
·
U32
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thoddnn/parakeet-tdt-0.6b-v3-mlx-4bit

Quantized
(96)
this model

Dataset used to train thoddnn/parakeet-tdt-0.6b-v3-mlx-4bit