See our collection for all Stable Diffusion 1.x checkpoints.

Run Stable Diffusion 1.x with Keras 3: JAX, PyTorch, or TensorFlow

GitHub Docs HuggingFace

zeromodels/stable-diffusion-v1-5

Paper: High-Resolution Image Synthesis with Latent Diffusion Models (arXiv:2112.10752) | HF Papers

Pure-Keras 3 conversion of stable-diffusion-v1-5/stable-diffusion-v1-5 for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX. The whole text-to-image model ships as one container: the UNet denoiser, the VAE and the CLIP ViT-L/14 text encoder in model.weights.h5 (1.07B parameters, 3.97 GB), plus zm_config.json (the three component configs, the checkpoint's PNDMScheduler schedule with its epsilon objective and the default generation settings) and the tokenizer as tokenizer.json. Weights are stored in float32, exactly as released. This checkpoint generates 512x512 images (a 64x64 latent).

For model details, intended use and limitations, see the upstream model card.

Architecture

Component zeromodels class Details
Denoiser UNet2DConditionModel (320, 640, 1280, 1280) channels, 2 ResNet blocks per level, 8 attention heads on the 768-d text context, 1x1 conv token projection, 64x64x4 latent
Autoencoder AutoencoderKL (128, 256, 512, 512) channels, x8 spatial compression to 4 latent channels, scaling_factor 0.18215
Text encoder CLIPTextModel CLIP ViT-L/14 text encoder: 768-d, 12 layers, 12 heads, 77 tokens, quick_gelu
Scheduler PNDMScheduler scaled_linear betas 0.00085 to 0.012 over 1000 steps, epsilon; DDIM / PNDM / Euler / Euler-ancestral are drop-in

Quick start

import os
os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from PIL import Image
from zeromodels.models.stable_diffusion import StableDiffusionTextToImage, StableDiffusionTokenizer

model = StableDiffusionTextToImage.from_weights("zeromodels/stable-diffusion-v1-5")
tokenizer = StableDiffusionTokenizer.from_weights("zeromodels/stable-diffusion-v1-5")

inputs = tokenizer("a photograph of an astronaut riding a horse")
images = model.generate(**inputs, num_inference_steps=50, guidance_scale=7.5, seed=0)
Image.fromarray(images[0]).save("astronaut.png")  # (512, 512, 3) uint8

generate takes the tokenizer's input_ids (batch them for several prompts), an optional negative_input_ids (tokenize the negative prompt), num_inference_steps, guidance_scale, a seed, or explicit latents of shape (batch, 64, 64, 4) for results that are identical across backends.

Load any Stable Diffusion 1.x checkpoint the same way with from_weights("zeromodels/<variant>"):

Variant Hub Training
stable-diffusion-v1-1 zeromodels/stable-diffusion-v1-1 237k steps at 256px on laion2B-en, then 194k steps at 512px on laion-high-resolution
stable-diffusion-v1-2 zeromodels/stable-diffusion-v1-2 v1-1 + 515k steps at 512px on laion-aesthetics v2 5+
stable-diffusion-v1-3 zeromodels/stable-diffusion-v1-3 v1-2 + 195k steps at 512px, 10% text-conditioning dropout (classifier-free guidance)
stable-diffusion-v1-4 zeromodels/stable-diffusion-v1-4 v1-2 + 225k steps at 512px, 10% text-conditioning dropout (classifier-free guidance)
stable-diffusion-v1-5 zeromodels/stable-diffusion-v1-5 v1-2 + 595k steps at 512px, 10% text-conditioning dropout (classifier-free guidance)

Tips

  • Set KERAS_BACKEND before importing Keras / zeromodels.

  • The graphs are built for 512px. Pass unet_sample_size=<px / 8>, vae_sample_size=<px> to from_weights to build for another multiple of 64px (the weights are resolution-independent).

  • Swap the sampler any time: model.scheduler = EulerDiscreteScheduler.from_config(model.config.scheduler_config) (zeromodels.base.base_scheduler).

  • StableDiffusionModel.from_weights(...) loads the same repo as the bare container (UNet / VAE / text encoder as .unet / .vae / .text_encoder) without the generation loop.

  • Both channels_last and channels_first are supported (keras.config.set_image_data_format before loading); generate always returns (batch, H, W, 3) uint8.

  • On-the-fly hf: conversion is not supported for diffusion models; the checkpoints are hosted here, converted once.

  • See the Stable Diffusion 1.x docs.

License

The weights are redistributed under the CreativeML OpenRAIL-M License of the upstream checkpoint, including its use-based restrictions. By using them you agree to those terms.

Notice

Modifications by zeromodels (https://github.com/IMvision12/ZeroModels): the checkpoint released at https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5 was converted to the Keras 3 weights layout of zeromodels (model.weights.h5, zm_config.json, tokenizer.json), stored in float32 as released. The model architecture and the parameter values are unchanged; the weight names and the file format differ from the release.

Special Thanks

Thank you to the CompVis group at LMU Munich, Runway and Stability AI for training and releasing Stable Diffusion, and to the Hugging Face diffusers team, whose implementation this port was verified against.

Downloads last month
1,528
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zeromodels/stable-diffusion-v1-5

Finetuned
(407)
this model

Collection including zeromodels/stable-diffusion-v1-5

Paper for zeromodels/stable-diffusion-v1-5