t2i-adapter-depth-zoe-sdxl-1.0 Open Source Model - Providing Depth Condition Control for StableDiffusionXL

T2i Adapter Depth Zoe Sdxl 1.0

Developed by TencentARC

T2I Adapter is a network that provides additional conditional control for stable diffusion models. This checkpoint offers depth condition control for the StableDiffusionXL model.

Image Generation OtherOpen Source License:Apache-2.0 #Depth Condition Control #High-Resolution Image Generation #Zoe Depth Estimation

Downloads 3,663

Release Time : 9/3/2023

Model Overview

A T2I Adapter trained based on Zoe depth estimation, designed to provide depth condition control for the Stable Diffusion XL model, generating images that conform to depth structures.

Model Features

Depth Condition Control

Provides precise depth information control through Zoe depth estimation, ensuring generated images conform to the input depth structure.

Lightweight Adapter

A lightweight adapter network with only 79M parameters, efficiently compatible with the 2.6B-parameter SDXL model.

High-Quality Generation

Trained on the LAION-Aesthetics V2 dataset, supporting high-resolution image generation.

Model Capabilities

Depth-Conditioned Image Generation

Text-to-Image Conversion

High-Resolution Image Generation

Image Structure Control

Use Cases

Creative Design

Product Concept Design

Generate product concept images based on depth maps and text descriptions.

Achieves creative designs while preserving the original depth structure.

Scene Reconstruction

Reconstruct 2D renderings of 3D scenes based on depth maps.

Accurately maintains the spatial relationships of scenes.

Artistic Creation

Depth-Guided Art Creation

Artists can control depth maps to guide AI in generating artworks with specific structures.

Achieves precise integration of artistic creativity and AI generation.

🚀 T2I-Adapter-SDXL - Depth-Zoe

T2I Adapter is a network that provides additional conditioning for Stable Diffusion. Each T2I checkpoint takes a different type of conditioning as input and is used with a specific base Stable Diffusion checkpoint. This checkpoint offers depth conditioning for the StableDiffusionXL checkpoint. It is a collaborative effort between Tencent ARC and Hugging Face.

🚀 Quick Start

To get started, first install the required dependencies:

pip install -U git+https://github.com/huggingface/diffusers.git
pip install -U controlnet_aux==0.0.7 timm==0.6.12 # for conditioning models and detectors
pip install transformers accelerate safetensors

Download images in the appropriate control image format.
Pass the control image and prompt to the StableDiffusionXLAdapterPipeline.

✨ Features

Additional Conditioning: T2I Adapter provides extra conditioning to Stable Diffusion, enabling more controllable image generation.
Depth Conditioning: This checkpoint offers depth conditioning for the StableDiffusionXL checkpoint.
Multiple Checkpoints: There are multiple checkpoints available, each taking a different type of conditioning as input.

📦 Installation

The installation steps are included in the Quick Start section.

💻 Usage Examples

Basic Usage

from diffusers import StableDiffusionXLAdapterPipeline, T2IAdapter, EulerAncestralDiscreteScheduler, AutoencoderKL
from diffusers.utils import load_image, make_image_grid
from controlnet_aux import ZoeDetector
import torch

# load adapter
adapter = T2IAdapter.from_pretrained(
  "TencentARC/t2i-adapter-depth-zoe-sdxl-1.0", torch_dtype=torch.float16, varient="fp16"
).to("cuda")

# load euler_a scheduler
model_id = 'stabilityai/stable-diffusion-xl-base-1.0'
euler_a = EulerAncestralDiscreteScheduler.from_pretrained(model_id, subfolder="scheduler")
vae=AutoencoderKL.from_pretrained("madebyollin/sdxl-vae-fp16-fix", torch_dtype=torch.float16)
pipe = StableDiffusionXLAdapterPipeline.from_pretrained(
    model_id, vae=vae, adapter=adapter, scheduler=euler_a, torch_dtype=torch.float16, variant="fp16", 
).to("cuda")
pipe.enable_xformers_memory_efficient_attention()

zoe_depth = ZoeDetector.from_pretrained(
    "valhalla/t2iadapter-aux-models", filename="zoed_nk.pth", model_type="zoedepth_nk"
).to("cuda")

Advanced Usage

# Condition Image
url = "https://huggingface.co/Adapter/t2iadapter/resolve/main/figs_SDXLV1.0/org_zeo.jpg"
image = load_image(url)
image = zoe_depth(image, gamma_corrected=True, detect_resolution=512, image_resolution=1024)

# Generation
prompt = "A photo of a orchid, 4k photo, highly detailed"
negative_prompt = "anime, cartoon, graphic, text, painting, crayon, graphite, abstract, glitch, deformed, mutated, ugly, disfigured"

gen_images = pipe(
  prompt=prompt,
  negative_prompt=negative_prompt,
  image=image,
  num_inference_steps=30,
  adapter_conditioning_scale=1,
  guidance_scale=7.5,  
).images[0]
gen_images.save('out_zoe.png')

📚 Documentation

Model Details

Property	Details
Developed by	T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
Model Type	Diffusion-based text-to-image generation model
Language(s)	English
License	Apache 2.0
Resources for more information	GitHub Repository, Paper.
Model complexity
Cite as	@misc{ title={T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models}, author={Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, Ying Shan, Xiaohu Qie}, year={2023}, eprint={2302.08453}, archivePrefix={arXiv}, primaryClass={cs.CV} }

Checkpoints

Model Name	Control Image Overview	Control Image Example	Generated Image Example
TencentARC/t2i-adapter-canny-sdxl-1.0 Trained with canny edge detection	A monochrome image with white edges on a black background.
TencentARC/t2i-adapter-sketch-sdxl-1.0 Trained with PidiNet edge detection	A hand-drawn monochrome image with white outlines on a black background.
TencentARC/t2i-adapter-lineart-sdxl-1.0 Trained with lineart edge detection	A hand-drawn monochrome image with white outlines on a black background.
TencentARC/t2i-adapter-depth-midas-sdxl-1.0 Trained with Midas depth estimation	A grayscale image with black representing deep areas and white representing shallow areas.
TencentARC/t2i-adapter-depth-zoe-sdxl-1.0 Trained with Zoe depth estimation	A grayscale image with black representing deep areas and white representing shallow areas.
TencentARC/t2i-adapter-openpose-sdxl-1.0 Trained with OpenPose bone image	A OpenPose bone image.

Training

Our training script was built on top of the official training script that we provide here. The model is trained on 3M high-resolution image-text pairs from LAION-Aesthetics V2 with:

Training steps: 25000
Batch size: Data parallel with a single gpu batch size of 16 for a total batch size of 256.
Learning rate: Constant learning rate of 1e-5.
Mixed precision: fp16

📄 License

This model is licensed under the Apache 2.0 license.

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご