October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
ThatPainter
AI art

How to Get Started With Stable Diffusion 3 Medium

A practical beginner’s guide to running the original Stable Diffusion 3 Medium locally or through a hosted option, with setup, hardware, prompting, and current API caveats.

By ThatPainter Team Updated 9 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ThatPainter is reader-supported. When you buy through links on our site, we may earn an affiliate commission. Learn More

Stable Diffusion 3 Medium (SD3 Medium) is an open-weight text-to-image model released by Stability AI on June 12, 2024. You can run its original weights locally with ComfyUI or Python’s Diffusers library, or use a hosted image-generation service. One important update: SD3 Medium is no longer Stability AI’s newest model family, and the company’s API documentation says SD3.0 API requests are rerouted to SD3.5. If you specifically need the original SD3 Medium model, use its downloadable weights rather than assuming an API call will return it.

Choose how you want to use SD3 Medium

SD3 Medium is a model, not a complete desktop art program. You need an interface or code library to load it, plus either suitable local hardware or a hosted service.

Route Best for Main trade-off
Hosted image-generation service Trying image generation without installing software or managing a GPU Service availability, model selection, usage limits, and pricing can change.
ComfyUI Artists who want a visual, reusable local workflow Its node-based interface is flexible but takes some learning.
Hugging Face Diffusers Developers, automation, and reproducible Python workflows You must manage Python packages, model access, and hardware compatibility.
Stability AI API Applications that need hosted, programmatic image generation Current SD3.0 API calls are rerouted to SD3.5, not the original SD3 Medium weights.

If you only want to experiment, start with a hosted interface. For local control, ComfyUI is a practical visual starting point; choose Diffusers if you are comfortable writing Python. Stability AI’s API getting-started guide explains account setup and API use. Check current service terms and pricing before committing: launch-era trials and product offers may no longer apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What SD3 Medium is—and what “Medium” means

SD3 Medium is a general-purpose text-to-image model from Stability AI. “Medium” identifies its place in the SD3 family and its roughly 2-billion-parameter scale; it does not describe image resolution. Its Multimodal Diffusion Transformer (MMDiT) architecture uses three text encoders: OpenCLIP-ViT/G, CLIP-ViT/L, and T5-XXL. Stability AI and the model card highlight improved prompt comprehension, composition, and text rendering compared with earlier generations. Those are strengths, not guarantees: a generated sign can still contain misspellings or awkward layout. See the model card and release announcement.

#1 Best Overall
Sale
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
  • AI Performance: 767 AI TOPS
  • OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode)
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • A 2.5-slot design maximizes compatibility and cooling efficiency for superior performance in small chassis

The weights are downloadable, but “open-weight” is more precise than implying unrestricted open-source use. Access is gated on Hugging Face, and use is subject to the applicable license and acceptable-use terms.

Run it locally with ComfyUI

ComfyUI is a node-based interface for connecting model components and image-generation steps. The official SD3 Medium repository recommends it for local or self-hosted inference and provides example workflows for basic text-to-image generation, multi-prompt generation, and upscaling.

  1. Install ComfyUI using its official project instructions or a trusted distribution. Follow the instructions for your operating system and GPU.
  2. Accept access terms. Sign in to Hugging Face, open the SD3 Medium model page, and accept the access conditions and license. A download or workflow will fail if your account has not been granted access.
  3. Choose a checkpoint package. The repository offers variants with different text encoders. sd3_medium.safetensors contains the core MMDiT and VAE weights but not the text encoders. sd3_medium_incl_clips.safetensors includes the CLIP encoders but not T5-XXL. The sd3_medium_incl_clips_t5xxlfp8.safetensors and sd3_medium_incl_clips_t5xxlfp16.safetensors variants include T5 in FP8 and FP16 respectively. Match the file to the workflow and your available memory; do not assume the smallest file is self-contained.
  4. Put model files where your chosen workflow expects them. Import an official SD3 Medium example workflow and follow its current node and file-placement instructions. ComfyUI menus, custom nodes, and folder conventions can change, so use the workflow’s accompanying instructions rather than relying on old screenshots.
  5. Enter a prompt and queue the workflow. Start with a basic text-to-image example at a moderate resolution. If ComfyUI reports missing encoders, nodes, or model files, check that the selected checkpoint and workflow agree before installing unrelated extensions.

ComfyUI offers control over generation and follow-up steps, but it is not a one-prompt-box application. Keep the first workflow simple; add upscaling or other processing only after a basic generation works.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Run it with Python and Diffusers

Diffusers is a Python library for loading and running diffusion pipelines. The steps below create an isolated environment, authenticate for the gated model, and generate one 1024 × 1024 image. Install a compatible PyTorch build separately using the official instructions for your operating system and CUDA or ROCm setup.

1. Create an environment and install the libraries

python -m venv .venv
source .venv/bin/activate
pip install --upgrade diffusers transformers accelerate safetensors

On Windows PowerShell, activate the environment with:

Rank #2
GIGABYTE GeForce RTX 4070 WINDFORCE OC 12G Graphics Card, 3X WINDFORCE Fans, 12GB 192-bit GDDR6X, GV-N4070WF3OC-12GD Video Card
  • Powered by NVIDIA DLSS 3, ultra-efficient Ada Lovelace architechture, and full ray tracing
  • 4th Generation Tensor Cores: Up to 4x performance with DLSS 3
  • 3rd Generation RT Cores: Up to 2x ray tracing performance
  • Powered by GeForce RTX 4070
  • Integrated with 12GB GDDR6X 192-bit memory interface
python -m venv .venv
.venvScriptsActivate.ps1

Diffusers and its dependencies change over time; use the current SD3 pipeline documentation if installation errors indicate a version mismatch.

2. Accept the model gate and authenticate

Sign in to Hugging Face and accept the conditions on the SD3 Medium model page. Then authenticate in the terminal:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
hf auth login

Older tutorials may show huggingface-cli login. If a download is denied, confirm you accepted the gate while signed into the same account used by the local token.

3. Generate a first image

import torch
from diffusers import StableDiffusion3Pipeline

model_id = "stabilityai/stable-diffusion-3-medium-diffusers"

pipe = StableDiffusion3Pipeline.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")

image = pipe(
    prompt="A cat holding a sign that says hello world",
    negative_prompt="",
    num_inference_steps=28,
    height=1024,
    width=1024,
    guidance_scale=7.0,
).images[0]

image.save("sd3_medium_first_image.png")

The pipeline downloads the model on first use, then saves the generated file as sd3_medium_first_image.png in the working directory. The 28 steps, 1024 × 1024 size, and guidance scale of 7.0 are documented starting values, not guaranteed best settings for every prompt or system.

Hardware and memory: plan for the text encoders

SD3 Medium is smaller than SD3.5 Large, but it can still be demanding compared with older Stable Diffusion checkpoints. Diffusers documentation notes that running the full pipeline in FP16 can be difficult on GPUs with less than 24 GB of VRAM, largely because of its three text encoders, including the 4.7-billion-parameter T5-XXL encoder. This is not a universal minimum: actual use depends on precision, resolution, batch size, whether T5 is loaded, offloading, GPU architecture, and other software using memory.

Rank #3
ASUS Dual GeForce RTX 4070 Super EVO OC Edition 12GB GDDR6X (PCIe 4.0, 12GB GDDR6X, DLSS 3, HDMI 2.1a, DisplayPort 1.4a, 2.5-Slot Design, Axial-tech Fan Design, 0dB Technology), 3 Year Warranty
  • Powered by NVIDIA DLSS3, ultra-efficient Ada Lovelace arch, and full ray tracing
  • 4th Generation Tensor Cores: Up to 4x performance with DLSS 3 vs. brute-force rendering
  • 3rd Generation RT Cores: Up to 2x ray tracing performance
  • OC edition: Boost Clock 2550 MHz (OC Mode)/ 2520 MHz (Default Mode)
  • Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • Full pipeline: preserves all three text encoders but has the greatest memory demand.
  • CPU offloading: moves components between CPU and GPU to reduce GPU pressure, usually at the cost of slower generation. In Diffusers, replace pipe.to("cuda") with pipe.enable_model_cpu_offload() after loading the pipeline.
  • Omit T5: set text_encoder_3=None and tokenizer_3=None when loading the pipeline. This lowers memory use but can weaken prompt understanding, particularly for detailed or complex descriptions.
  • Quantize T5: Diffusers documents lower-memory options such as 8-bit quantization with bitsandbytes. This is an advanced route; compatibility varies with operating system, GPU, and software stack.

For an aggressive low-memory attempt, load the pipeline without T5 as shown in the current Diffusers documentation, or use CPU offloading. These techniques can make a run possible, but they do not promise a particular speed or fit on every GPU. A hosted service avoids local VRAM management, though it adds service dependency and potentially usage costs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Write prompts that give the model useful direction

SD3 Medium does not require a special prompt syntax. A clear natural-language description is a good starting point:

[subject] + [action or pose] + [environment] + [lighting] +
[composition] + [medium or visual style] + [specific text, if needed]

For example:

A red fox reading a newspaper at a rainy café window, three-quarter view, warm tungsten light, shallow depth of field, editorial illustration, muted teal and orange palette, the newspaper headline clearly reads “GOOD MORNING”

Put the main subject and action early. Describe spatial relationships explicitly—such as “a small blue cup beside a larger white plate”—and add framing, angle, lighting, materials, or palette when they matter. For signs and labels, write the desired wording plainly. Generate several outputs or seeds before deciding that a prompt is failing; a single image is not a reliable test of a prompt or model.

SD3 Medium was designed to improve typography, but it cannot guarantee exact spelling or layout. Inspect all generated text. For a logo, poster, label, or other business-critical design, plan to correct or typeset the wording in an editing or design tool.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
QTHREE GeForce GT 730 4GB Graphics Card,2X HDMI, DP,VGA,DDR3,64 Bit,Low Profile Video Card for PC,Computer GPU,PCI Express X8,SFF,DirectX 12,Support Winows 11
  • NVIDIA GT 730 graphics cards offer basic display capabilities for office work and light multimedia,which with 1000 MHz Memory Clock 4GB DDR3 on Kepler architecture, support multiple monitors and HD video playback,easily upgrading for convenient usage to save your budget for your old pc
  • The low-profile design of the PC graphics card saves installation space, easy to install,plug &play,making it easy to build a compact computer system, even compatible with ITX chassis.
  • The 4x outputs enables multi-monitor productivity on up to 4 monitors simultaneously,including 2x HDMI,VGA,DP.Designed for full-size chassis and small case installations.
  • PCI Express based PC is required with one X8 lane graphics slot available on the motherboard. 300 Watt or greater power supply. This video card can automatically install new drivers and support Win11,DirectX 12.
  • 30W low power,no external power supply and the all-solid-state capacitor keeps low power consumption and high performance.If you have any problems about this card,please contact us via amazon messages.

SD3 Medium, SD3.5, and the Stability API

SD3 Medium was a new release in June 2024, not the current newest Stability model family. Stability AI later released SD3.5 models. Its current API documentation says SD3.0 APIs were deprecated on April 17, 2025, and requests are automatically rerouted to SD3.5 models at no extra cost. In other words, an API request associated with SD3.0 should not be treated as a way to reproduce the original SD3 Medium checkpoint.

Choose the original SD3 Medium weights when compatibility with a specific checkpoint, tutorial, workflow, or reproducibility matters. Consider SD3.5 if you want the newer family and current Stability API support, but do not assume every newer model is automatically the better fit: hardware, workflow compatibility, speed, and the output you need all matter. SD3.5 Medium, SD3.5 Large, and SD3.5 Large Turbo are distinct options, with different model sizes and generation behavior.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Licensing and commercial use

The Hugging Face model card describes the Stability Community License as allowing commercial use for individuals or organizations with annual revenue below US$1 million. Entities above that threshold need to review Stability AI’s Enterprise licensing requirements when using its models in commercial products or services. Read the current Stability AI license, the model’s accompanying terms, and the acceptable-use policy information before relying on the model for business use.

Accepting the Hugging Face gate is not a waiver of license obligations. Revenue, enterprise use, a product embedding the model, API-provider use, and derivative models can raise different questions. The model license also does not settle copyright, trademark, publicity-rights, or platform-policy questions about a particular image. Check the terms that apply to your actual use and seek legal advice for consequential commercial decisions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Local weights and hosted services may also have different operational safeguards. Do not assume that a self-hosted workflow applies exactly the same moderation or filtering as an online service.

Best Value
ASUS Dual GeForce RTX 4070 OC Edition 12GB GDDR6X, IP5X, Auto-Extreme Technology, 144-Hour Validation Program, HDMI 2.1a, DP 1.4a, 3 Year Warranty
  • Powered by NVIDIA DLSS3, ultra-efficient Ada Lovelace architecture, and full ray tracing.
  • 4th Generation Tensor Cores: Up to 4x performance with DLSS 3 vs. brute force rendering
  • 3rd Generation RT Cores: Up to 2x ray tracing performance
  • OC mode: 2505 MHz / Default Mode: 2475 MHz
  • Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure.

Troubleshooting

“Access denied” or download failure

Check that you accepted the gate on the correct Hugging Face account and authenticated locally with that account. Verify your login with hf auth whoami; if necessary, log in again with hf auth login. Confirm the repository identifier is exactly stabilityai/stable-diffusion-3-medium-diffusers for the Diffusers example.

CUDA out of memory

Reduce the batch size to one, use FP16 where supported, close other GPU-heavy applications, and try CPU offloading. If that is still too demanding, omit T5 or try a documented quantized T5 option; lowering resolution may also help. The text encoders can account for substantial memory use, so reducing image size alone may not solve the problem. Restarting the Python process can help after repeated failed runs or fragmented memory.

Missing encoders, distorted images, or a failed ComfyUI workflow

Check that the checkpoint package includes—or is paired with—the text encoders expected by the workflow. With Diffusers, start from the official model identifier and documented pipeline. With ComfyUI, re-import the official example and verify that required nodes and model files are present. Re-download an incomplete or corrupted file, and test without LoRAs, custom VAEs, ControlNets, or extensions before adding them back one at a time.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generation works but is very slow

CPU offloading trades speed for lower GPU memory use. A low-memory GPU, a text encoder running on the CPU, first-run initialization, or suboptimal attention support can also slow a generation. A successful run is not necessarily a practical-speed setup; if it is consistently too slow, use a hosted service or a machine better suited to the workload.

The API output does not look like SD3 Medium

This is expected if you are using Stability AI’s current API: SD3.0 API calls are rerouted to SD3.5 according to the current documentation. Use the gated local SD3 Medium weights when the original model is essential.

Quick Recap

SaleBestseller No. 1
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
AI Performance: 767 AI TOPS; OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode); Powered by the NVIDIA Blackwell architecture and DLSS 4
$790.99
Bestseller No. 2
GIGABYTE GeForce RTX 4070 WINDFORCE OC 12G Graphics Card, 3X WINDFORCE Fans, 12GB 192-bit GDDR6X, GV-N4070WF3OC-12GD Video Card
GIGABYTE GeForce RTX 4070 WINDFORCE OC 12G Graphics Card, 3X WINDFORCE Fans, 12GB 192-bit GDDR6X, GV-N4070WF3OC-12GD Video Card
Powered by NVIDIA DLSS 3, ultra-efficient Ada Lovelace architechture, and full ray tracing
$839.00
Bestseller No. 3
ASUS Dual GeForce RTX 4070 Super EVO OC Edition 12GB GDDR6X (PCIe 4.0, 12GB GDDR6X, DLSS 3, HDMI 2.1a, DisplayPort 1.4a, 2.5-Slot Design, Axial-tech Fan Design, 0dB Technology), 3 Year Warranty
ASUS Dual GeForce RTX 4070 Super EVO OC Edition 12GB GDDR6X (PCIe 4.0, 12GB GDDR6X, DLSS 3, HDMI 2.1a, DisplayPort 1.4a, 2.5-Slot Design, Axial-tech Fan Design, 0dB Technology), 3 Year Warranty
Powered by NVIDIA DLSS3, ultra-efficient Ada Lovelace arch, and full ray tracing; 4th Generation Tensor Cores: Up to 4x performance with DLSS 3 vs. brute-force rendering
$839.22
Bestseller No. 5
ASUS Dual GeForce RTX 4070 OC Edition 12GB GDDR6X, IP5X, Auto-Extreme Technology, 144-Hour Validation Program, HDMI 2.1a, DP 1.4a, 3 Year Warranty
ASUS Dual GeForce RTX 4070 OC Edition 12GB GDDR6X, IP5X, Auto-Extreme Technology, 144-Hour Validation Program, HDMI 2.1a, DP 1.4a, 3 Year Warranty
Powered by NVIDIA DLSS3, ultra-efficient Ada Lovelace architecture, and full ray tracing.; 4th Generation Tensor Cores: Up to 4x performance with DLSS 3 vs. brute force rendering
$1,195.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

More from the Paint Desk

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.