This repository has been archived on 2026-08-19. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
ROCm-Research-Archive/_TestScripts/Danis ROCm Kernel Patch Research/ComfyUI_ZImage_Documentation.md
T
2026-08-20 00:45:43 +02:00

638 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ComfyUI + Z-Image Turbo on AMD BC-250 — Complete Setup Guide
> **Hardware**: AMD BC-250 (Cyan Skillfish, gfx1013→gfx1010, 24 CUs, shared RAM)
> **Backend**: ROCm 7.2.0 / PyTorch 2.5.1+rocm6.2
> **OS**: CachyOS, kernel 6.18.8-3-cachyos
> **ComfyUI Version**: 0.15.1
> **Date**: 2026-03-02
---
## Table of Contents
1. [Overview](#1-overview)
2. [Architecture](#2-architecture)
3. [Prerequisites](#3-prerequisites)
4. [Installation — Step by Step](#4-installation--step-by-step)
5. [Model Setup](#5-model-setup)
6. [Launch Script](#6-launch-script)
7. [ComfyUI Workflow — Z-Image Turbo](#7-comfyui-workflow--z-image-turbo)
8. [BC-250 Specific Tuning](#8-bc-250-specific-tuning)
9. [Troubleshooting](#9-troubleshooting)
10. [File Inventory](#10-file-inventory)
11. [Performance Notes](#11-performance-notes)
---
## 1. Overview
ComfyUI is a node-based Stable Diffusion GUI that runs Z-Image Turbo (a Lumina2-architecture model) via PyTorch with ROCm/HIP on the AMD BC-250 GPU. The model uses a GGUF-quantized diffusion model (Q5_K_S) loaded via the ComfyUI-GGUF custom node, with a Gemma 2 2B text encoder and a Flux-compatible VAE.
### What's Running
| Component | File | Size | Format |
|-----------|------|------|--------|
| Diffusion Model | `z_image_turbo-Q5_K_S.gguf` | 5.2 GB | GGUF Q5_K_S |
| Text Encoder | `gemma2_2b_lumina2.safetensors` | 9.8 GB | Safetensors (f32) |
| VAE | `ae.safetensors` | 335 MB | Safetensors (f32) |
### Pipeline
```
[ComfyUI WebUI :8188] → [PyTorch] → [ROCm/HIP] → [AMD BC-250 GPU]
[Gemma 2 2B Text Encoder] → CLIP Encode → [Z-Image Turbo Diffusion] → [VAE Decode] → Image
```
---
## 2. Architecture
### Z-Image Turbo Details
- **Architecture**: Lumina2 (Lumina-Image 2.0 family)
- **Base**: Z-Image by Freepik, turbo-distilled variant
- **Text Encoder**: Gemma 2 2B (Google, 2304-dim embeddings)
- **VAE**: Flux-compatible autoencoder (`ae.safetensors`)
- **Sampler**: Euler with SGM Uniform scheduler, 8 steps (turbo)
- **CFG Scale**: 3.0 (turbo models use low CFG)
- **Latent Format**: Flux-style latent space
### Why GGUF?
The BC-250 has ~14.7 GB shared system RAM. The full FP16 diffusion model would be too large. GGUF Q5_K_S quantization reduces the model from ~12+ GB to 5.2 GB, making it feasible alongside the text encoder and VAE.
---
## 3. Prerequisites
Before starting, you need ROCm working on the BC-250. See `ROCm_BC250_Documentation.md` for the full ROCm setup.
### Required
- ROCm 7.2.0 installed and working (`rocminfo` detects BC-250)
- Python 3.11 (`/usr/bin/python3.11`)
- Git
- ~30 GB free disk space
### Verify ROCm
```bash
rocminfo | grep "Name:"
# Should show: gfx1010 and AMD BC-250
```
---
## 4. Installation — Step by Step
### 4.1 Clone ComfyUI
```bash
cd ~
git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUI
```
### 4.2 Create Python 3.11 Virtual Environment
Python 3.11 is required — Python 3.14 (system default) is too new for PyTorch ROCm wheels.
```bash
python3.11 -m venv venv
source venv/bin/activate
```
### 4.3 Install PyTorch with ROCm Support
```bash
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/rocm6.2
```
This downloads ~4 GB. The ROCm 6.2 PyTorch wheel is compatible with the ROCm 7.2 runtime.
**Verify installation:**
```bash
python -c "import torch; print(torch.version.cuda); print(torch.cuda.is_available())"
# Should print: 6.2 and True
```
### 4.4 Install ComfyUI Dependencies
```bash
pip install -r requirements.txt
```
### 4.5 Install ComfyUI-GGUF Custom Node
This enables loading GGUF-quantized models in ComfyUI.
```bash
cd ~/ComfyUI/custom_nodes
git clone https://github.com/city96/ComfyUI-GGUF.git
source ~/ComfyUI/venv/bin/activate
pip install gguf
```
### 4.6 Install huggingface-hub (for model downloads)
```bash
pip install huggingface-hub
```
---
## 5. Model Setup
### 5.1 Directory Structure
ComfyUI looks for models in `~/ComfyUI/models/`. Our models live in `~/sd-models/` and are symlinked.
```
~/ComfyUI/models/
├── unet/
│ └── z_image_turbo-Q5_K_S.gguf → ~/sd-models/diffusion_models/z_image_turbo-Q5_K_S.gguf
├── text_encoders/
│ └── gemma2_2b_lumina2.safetensors (merged from 3 shards, 9.8 GB)
├── vae/
│ └── ae.safetensors → ~/sd-models/vae/ae.safetensors
└── ...
```
### 5.2 Symlink Diffusion Model (GGUF)
The Z-Image Turbo GGUF model must go in `models/unet/` (ComfyUI-GGUF's `UnetLoaderGGUF` node reads from there):
```bash
ln -sf /home/dars/sd-models/diffusion_models/z_image_turbo-Q5_K_S.gguf \
~/ComfyUI/models/unet/z_image_turbo-Q5_K_S.gguf
```
### 5.3 Text Encoder — Gemma 2 2B
Z-Image Turbo uses the Gemma 2 2B text encoder from the Lumina-Image-2.0 family. The original model is sharded into 3 safetensors files. We merge them into a single file for ComfyUI.
**Download from Alpha-VLLM (not gated, no login required):**
```bash
source ~/ComfyUI/venv/bin/activate
python3 -c "
from huggingface_hub import hf_hub_download
import os
repo = 'Alpha-VLLM/Lumina-Image-2.0'
dest = os.path.expanduser('~/sd-models/text_encoders/lumina2_gemma2_2b')
os.makedirs(dest, exist_ok=True)
files = [
'text_encoder/config.json',
'text_encoder/model.safetensors.index.json',
'text_encoder/model-00001-of-00003.safetensors',
'text_encoder/model-00002-of-00003.safetensors',
'text_encoder/model-00003-of-00003.safetensors',
]
for f in files:
print(f'Downloading {f}...')
hf_hub_download(repo, f, local_dir=dest)
print('Done!')
"
```
**Merge shards into single file:**
```bash
source ~/ComfyUI/venv/bin/activate
python3 << 'EOF'
import safetensors.torch
import torch
import os, json
base_dir = os.path.expanduser("~/sd-models/text_encoders/lumina2_gemma2_2b/text_encoder")
output = os.path.expanduser("~/ComfyUI/models/text_encoders/gemma2_2b_lumina2.safetensors")
os.makedirs(os.path.dirname(output), exist_ok=True)
with open(os.path.join(base_dir, "model.safetensors.index.json")) as f:
index = json.load(f)
all_tensors = {}
shards = set(index["weight_map"].values())
print(f"Loading {len(shards)} shards with {len(index['weight_map'])} tensors...")
for shard in sorted(shards):
path = os.path.join(base_dir, shard)
print(f" Loading {shard}...")
tensors = safetensors.torch.load_file(path, device="cpu")
all_tensors.update(tensors)
print(f"Total tensors: {len(all_tensors)}")
print(f"Saving merged file...")
safetensors.torch.save_file(all_tensors, output)
print(f"Done! Size: {os.path.getsize(output)/1e9:.2f} GB")
EOF
```
**Clean up shards (optional):**
```bash
rm -rf ~/sd-models/text_encoders/lumina2_gemma2_2b/
```
### 5.4 VAE
```bash
ln -sf /home/dars/sd-models/vae/ae.safetensors \
~/ComfyUI/models/vae/ae.safetensors
```
### 5.5 Verify All Models in Place
```bash
ls -lh ~/ComfyUI/models/unet/*.gguf \
~/ComfyUI/models/text_encoders/*.safetensors \
~/ComfyUI/models/vae/*.safetensors
```
Expected output:
```
9.8G ~/ComfyUI/models/text_encoders/gemma2_2b_lumina2.safetensors
5.2G ~/ComfyUI/models/unet/z_image_turbo-Q5_K_S.gguf (symlink)
335M ~/ComfyUI/models/vae/ae.safetensors (symlink)
```
---
## 6. Launch Script
### Location: `~/start-comfyui.sh`
```bash
#!/bin/bash
# =============================================================
# ComfyUI Launch Script for AMD BC-250 (ROCm / gfx1013)
# =============================================================
set -euo pipefail
echo "=========================================="
echo " ComfyUI — BC-250 ROCm Launcher"
echo "=========================================="
# --- GPU Health Check ---
if dmesg 2>/dev/null | tail -50 | grep -qi "KIQ fence timeout"; then
echo "[ABORT] KIQ fence timeout detected in dmesg — reboot required!"
exit 1
fi
echo "[OK] GPU health check passed"
# --- ROCm Environment for BC-250 (gfx1013 → gfx1010 spoof) ---
export HSA_OVERRIDE_GFX_VERSION=10.1.0
export HSA_ENABLE_SDMA=0
export HIP_VISIBLE_DEVICES=0
export ROCM_PATH=/opt/rocm
export HSA_TOOLS_LIB=""
export HSA_TOOLS_REPORT_LOAD_FAILURE=0
export PATH="/opt/rocm/bin:$PATH"
export LD_LIBRARY_PATH="/opt/rocm/lib"
# --- Unset old workaround variables that destroy performance ---
unset GPU_MAX_HW_QUEUES 2>/dev/null || true
unset HIP_LAUNCH_BLOCKING 2>/dev/null || true
unset GGML_CUDA_ENABLE_UNIFIED_MEMORY 2>/dev/null || true
unset GGML_HIP_HOST_ALLOC 2>/dev/null || true
unset GGML_CUDA_NO_PINNED 2>/dev/null || true
unset GGML_HIP_NO_COARSE_GRAIN 2>/dev/null || true
unset HSA_DISABLE_FRAGMENT_ALLOCATOR 2>/dev/null || true
# --- PyTorch ROCm tuning ---
export PYTORCH_HIP_ALLOC_CONF="expandable_segments:False"
echo "[OK] ROCm environment configured"
# --- Activate venv ---
cd ~/ComfyUI
source venv/bin/activate
# --- Launch ComfyUI ---
echo "[START] Launching ComfyUI on http://0.0.0.0:8188"
echo "=========================================="
python main.py \
--listen 0.0.0.0 \
--port 8188 \
--force-fp32 \
--lowvram \
"$@"
```
### Usage
```bash
# Foreground (see logs):
bash ~/start-comfyui.sh
# Background with logging:
nohup bash ~/start-comfyui.sh > /tmp/comfyui.log 2>&1 &
# Check if running:
curl -s http://localhost:8188/system_stats | python3 -m json.tool
```
### CLI Flags Explained
| Flag | Why |
|------|-----|
| `--listen 0.0.0.0` | Accept connections from any interface (access from other machines) |
| `--port 8188` | Default ComfyUI port |
| `--force-fp32` | BC-250 gfx1010 has limited FP16 support in PyTorch ROCm; FP32 prevents crashes |
| `--lowvram` | Enables aggressive model offloading — essential for 14.7 GB shared RAM |
---
## 7. ComfyUI Workflow — Z-Image Turbo
### Access the WebUI
Open in browser: **http://localhost:8188** (or `http://<machine-ip>:8188` from another machine)
### Pre-made Workflow
A ready-to-use workflow is saved at:
```
~/ComfyUI/workflows/z_image_turbo_bc250.json
```
Load it via: **Menu → Load → select `z_image_turbo_bc250.json`**
### Manual Node Setup
If building the workflow from scratch, create these nodes:
#### Node 1: UnetLoaderGGUF
- **Type**: `UnetLoaderGGUF` (from ComfyUI-GGUF custom node, category: bootleg)
- **unet_name**: `z_image_turbo-Q5_K_S.gguf`
- **Output**: MODEL → connect to KSampler's "model" input
#### Node 2: CLIPLoader
- **Type**: `CLIPLoader` (built-in, category: advanced/loaders)
- **clip_name**: `gemma2_2b_lumina2.safetensors`
- **type**: `lumina2`**CRITICAL: must be set to lumina2**
- **Output**: CLIP → connect to both CLIP Text Encode nodes
#### Node 3: CLIP Text Encode (Positive)
- **Type**: `CLIPTextEncode`
- **text**: Your prompt (e.g., "a beautiful sunset over the ocean")
- **Input**: clip ← from CLIPLoader
- **Output**: CONDITIONING → connect to KSampler's "positive" input
#### Node 4: CLIP Text Encode (Negative)
- **Type**: `CLIPTextEncode`
- **text**: Empty string `""` (turbo models work best with empty negative)
- **Input**: clip ← from CLIPLoader
- **Output**: CONDITIONING → connect to KSampler's "negative" input
#### Node 5: Empty Latent Image
- **Type**: `EmptyLatentImage`
- **width**: `512`
- **height**: `512`
- **batch_size**: `1`
- **Output**: LATENT → connect to KSampler's "latent_image" input
#### Node 6: KSampler
- **Type**: `KSampler`
- **seed**: Any number (42)
- **control_after_generate**: `fixed` (or `randomize` for variety)
- **steps**: `8` (turbo — more steps won't improve quality)
- **cfg**: `3.0` (turbo models use low CFG guidance)
- **sampler_name**: `euler`
- **scheduler**: `sgm_uniform`
- **denoise**: `1.0`
- **Inputs**: model, positive, negative, latent_image
- **Output**: LATENT → connect to VAEDecode
#### Node 7: VAELoader
- **Type**: `VAELoader`
- **vae_name**: `ae.safetensors`
- **Output**: VAE → connect to VAEDecode's "vae" input
#### Node 8: VAE Decode
- **Type**: `VAEDecode`
- **Inputs**: samples (from KSampler), vae (from VAELoader)
- **Output**: IMAGE → connect to SaveImage
#### Node 9: Save Image
- **Type**: `SaveImage`
- **filename_prefix**: `ComfyUI`
- **Input**: images ← from VAEDecode
- Output images saved to: `~/ComfyUI/output/`
### Wiring Summary
```
UnetLoaderGGUF ───MODEL──→ KSampler
CLIPLoader ───CLIP──→ CLIPTextEncode (positive) ──CONDITIONING──→ KSampler
CLIPLoader ───CLIP──→ CLIPTextEncode (negative) ──CONDITIONING──→ KSampler
EmptyLatentImage ──LATENT──→ KSampler
KSampler ──LATENT──→ VAEDecode
VAELoader ──VAE──→ VAEDecode
VAEDecode ──IMAGE──→ SaveImage
```
---
## 8. BC-250 Specific Tuning
### Environment Variables (set in launch script)
| Variable | Value | Why |
|----------|-------|-----|
| `HSA_OVERRIDE_GFX_VERSION` | `10.1.0` | BC-250 (gfx1013) needs gfx1010 spoof for ROCm |
| `HSA_ENABLE_SDMA` | `0` | SDMA engine has hardware bugs on gfx1013 |
| `HIP_VISIBLE_DEVICES` | `0` | Select the BC-250 GPU |
| `ROCM_PATH` | `/opt/rocm` | ROCm installation path |
| `HSA_TOOLS_LIB` | `""` | Disable profiling tools (stability) |
| `HSA_TOOLS_REPORT_LOAD_FAILURE` | `0` | Suppress tool warnings |
| `PYTORCH_HIP_ALLOC_CONF` | `expandable_segments:False` | Prevent memory fragmentation |
### Variables to NEVER Set
These old workarounds **destroy performance** and must NOT be set:
| Variable | Why it's bad |
|----------|-------------|
| `GPU_MAX_HW_QUEUES=1` | Serializes all GPU ops to 1 queue |
| `HIP_LAUNCH_BLOCKING=1` | Forces synchronous kernel launches |
| `GGML_CUDA_ENABLE_UNIFIED_MEMORY=1` | Page fault overhead |
| `GGML_HIP_HOST_ALLOC=1` | Zero-copy over PCIe is slow |
### Memory Considerations
- Total available: ~14.7 GB shared system RAM
- Diffusion model (GGUF Q5_K_S): ~5.2 GB
- Text encoder (Gemma 2 2B f32): ~9.8 GB
- VAE: ~335 MB
- Total model footprint: ~15.3 GB — exceeds available RAM
- **`--lowvram` is essential**: it offloads models to CPU when not in active use
- Only one component is on GPU at a time during inference
### Resolution Recommendations
| Resolution | Latent Size | Notes |
|-----------|-------------|-------|
| 512×512 | 64×64 | Fastest, recommended for testing |
| 768×768 | 96×96 | Good quality, slower |
| 1024×1024 | 128×128 | May OOM on BC-250 |
---
## 9. Troubleshooting
### "KIQ fence timeout" in dmesg → Reboot
```bash
sudo dmesg | grep -i "KIQ fence timeout"
```
If this appears, the GPU is in a bad state. **Reboot the machine.**
### ComfyUI won't start — "No module named torch"
Make sure you activated the venv:
```bash
source ~/ComfyUI/venv/bin/activate
python -c "import torch; print(torch.__version__)"
```
### "CLIP type not found" or wrong model type
Make sure the CLIPLoader node type is set to **`lumina2`** — NOT `stable_diffusion`.
### OOM (Out of Memory) during generation
1. Reduce resolution to 512×512
2. Ensure `--lowvram` is set
3. Close other programs using RAM
4. Try `--use-split-cross-attention` flag
### Model not showing in dropdown
Verify symlinks are not broken:
```bash
ls -la ~/ComfyUI/models/unet/z_image_turbo-Q5_K_S.gguf
ls -la ~/ComfyUI/models/vae/ae.safetensors
ls -la ~/ComfyUI/models/text_encoders/gemma2_2b_lumina2.safetensors
```
### "UnetLoaderGGUF" node not found
Ensure ComfyUI-GGUF is installed:
```bash
ls ~/ComfyUI/custom_nodes/ComfyUI-GGUF/
pip list | grep gguf
```
### PyTorch ROCm version mismatch
```bash
python -c "import torch; print(torch.version.cuda)"
# Should print: 6.2
```
---
## 10. File Inventory
### Installation Files
| File | Purpose |
|------|---------|
| `~/ComfyUI/` | ComfyUI installation directory |
| `~/ComfyUI/venv/` | Python 3.11 virtual environment |
| `~/ComfyUI/custom_nodes/ComfyUI-GGUF/` | GGUF model loader custom node |
| `~/ComfyUI/workflows/z_image_turbo_bc250.json` | Pre-made Z-Image Turbo workflow |
| `~/start-comfyui.sh` | Launch script with ROCm env vars |
### Model Files
| File | Size | Format |
|------|------|--------|
| `~/ComfyUI/models/unet/z_image_turbo-Q5_K_S.gguf` | 5.2 GB | Symlink → `~/sd-models/diffusion_models/` |
| `~/ComfyUI/models/text_encoders/gemma2_2b_lumina2.safetensors` | 9.8 GB | Merged from Alpha-VLLM/Lumina-Image-2.0 |
| `~/ComfyUI/models/vae/ae.safetensors` | 335 MB | Symlink → `~/sd-models/vae/` |
### Output
| File | Purpose |
|------|---------|
| `~/ComfyUI/output/` | Generated images saved here |
### Python Packages (key ones)
| Package | Version |
|---------|---------|
| torch | 2.5.1+rocm6.2 |
| torchvision | 0.20.1+rocm6.2 |
| torchaudio | 2.5.1+rocm6.2 |
| pytorch-triton-rocm | 3.1.0 |
| transformers | 5.2.0 |
| safetensors | 0.7.0 |
| gguf | 0.18.0 |
| comfyui-frontend-package | 1.39.19 |
---
## 11. Performance Notes
### Startup Output (successful launch)
```
Total VRAM 14750 MB, total RAM 15205 MB
pytorch version: 2.5.1+rocm6.2
AMD arch: gfx1010
ROCm version: (6, 2)
Forcing FP32
Set vram state to: LOW_VRAM
Device: cuda:0 AMD Radeon Graphics : native
ComfyUI version: 0.15.1
ComfyUI-GGUF: Partial torch compile only, consider updating pytorch
```
### Expected Timing (BC-250, 512×512, 8 steps)
| Phase | Estimated Time |
|-------|---------------|
| Model Loading (first run) | 30-60s |
| Text Encoding (Gemma 2 2B) | ~2-5s |
| Sampling (8 steps, Euler) | ~60-90s |
| VAE Decode | ~10-15s |
| **Total (first image)** | **~2-3 min** |
| **Total (subsequent)** | **~1-2 min** |
### Comparison with sdcpp-restapi
| | ComfyUI + PyTorch | sdcpp-restapi |
|---|---|---|
| Frontend | Full node-based GUI | REST API + simple WebUI |
| Model format | GGUF + safetensors | GGUF only |
| Memory management | PyTorch (--lowvram) | ggml manual |
| Flexibility | Full workflow customization | Fixed pipeline |
| Turbo steps | Configurable per-run | Config-based |
---
## Appendix: Quick Start Cheatsheet
```bash
# 1. Launch ComfyUI
bash ~/start-comfyui.sh
# 2. Open browser
# http://localhost:8188
# 3. Load workflow
# Menu → Load → z_image_turbo_bc250.json
# 4. Click "Queue Prompt" to generate
# 5. Images saved in ~/ComfyUI/output/
```
---
*Document generated: 2026-03-02 | System: CachyOS + AMD BC-250 + ROCm 7.2.0*