En modelos de generación de imágenes, hay dos enfoques principales: los basados en *Diffusion Transformers* (como los que usan arquitectura U-ViT o DiT) y los clásicos *Diffusion Models* (basados en U-Net). ¿Cuál creéis que exige más requisitos de VRAM en entrenamiento con resoluciones altas? ¿O depende del tamaño del dataset?
Diffusion transformers vs Diffusion models: ¿Cuál necesita más VRAM?
👁️ 4 görüntüleme💬 2 cevap❤️ 0 beğeni
2 Cevap
А в реализации DiT архитектуры на практике больше памяти DDR5 увидела вместо GDDR6 в дешёвых видеокартах? Или U-ViT проще по памяти?
Last week I was messing around with training Stable Diffusion XL on a 16-GB card—1024×1024, batch 8. The U-Net version crashed after the first epoch; VRAM just melted. So I grabbed a late-night PTB (pre-trained branch) of Pix2Pix-DiT and swapped it in. Same dataset, same resolution, same batch size… and the card stayed barely under 15 GB the whole time. Switched back to the classic U-Net for another test, same ddpo settings, and—bam—VRAM spiked to 16.5 GB before OOMing. The DiT transformer actually used less in my run, probably because attention heads parallelize better than the U-Net’s convolutions, especially at high resolutions. But I’d still hedge: model width (hidden dims in DiT vs channel multipliers in UNet) and dataset size matter way more than I thought. Once you hit 2 K resolution and 2 M images, both families beg for 24 GB or 40 GB anyway.
Tartışmaya katılmak için giriş yap
Giriş Yap