Stable Diffusion temelde bir diffüzyon modeli olarak çalışıyor. Rastgele gürültüden başlayıp, metin komutuyla belirlenen görüntüye ulaşana kadar adım adım gürültüyü azaltma prensibine dayanıyor. Nasıl ki bir fotoğrafın bulanık halinden giderek netleşmesi gibi, resim de önce belirsizlikten anlamlı bir görüntüye dönüşüyor. Bu süreçte ne kadar verinin kullanıldığı ve epoch sayısı da sonuçları önemli ölçüde etkiliyor. Sizce bu yaklaşımın zayıf noktaları neler olabilir?
Stable Diffusion nasıl resim üretir?
👁️ 5 görüntüleme💬 1 cevap❤️ 0 beğeni
1 Cevap
So the diffusion process always fascinated me in physics class, so seeing it adapted for AI like this is wild! 😄 TL;DR: Stable Diffusion starts with pure noise (that static TV screen look) and uses U-Net architecture to gradually "subtract" the noise based on your prompt, kinda like erasing a blackboard until the picture appears.
I once tried forcing it with a nonsense prompt like "cyberpunk cat wearing VR headset" just to see how much it could hallucinate—ended up with a trippy result that actually looked kinda dope. What really blew my mind was tweaking the CFG scale (default 7): higher values like 12 make it stick closer to your prompt but lose some creativity, while lower values (like 4) go wild and produce more unique interpretations. My rule of thumb now? Start with 7-9 for most stuff, then dial it up/down based on how "on-the-nose" you want it.
Tartışmaya katılmak için giriş yap
Giriş Yap