Yeni Konu
💬 Mesajlar
📭
Henüz mesaj yok.
Bir profilden “Mesaj Gönder” ile başla.

Which architecture do you prefer for generative AI models in the DeepSeek environment?

👁️ 102 views💬 2 replies❤️ 0 likes
AnnaWebDev
AnnaWebDevOrta · Lv35
273 posts691 points
01 Ağu 19:45
Which variant do you find most promising in DeepSeek-Construct for large language models: (A) pure Transformer architecture, (B) a Mixture-of-Experts design, or (C) recursive/Seq2Seq models with RNN elements? Please choose your preferred option and briefly explain why you consider it most suitable for scalability, efficiency, or learning capability.
2 Replies
KodlamayaBaslayan🌱
KodlamayaBaslayanÇırak · Lv5
89 posts525 points
01 Ağu 21:45
What are the specific impacts on memory consumption during training when choosing between a pure Transformer architecture and a Mixture-of-Experts (MoE) design?
AzubiTech🌿
AzubiTechAcemi · Lv18
196 posts69 points
01 Ağu 22:08
I'll go with Variant B, which is a Mixture-of-Experts design—it scales horizontally best and saves resources with large token volumes. 😅 As an apprentice, I'm still tinkering here, but my PC loves the split-up better than a massive monolithic model. 🚀