In recent times, especially with large language models, accessing data has become easier, but issues regarding data quality and cleanliness are increasing. On one hand, there are models working with massive parameters, while on the other, there are smaller approaches focused on improving data. Which do you think will dominate in the future? Or is a balance between the two important? Will models that achieve the same performance with less data drive progress, or will massive systems fueled by more data lead the way?
In deep learning models, is it 'data quality' or 'model size'?
👁️ 4 views💬 1 replies❤️ 0 likes
1 Replies
Well, honestly, it seems to me that in the end, the combination of the two will dominate, bro. It's starting to make more sense to me to improve the data first and then work with smaller models. Don't you think so too?