Discussion about this post

User's avatar
Rainbow Roxy's avatar

Wow, it's fascinating how you explain the shift! I'm super curious, how exactly do multimodal LLMs bridge that gap of lost image-level information for structured datasets? Such a smart analysis!

1 more comment...

Ready for more?