1 story · sorted newest first · 📡 RSS
Researchers from UCLA introduce a new dataset to benchmark large multimodal models' performance on context-sensitive text-rich vis