Multimodal models still can't ground language in embodied experience · Chi tiết bài viết