Speaker:Xingping Dong (Wuhan University)
Time:2023-12-15 15:00
Location:Conference Room 686 at the 6th floor of Shuli Building at Haiyun Campus
Abstract:
With the vigorous development of deep learning, many artificial intelligence algorithms have achieved tremendous success in machine learning tasks, such as speech recognition, machine translation, and facial recognition. However, most of these tasks are based on single-modal data, while real-world data is often multimodal, as seen in movie data that includes video, audio, and text subtitles. The development of multimodal learning not only enables the effective utilization of the growing multimodal data but also provides a more comprehensive understanding of real-world data, enhancing the efficiency and capabilities of models. It can also handle more complex tasks, such as the recently highlighted text-to-image generation task.
This presentation will briefly introduce the basic concepts and development of multimodal learning, along with its applications in the visual and language domains. Through the analysis of specific tasks, including visual language navigation, language referring object segmentation and tracking, we will explore the challenges and solutions that multimodal learning faces in practical applications, providing initial insights for the practical implementation of multimodal learning algorithms.