Multimodal AI
Multimodal AI describes systems that can process several input and output formats at the same time — text, image, audio and video, for example — instead of being limited to a single modality such as plain text.
In practice
Multimodal AI search systems are increasingly able to evaluate the images, alt text and video transcripts on a website as well. Meaningful image descriptions and well-structured video subtitles thereby become usable for machine analysis too – and they are needed for accessibility and usability regardless.