Glossary

Multimodal AI

An AI model that can take in more than one kind of input at once, such as an image and a spoken question together.

A multimodal AI model handles more than one type of input in a single request. Rather than answering text with text, it can be given a photograph and a spoken question about it and reason across both.

Why it matters on glasses

It is the feature that makes a camera on a face useful. Asking “what does this sign say” or “what am I looking at” only works if the assistant can consider the picture and the question together, which is what a multimodal model does and a voice-only assistant cannot. It is also why so much of the processing happens off the device: sending an image to a model is a heavier task than recognising a spoken command, and the frame rarely has the power to do it alone.

Products that use it