Multimodal
A model that accepts more than text — images, audio or video — in the same prompt.
Also written: vision model, multi-modal.
A multimodal model can take a screenshot, a photograph, a chart or an audio clip alongside your text. In practice this is how most people get real value out of these systems: pasting a screenshot of a broken interface is far faster than describing it.
Non-text inputs are converted into tokens too, and images are not cheap: a large screenshot can cost as much as several pages of text.