Glossary · Technical

Multi-Modal Model

An AI model that can process and generate more than one type of content — text, images, audio or video — within the same system.

What it is

Rather than handling only text, a multi-modal model can take an image, a chart, or an audio clip as input and reason about it alongside written language, or produce output that mixes formats.

Why it matters for AI visibility

Multi-modal models can read text embedded in images — screenshots, product photos, infographics — and describe or reference what they see. That means visual content, not just written copy, can factor into whether and how a brand gets mentioned, especially as more platforms accept image and screen-based queries.

See more terms in the full glossary, or read the AI Visibility 101 guide for the full picture.