Glossary · Technical
Multi-Modal Model
An AI model that can process and generate more than one type of content — text, images, audio or video — within the same system.
What it is
Rather than handling only text, a multi-modal model can take an image, a chart, or an audio clip as input and reason about it alongside written language, or produce output that mixes formats.
Why it matters for AI visibility
Multi-modal models can read text embedded in images — screenshots, product photos, infographics — and describe or reference what they see. That means visual content, not just written copy, can factor into whether and how a brand gets mentioned, especially as more platforms accept image and screen-based queries.
See more terms in the full glossary, or read the AI Visibility 101 guide for the full picture.
