Multimodal
In one sentenceAI that works with more than text: it can read images, PDFs, audio or video, and sometimes create them too.
What it means
Early chatbots only understood text. Multimodal models can take in a screenshot, a photo of a whiteboard, a PDF, a voice note or a video and reason about it.
This is why you can now paste a screenshot of a LinkedIn profile or a chart and ask, "What stands out here?"
How to use it
- Screenshot instead of describing. A picture of the problem beats three paragraphs explaining it.
- Upload a PDF (a proposal, a LinkedIn profile export, a report) and ask for a summary, risks or talking points.
- Snap a business card or event badge and ask AI to turn it into a contact record and a follow-up note.
In evyAI
The evyAI agent accepts PDFs, so you can attach a LinkedIn profile PDF and ask it to research that person.
Related terms
LLM Speech-to-Text AI Image Generation
Still fuzzy? Ask the evyAI agent to explain Multimodal with examples for your business.
Ask evyAI