Frontiers · 3 MIN READ

Multimodal AI: the model can see. Did it look carefully?

Images, speech, video, and text create richer interfaces—and more ways to misunderstand the evidence.

Original SINLP multimodal schematic illustration
Original conceptual illustration by SINLP · not a data chart

Multimodal AI works across more than one kind of input or output: text, images, audio, video, and sometimes other representations. It can connect a spoken question to a photo, explain a chart, or help navigate a recorded demonstration. That is a broader interface to information, not a guarantee of correct perception.

The Gemini technical report is an early reference for a family of multimodal models. Google’s September 2026 roundup shows the product direction continuing across voice and other interactions. These sources describe systems and releases; SINLP has not independently benchmarked all their modes.

Perception errors propagate

Imagine a system answering a question about a chart. It must identify the axes, read labels, associate marks with a legend, and infer the right relationship. An error at the label stage can turn an otherwise sensible explanation into fiction with units.

Ask for observable details before conclusions. Have the system identify the relevant region or timestamp. Keep the original image or audio available for inspection. If the source is blurry, cropped, or silent at a crucial point, allow the answer to be incomplete.

Audio has several layers

Transcription is not speaker identification, and speaker identification is not an accurate summary of intent. Accents, background sound, domain vocabulary, and overlapping speech can change the result. A meeting summary can omit uncertainty that was audible in the original conversation.

For an important decision, retain the connection between a summary and the recording or verified transcript. Mark uncertain names and numbers. A polished sentence should not quietly resolve an unclear remark into a firm commitment.

Video is a time problem

A clip contains order, duration, and changes between frames. Sampling a few frames may miss the moment that matters. “The worker touched the switch” is different from “the worker turned it off before maintenance.” Temporal evidence belongs in the evaluation.

Test event order and references to specific timestamps. Do not assume a successful description of a still frame implies dependable understanding of an entire procedure.

More inputs mean more boundaries

Images can contain private addresses, documents, faces, or embedded instructions. Audio can contain people who did not expect their speech to be processed. Decide what is necessary for the task, which service receives it, and what is retained.

External content remains evidence rather than authority. A screenshot containing “ignore the user” is not a new instruction from the user. This is the same trust problem discussed in the agent guide, with extra pixels.

An evaluation recipe

Build cases where perception is easy, cases where it is ambiguous, and cases with no answer. Score extracted details separately from conclusions. Include low-resolution images, uncommon labels, confusing charts, and speech interruptions. Compare against human-verified references where appropriate.

Multimodality earns its keep when it reduces friction while preserving evidence. A model that can hear and see is an impressive collaborator. A model that admits it cannot read the tiny axis label is often the more useful one.

KEEP EXPLORING

Spot an error? See our corrections channel and editorial policy.

Pull another thread.