Free account = 1 chapter of every course unlocked
No credit card · Google sign-in in 30 seconds · 25 free chapters, one per course
Start free →
All Courses/Multimodal AI Systems
A camera lens, a microphone and a single keyboard key together on a desk

Multimodal AI Systems

Build systems that process text, images, audio, and video together. Cross-modal attention and fusion architectures.

8 chaptersabout 9 hoursFirst chapter free with a free accountFull access: Pro $8.99/month or Lifetime $149 once

After this course, you'll be able to:

✓Understand architectures like BLIP-2, Flamingo, Gemini, GPT-4o
✓Build systems that process text + images + audio together
✓Implement cross-modal attention and fusion
✓Deploy multimodal models locally

Who this is for

  • →Engineers who have used a vision-language API and now need to know why it fails on dense documents, small text or fine spatial detail.
  • →Machine learning practitioners who understand language models and want the mechanics of how images and audio get into one in the first place.
  • →Anyone evaluating or fine-tuning an open multimodal model and needing to judge which architectural family fits their constraints.
  • →Robotics developers curious about vision-language-action models and what changes when the output is a motor command rather than a sentence.
  • →Not for you if you want to generate images. Text-to-image diffusion is adjacent but a different craft with different tooling.
  • →Not for you if you have not worked with transformers or language models before. This material starts above that line and does not re-teach attention from scratch.

What you need first

  • ·A working understanding of transformer language models: tokens, embeddings, self-attention, autoregressive decoding, context windows.
  • ·Some prior contact with image models. You should know roughly what an image encoder produces and what a patch embedding is.
  • ·Python and PyTorch or an equivalent framework, at the level of loading a pretrained model and running inference on your own data.
  • ·Enough hardware to run a small vision-language model for inference, or cloud access. Training a connector layer is feasible on rented hardware; pretraining is not a realistic exercise.
  • ·Willingness to read papers. This subject moves through published architectures, and the primary sources are where the design decisions are actually explained.

The real question: how does a picture get into a language model

A language model consumes a sequence of vectors. An image is a grid of pixels. Everything interesting about multimodal AI lives in the translation between those two facts, and the different answers people have given to it produce systems with visibly different strengths and weaknesses.

The naive framing is that a multimodal model "understands images". A more useful framing is that some component converts an image into a sequence of vectors that occupy the same space the language model already reads from, and the language model then does what it always did: predict the next token conditioned on a sequence. Almost every practical consequence follows from how that conversion is done and how much of the model was trained with it in place.

Three questions separate the architectures. First, does the image encoder stay frozen or get trained alongside the language model? Freezing preserves a strong pretrained visual representation and is far cheaper, but it caps what visual detail can ever reach the language model, because the encoder was optimized for a different objective. Second, do visual vectors get inserted into the input sequence alongside text tokens, or does the language model reach out to them through dedicated cross-attention layers? Third, was the model trained on interleaved multimodal data from early in its training, or was vision attached to a finished text model afterward?

Those three choices explain most of the behavior you will observe. A model with a frozen low-resolution encoder will read a headline and fail on a footnote. A model that inserts a large number of image tokens into the context will be accurate on fine detail and expensive per request. A model with vision bolted on late will be excellent at describing a photograph and mediocre at reasoning that requires tightly interleaving what it sees with what it knows.

The course is organized around named architectures rather than abstractions, because the abstractions only become concrete when you can see how a specific system resolved these tensions. The design space is small enough to hold in your head once you have walked through several points in it, and that is what makes model cards readable rather than mysterious.

Three architectural families and what each one costs

The projector approach. Take a pretrained image encoder, take a pretrained language model, and train a small module — a linear layer or a shallow MLP — that maps encoder outputs into the language model's embedding space. The projected vectors are inserted into the input sequence as if they were tokens. LLaVA, from a collaboration involving the University of Wisconsin-Madison, Microsoft Research and Columbia University, popularized this design, and it dominates the open-weight ecosystem because it is by far the cheapest path to a working vision-language model. Both large components stay frozen, at least initially, and the trained part is tiny.

Its costs are equally clear. The number of visual vectors scales with encoder output, so higher resolution means more tokens, which consumes context and inflates cost per request. Quality is bounded by the frozen encoder: if it never encoded the small text in the corner of the image, no amount of language modeling recovers it.

The resampler and query approach. Instead of passing everything through, compress. Flamingo, from DeepMind, introduced a perceiver-style resampler that maps a variable number of visual features into a fixed, small number of vectors, then feeds them to the language model through gated cross-attention layers inserted between the frozen language layers. BLIP-2, from Salesforce Research, took a related route with the Q-Former, a lightweight transformer holding a set of learned query vectors that extract the visually relevant information and produce a fixed-size output.

The gain is a bounded token cost regardless of input size, which makes multi-image and interleaved input tractable. The loss is that a fixed-size bottleneck must discard something, and what it discards was decided during training. These architectures tend to be strong at holistic understanding and weaker at tasks demanding exhaustive detail.

Native multimodal training. Train the model on multiple modalities from the outset, with a shared representation and, in some designs, the ability to generate as well as consume non-text output. Publicly described systems in this family include Google's Gemini and OpenAI's GPT-4o, and in the open literature Show-O, from a collaboration between the National University of Singapore's Show Lab and ByteDance, unifies understanding and generation in a single transformer.

This gives the tightest integration and the lowest-latency path for interactive audio and video, since there is no chain of separate models to traverse. The cost is that it is a pretraining-scale undertaking, which puts it out of reach for anyone not operating at that scale. For most practitioners these systems are things to use and evaluate rather than build.

Token budgets, resolution, and the cost you did not plan for

The single most practical thing to understand about multimodal systems is that an image is not free, and its price is measured in tokens.

An image passed to a projector-style model is converted into some number of vectors that occupy context exactly like text does. Increase the input resolution and that number rises, often steeply, because patch-based encoders produce a count that scales with area. A conversation with several images can consume more context than a long document, and each turn re-sends them unless the implementation caches.

This creates a genuine trade-off rather than a tuning knob. Many widely used image encoders were pretrained at modest fixed resolutions, so an image is downscaled before the model ever sees it. That is why a model can describe a scene accurately and then misread a serial number, a chart axis label or a line in a table: the pixels carrying that information were destroyed before encoding. Tiling strategies work around it by cutting a high-resolution image into patches, encoding each separately, and usually adding a downscaled full view for global context. Accuracy on dense material improves substantially. Token count multiplies accordingly.

The engineering consequences are worth planning for rather than discovering. Cost per request for image-heavy workloads can dominate a budget in a way text workloads rarely do. Time to first token grows because the encoder must run before decoding starts. Batching is harder because different images produce different token counts. And prompt caching, which works well for repeated text prefixes, needs explicit support to help with repeated images.

Video makes all of this sharper. A video is a sequence of images, and encoding every frame is immediately unaffordable. Real systems sample: uniformly, at keyframes, or by selecting frames adaptively based on change. Sampling is where most video understanding failures originate, because an event shorter than the sampling interval simply does not exist as far as the model is concerned. When a video model misses something obvious, check the sampling rate before questioning the model.

The correct discipline is to decide the resolution and frame policy from the task. Reading a dense document needs high resolution and probably tiling. Describing a scene does not. Counting objects in a crowd needs resolution and will still be unreliable. Matching the policy to the requirement is where most of the achievable cost saving lives.

Hallucination in multimodal models is a specific, diagnosable thing

Multimodal hallucination has a characteristic signature that distinguishes it from ordinary language model confabulation, and recognizing it changes how you debug.

The dominant failure is the language prior overriding the visual evidence. Ask whether a kitchen photograph contains a refrigerator and the model may answer yes because kitchens usually contain refrigerators, not because it saw one. Ask about an object commonly co-occurring with something present in the image and you will get a confident affirmative. This is object hallucination, and it is measurable: the POPE evaluation approach, introduced by researchers at Renmin University of China and collaborators, probes it directly by asking balanced yes-or-no questions about objects that are and are not present, including ones the training distribution would lead the model to expect.

A second failure mode is spatial and relational. Models frequently identify every object in a scene correctly and then get left and right, above and below, or which object is holding which, wrong. The visual representation is often closer to a bag of recognized content than a structured scene, particularly in architectures where a compression bottleneck discarded spatial arrangement.

A third is counting. Reliable counting requires attending to each instance separately, and pooled or compressed representations are poorly suited to it. Treat any count from a vision-language model as an estimate.

Evaluation has to account for all of this. Aggregate scores on standard visual question answering benchmarks are weak signals because many questions are answerable from the language prior alone, and because widely used benchmarks are old enough that contamination is a live concern. Better practice is to build a small evaluation set from your own images with questions where the answer cannot be guessed, include negative cases where the correct answer is that something is absent, and check that the model can say it does not know. Free-form description outputs need human or model-assisted grading against the actual image content, and if you use a model as judge, remember it inherits the same priors you are trying to detect.

The practical rule: never accept a multimodal output as ground truth for a downstream automated action without a verification path. Extraction tasks should be checked against a schema and, where possible, against a deterministic source.

Beyond images: audio, video, and action as a modality

The same architectural logic extends past vision, and the extensions are where the field is currently most unsettled.

Audio. There are two distinct designs and they behave differently. The pipeline approach transcribes speech to text with a separate model and feeds the text to a language model, which is simple, debuggable, and throws away everything that was not words: tone, hesitation, emotion, overlapping speakers, background sound. The native approach encodes audio directly into the model's representation space, preserving those signals and removing a round trip, which matters enormously for conversational latency. Native audio also enables the model to produce speech directly rather than generating text for a separate synthesizer, which is what makes interruption and natural turn-taking feasible.

Video. Beyond frame sampling, the open problem is temporal reasoning: understanding that one event caused another, tracking an object through occlusion, or summarizing a long recording without processing every frame. Current approaches compress along the time axis with pooling or learned temporal modules, and all of them trade temporal precision for tractable cost. Long-form video understanding remains genuinely hard, and claims about it deserve scrutiny about the sampling and context strategy behind them.

Action. The most interesting recent direction treats robot control as another output modality. Vision-language-action models take camera input and a natural language instruction and emit actions directly. OpenVLA, released as an open model by a group including Stanford and UC Berkeley researchers, follows this pattern by building on a vision-language backbone and training it on large-scale robot demonstration data, with actions represented in a form the model can produce autoregressively.

The reason this is worth studying even if you never touch a robot is that it exposes what the architecture is really doing. A model producing motor commands cannot get away with a plausible-sounding answer. It either grounded the instruction in what the camera sees or it did not, and the consequence is immediate and physical. Control also introduces requirements that text generation never faced: a fixed control frequency, safe behavior under uncertainty, and a distribution shift problem where the model's own errors move it into states no demonstration ever covered.

Deploying multimodal systems, and the security surface nobody expects

Running a multimodal system introduces operational and security problems that a text-only pipeline does not have.

Preprocessing parity. Image models are unusually sensitive to the exact preprocessing they were trained with: resize method, interpolation, normalization constants, channel order, and how aspect ratio is handled. Mismatched preprocessing does not raise an error, it produces a model that seems mildly worse for no discoverable reason. This is the first thing to check when a locally run model underperforms its published description.

Memory and hardware. A vision-language model carries an encoder alongside the language model, and the encoder's activations at high resolution can be a meaningful share of peak memory. Quantization behaves differently across the two components, and a quantization scheme that leaves the language model intact can degrade the vision tower noticeably. Evaluate the quantized artifact on visual tasks specifically, not only on text ones.

Latency composition. Image encoding runs before the first output token. For interactive use, that fixed cost sits in front of every request and is not hidden by streaming. High resolution and tiling make it worse. If responsiveness matters, resolution policy is a latency decision as much as an accuracy one.

Prompt injection through images. This is the genuinely underappreciated risk. Any model that reads text in images will read text an attacker placed there. Instructions embedded in a screenshot, a PDF, a document photograph or even faintly rendered in an image can be picked up and followed, and unlike a text prompt the payload is not visible to a casual reviewer of the request log. The mitigations are architectural rather than clever prompting: treat all model output derived from untrusted images as untrusted data, never let it directly trigger a privileged tool call, and keep a human or a deterministic validator between the model and any consequential action.

Data handling. Images and video carry more incidental personal information than text does, including faces, documents, locations and metadata, and in most jurisdictions that raises the compliance stakes. This is a substantial part of why organizations run multimodal models on their own infrastructure, and open vision-language models are now capable enough that doing so is realistic rather than a compromise.

Whether this is the right thing to study now

The honest readiness test is whether you have already hit the API ceiling. If you have called a vision-language endpoint, watched it confidently misread a table, and found that no amount of prompt rewriting fixed it, you are exactly the reader this material is written for, because the explanation is architectural and lives below the interface. If you have not yet worked with transformers or run an image model at all, the connector-level discussion will float past you; spend time on language models and on vision fundamentals first, and both roads arrive here.

The other useful signal is having a real input to point at. Multimodal behavior is idiosyncratic in a way that only shows up on your own material: your screenshots, your scanned forms, your camera feed, your recordings. Two models with similar published descriptions can behave very differently on the same page of a document, and the only way to find out is to run both against material you know well enough to grade.

Common questions

What is the difference between a multimodal model and a model with a vision plugin?

A pipeline runs a separate model first — an image captioner or an OCR engine — and hands text to a language model, which means the language model only ever sees a description written by another model. A multimodal model receives a representation of the image itself, projected into the space it reads from, so the language model can attend to visual detail the captioner would have omitted. Pipelines are easier to debug and cheaper; integrated models are better at anything requiring detail the intermediate step did not think to mention.

Why does my vision-language model misread text in images?

Almost always resolution. Many image encoders were pretrained at modest fixed resolutions, so a high-resolution screenshot is downscaled before encoding and small text is destroyed before the model ever sees it. Models that handle dense documents well typically use a tiling strategy that encodes crops separately at higher effective resolution. If your model does not, either enable that mode if it has one, crop the region of interest yourself, or use a dedicated document model. Turning up the prompt specificity will not recover pixels that were discarded.

Can I run multimodal models locally?

Yes. Open vision-language models in the projector-style family are widely available in sizes that run on a single consumer GPU, and quantized versions run on less. The considerations are that the vision encoder adds memory beyond the language model weights, that high-resolution or tiled input increases activation memory substantially, and that quantization can affect the vision tower differently from the language model. Verify quality on your own visual task after quantizing rather than assuming published text benchmarks carry over.

Is fine-tuning a multimodal model realistic on a modest budget?

Training the connector or applying a low-rank adaptation to a projector-style model is genuinely achievable on rented hardware, and it is the standard route for adapting an open model to a specific visual domain. What is not achievable outside a large lab is pretraining a natively multimodal system, which requires the scale of data and compute used to train a frontier model. Be clear which one a tutorial is describing, because the resource requirements differ by orders of magnitude.

How do I stop a multimodal model from hallucinating objects?

You reduce it rather than eliminate it. Include negative questions in your evaluation so you can measure the rate rather than guess at it. Prompt explicitly for uncertainty and permit an answer of not visible or not determinable. Prefer higher resolution for tasks where the evidence is small. Constrain the output to a schema so a spurious detail has nowhere to go. Most importantly, do not let an unverified visual claim trigger an automated action; put a validator or a person between the model and the consequence.

Are vision-language-action models relevant if I do not work in robotics?

They are worth understanding for what they reveal. Because the output is a physical action rather than a sentence, a model cannot bluff, so this family exposes very clearly whether an architecture actually grounds language in perception. The design questions — how much visual detail survives the connector, how the model handles states it never saw in training, how you evaluate without a reference answer — are the same ones facing any multimodal system where the output drives something automatic.

Related reading

Full syllabus

1

Multimodal AI Complete Guide

Free preview
Read free →
2

BLIP-2 & Q-Former

3

Flamingo Architecture

4

Gemini Multimodal

5

GPT-4o Native Multimodal

6

LLaVA Evolution

7

OpenVLA Implementation

8

Show-O Unified Model

Unlock all 8 chapters

Plus 24 other courses — 553 more chapters included.

Every course, every future course, the Python Lab and eight downloadable kits, nothing to renew. Or subscribe: Pro $8.99/month

Free Tools & Calculators