Agentic AI Notebook
Phase 27

Video Models

~3 min read

Concept & How It Works

  • Key points are in the visual diagram above.

Why Does It Exist?

Real-world tasks involve multiple data types: analyze a chart in a PDF, describe a video, answer questions about a photo. Multimodal models handle all inputs in one model, eliminating brittle multi-model pipelines.

Real-World Analogy

A doctor who can listen to symptoms (audio), read lab results (text), examine X-rays (images), and watch patient movement (video) — all integrated into one diagnosis, not four separate specialists.
Loading diagram...

Visual Workflows

What is Video Models?

Loading diagram...

Example

Scenario

Product support agent: customer sends screenshot of error + describes problem in text. GPT-4o analyzes screenshot (reads error message, identifies UI state) combined with text description → provides targeted fix. Single model, no separate OCR + LLM pipeline.

Solution

In Model Engineering (Awareness), apply Video Models to this scenario: Product support agent: customer sends screenshot of error + describes problem in text. Identify the inputs, run the technique, validate the output, and note one thing you would monitor in production.

Practice Task

Do this before moving to the next module — reading alone is not enough.

Open the Code Walkthrough below and run it locally. Change one parameter related to Video Models (e.g. model, temperature, top_k, or tool name), observe the difference in output, and write 2–3 sentences explaining what changed.

Code Walkthrough

Highlighted lines show where Video Models happens in the code.

Video Models
1from openai import OpenAI  # import dependencies2import base64  # import dependencies3
4client = OpenAI()  # create API client5
6# Multimodal: text + image + structured request7response = client.chat.completions.create(  # call the API8    model="gpt-4o",9    messages=[{10        "role": "user",11        "content": [12            {"type": "text", "text": "Analyze this dashboard screenshot. What trends do you see? Any anomalies?"},13            {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{base64.b64encode(open('dashboard.png','rb').read()).decode()}"}},14        ],15    }],16)17
18# Gemini multimodal: text + video19# import google.generativeai as genai20# model = genai.GenerativeModel("gemini-2.0-flash")21# response = model.generate_content(["Summarize this video:", video_file])

Commands to Remember

Commands to Remember

  • pip install peft transformers # LoRA / QLoRA fine-tuning
  • pip install bitsandbytes # quantized training

Common Mistakes

  • Using separate OCR + LLM when native multimodal suffices
  • Sending full-resolution images — expensive and unnecessary
  • Not accounting for multimodal token costs in pricing
  • Assuming all models handle all modalities equally

Cheat Sheet

Quick recap — the most important points from this module.

Cheat Sheet

quick ref
  • content: [{type:text}, {type:image_url}]
  • GPT-4o: text + image + audio
  • Gemini: + video natively
  • Images ≈ 765 tokens each
  • Multimodal RAG: CLIP + text embeddings
  • Resize images to reduce cost