Section 1 · Lesson 1 of 6

What Is AIGC? How AI Generates Images and Video

AIGC (AI-generated content) is AI that creates something new, such as an image, a video, a song or a piece of text, instead of only recognizing what already exists. An image model first finds the idea in a "library of concepts", then removes random noise step by step until a matching picture appears. A video model does the same across time, so the whole clip stays consistent from the first frame to the last.

AIGC vs traditional AI: the food critic and the chef

Traditional AI recognizes things that already exist. It works like a food critic: show it a dish and it tells you "this is pasta, this is a burger, this is a cake." Unlocking your phone with your face works the same way.

AIGC works like a chef. You don't show it any food; you give it an order, such as "a spicy pizza with pineapple and extra cheese", and it cooks a new dish from scratch instead of handing you a frozen pizza from the fridge. In the same way, it can draw a face that has never existed.

A cartoon food critic at a table, labeling spaghetti, a cheeseburger and a cheesecake
Traditional AI: a critic that names what it sees
A cartoon chef stretching pizza dough in a kitchen
AIGC: a chef that cooks something new from your order
Comic panels of a phone camera scanning a face to unlock the phone
Recognizing: face unlock checks a face that exists
Comic panels of a computer generating a wall of different faces
Creating: AI draws faces that have never existed

Step 1: Imagination, a library of concepts

Before it draws anything, the model finds the idea in what is called latent space. Think of a huge library with no books, only properties, where similar things sit close together: a furry corner, a mechanical corner, a food corner.

An illustrated library room with a furry corner, a mechanical corner and a food corner, labeled Latent space: a library of concepts
Latent space, pictured as a library of concepts.

Ask for a cat, and the model goes to the spot where "furry", "pointy ears" and "whiskers" meet. Ask for a robot cat, and it goes halfway between the mechanical corner and the furry corner. It doesn't copy a cat from memory; it builds a new one from the recipe it finds there.

Three photos side by side: furry fur, pointy ears and whiskers
A cat is where furry, pointy ears and whiskers meet.

Step 2: Creation, from noise to a picture

Image models use a process called diffusion. They start from a screen of random noise, like a TV with no signal. For a prompt such as "a cat eating pizza", the model asks: if I change these dots slightly, will it look more like a cat? It repeats that many times, removing noise step by step, until a clear picture appears.

Four steps from raw noise to a sharp picture of a cat eating pizza
Diffusion: from pure noise to a finished picture, one step at a time.

That is why it is closer to sculpting than to painting. A painter adds paint to a blank canvas; the model chips away at noise until the picture is revealed.

A painter adding paint to a large canvas in a studio
A painter adds paint to a blank canvas
A sculptor chiseling a block of stone
Image AI works like a sculptor: it removes what does not belong

How AI generates video

Video is harder because it adds time. Generate 24 separate pictures of a cat and you get 24 slightly different cats; played in a row, they flicker. A video model draws all the pages of the flipbook at once: while it draws one frame, it checks the frames around it, so the fur color stays the same and the movement flows.

It works in small blocks called spacetime patches. Picture a five-second clip as a block of ice: the front is second 0 and the depth is time. The model cuts the block into thousands of small cubes, and each cube knows what a small part of the picture looks like and how it moves. Fitting the moving pieces together is what stops a character from suddenly changing clothes or disappearing.

A clear block of ice with a walking man inside it, cut into small cubes labeled spacetime patches
A clip as a block of ice, cut into spacetime patches.

How does AI know what a cat looks like?

It learned from a very large number of labeled pictures, like a photo of a cat labeled "cat". Put simply, it doesn't keep the photos; it learns the recipe: cats usually have two pointed ears, whiskers and fur.

What this means when you make AI videos

  • Every result starts from new noise, so the same prompt gives a different picture each time. Generate a few versions and keep the best one.
  • Describe the properties you want, such as "orange tabby cat, blue scarf, lantern festival at sunset". That is how the model finds the right place in its library.
  • For the same character in every shot, give the model references instead of hoping it lands on the same face twice. Lesson 4 and Lesson 6 show how.

Key terms

  • AIGC: AI-generated content, meaning images, video, music or text that an AI model creates.
  • Latent space: the model's map of concepts, where similar ideas sit close together.
  • Diffusion: making a picture by starting from random noise and removing it step by step.
  • Spacetime patch: a small piece of a video that holds both part of the picture and how it moves over a short time.

Frequently asked questions

What does AIGC stand for?

AI-generated content: images, video, music and text created by an AI model rather than only analyzed by it.

What is the difference between AIGC and traditional AI?

Traditional AI recognizes and labels things that already exist, like face unlock on a phone. AIGC creates something new from a request, like a picture of a face that has never existed.

How does an AI image generator work?

It finds the idea in latent space, a map of concepts, and then uses diffusion: it starts from random noise and removes it step by step until the picture matches the prompt.

Why is AI video harder than AI images?

Video adds time. The model has to keep every frame consistent, so it generates the whole clip together in small spacetime patches instead of drawing each frame on its own.

Transcript

0:00 · What AIGC is

What is AIGC? AI-generated content refers to artificial intelligence that can create new content rather than just analyzing existing data. Well, traditional AI focuses on recognizing patterns, like unlocking your phone with your face. AIGC focuses on creation, generating a new face that doesn't exist in the real world before.

0:23 · The food critic and the chef

Traditional AI was like a food critic. You show it a dish and it analyzes it: “This is the pasta. This is a burger. Or this is a cake.” It identifies and labels things that already exists. AIGC is like a chef. You don't show it food. You give it an order: “I want a spicy pizza with pineapples and extra cheese.” It then goes into the kitchen and cooks a brand-new dish from scratch. It doesn't just go to the fridge and hand you a frozen pizza. It makes it a fresh one that has never existed before.

0:58 · Step 1: imagination in latent space

How AI generate images The first step is imagination. Think of latent space as a giant library of corners, concepts. How it works: in this library, there aren't books, but properties. Every object is placed based on how similar it is to others. Imagine a giant library room. One corner is a furry corner, one corner is a mechanical corner, one corner is a food corner. If you ask for a cat, the AI flies to the point in the library where furry, pointy ears and whiskers all meet. If you ask for a robot cat, the AI finds a spot exactly halfway between the mechanic corner and the furry corner. Why it's important: the AI doesn't copy a cat from its memory. It goes to the cat coordinates in its brain, builds a brand-new one based on the recipes it finds at those coordinates.

1:56 · Step 2: creation with diffusion

The second step is creation. Once AI has an idea, it needs to turn it into a visual. It uses process called diffusion. The AI starts with a screen full of static noise, like a TV with no signal. You tell the AI: “Show me a cat eating pizza.” It looks at the noise and asks: “If I change these dots slightly, will it look more like a cat?” It repeats these thousands of times, denoising the static until a clear high-definition image emerges. It's not painting like a human adding paint to a blank canvas; it is sculpting order out of chaos.

2:37 · How AI generates video

How AI generates videos Video generation like Sora, Runway, Kling is much harder because it adds a new dimension: time. If you ask AI to generate 24 pictures of a cat, you'll get 24 different cats. When you combine them together as a video, the video will flicker and look crazy. Video AI generates all the pages of a flipbook at once. When it draws page 2, it looks at page 1 to make sure the cat's fur color hasn't changed. It looks at page 3 to make sure the movement flows smoothly. Instead of just working with pixels, tiny dots, the AI works with patches that represent both space part of the image, and time, how that part moves over several seconds. It creates the whole timeline together, ensuring the story makes sense from start to finish.

3:29 · How AI learns what a cat looks like

You might wonder: how does it know what a cat looks like? The AI has looked at billions of images on the internet with labels, for example, a photo of cat labeled “cat.” It doesn't memorize photos. Instead, it learns mathematical recipes for things. It does not store a JPEG of cat. It stores the concept: cats usually have two pointy ears, whiskers and fur.

3:56 · The hardest part: spacetime patches

This is the hardest part. To make a video, the AI has to ensure the movement is smooth and consistent over time, breaking a video into 3D blocks of data that contain both the image and the time. Imagine a five-second video is a solid block of clear ice. The front of the block is the start, 0 seconds; the depth of the block is time, reaching 5 seconds at the back. The AI cuts this giant ice block into thousands of tiny individual ice cubes. Each ice cube, a space-time patch, is a moving puzzle piece. It knows exactly what a small part of the image looks like and how it should move during those few seconds. Why this works: by snapping these moving cubes together, the AI ensures that a character doesn't suddenly change clothes or disappear when they move, because the time and the movements are built into the block.

4:50 · Recap: the sculptor and the flipbook artist

Image AI is like a sculptor chipping away at a block of static noise to reveal a picture. Video AI is like flipbook artist who draws all the pages at once to make sure story moves smoothly.