Unit 2 / 11

Visual Production Fundamentals: How Image AI Works and Prompt Anatomy

Gains:

  • Understanding how visual artificial intelligence (diffusion) works and how it cannot produce what is not described and cannot draw text well.
  • Ability to write a strong visual prompt consisting of subject, context, style, composition, light and technical components.
  • Gaining the ability to determine the aspect ratio, prevent the readable text from being drawn by artificial intelligence, and reduce the defect with a negative description.

A designer producing visuals with AI is like giving an order to an artist: the clearer and more accurate you explain it, the more accurate the result will be. In this unit, we will understand in plain language how image-generating artificial intelligence (image AI) works and learn the anatomy of writing a good prompt (written instruction for a visual), step by step. The aim is to be able to consciously describe the image you want, rather than "trying randomly and relying on luck".

How does Image AI work? (Simple explanation)

Most image producers today use a method called diffusion. Simply: the model has learned "which words correspond to which images" relationships from millions of image-text pairs. When producing visuals, it starts from an image with random noise (such as a snowy television screen) and cleans this noise step by step to suit your prompt and turns it into a meaningful image. In other words, the model does not "draw" the picture, it statistically constructs the image that best suits your prompt.

This has three practical consequences. First: The same prompt gives a different result every time, because the initial noise is random (we will see how to control this with "seed" in the next unit). Second: The model cannot know what you cannot describe; It can't read the picture in your head, it just interprets what you write. Third: The model cannot draw text and number well, because it imitates letters in shape and not in meaning; That's why it's risky to leave the writing in logos to AI.

Tip: Do not let AI produce the clear text (brand name, slogan) in the image. You produce an image without text and then add the text (as a vector) in a design program. This makes it both readable and editable.

Anatomy of a good prompt

It helps to break a strong visual prompt into six components. You don't need to use them all every time, but choosing consciously will determine the outcome.

  1. Subject/subject: What is the main element of the image? ("a ceramic coffee cup", "portrait of a young woman").
  2. Action/context: What is he doing, where is he? (“at the wooden table, in the morning light”).
  3. Style/aesthetics: What is the visual language? (“minimal product photo”, “watercolor illustration”, “isometric 3D rendering”). This is the most decisive component.
  4. Composition / angle: How is the frame? ("close-up", "bird's eye view", "wide angle", "in the middle, with space").
  5. Light and color: ("soft natural light", "earth tone palette", "high contrast").
  6. Technique / quality: ("high resolution", "1:1 frame rate", "background plain white").

There is also a negative recipe: saying things you don't want ("no text, no people, no messy background"). We will deepen this with "negative prompt" in the 4th unit.

Glossary of terms

  • Aspect ratio: The aspect ratio of the image; 1:1 frame (Instagram post), 9:16 portrait (story/reels), 16:9 landscape (cover).
  • Resolution: Number of pixels in the image; High is enough for printing, medium is enough for screen.
  • Style reference: Giving a sample to the model saying "look like this".
  • Render: The computer calculates and produces an image; "3D rendering" means realistic volumetric view.

Step by step: describing an image

Let's say we want a product image for an olive oil brand.

Step 1 — Focus on subject: “extra virgin olive oil in a dark green glass bottle.”

Step 2 — Add context: “on the rustic wooden counter, with a few fresh olive branches next to it.”

Step 3 — Choose style: “premium product photo, realistic”.

Step 4 — Give the composition: “product in the middle, space for text at the top”.

Step 5 — Light/colour: “soft side natural light, warm earth tones”.

Step 6 — Technique: “1:1 frame rate, high resolution, plain background”.

When you combine them all, you get a workable draft that suits the brand's identity.

three mini cases

Case 1 — From uncertainty to clarity. A designer wrote “image of a modern office” and made 12 attempts, none of which worked (lost 30 minutes). He rewrote the prompt into six components: "bright, minimal office; wooden desk; natural light from large window; warm neutral tones; wide angle; text space at top". 2 of the first 4 attempts were usable; The time was reduced to 10 minutes.

Case 2 — Text trap. One intern had an AI produce a cafe poster from start to finish; The text "COFFEE" in the image turned out to be "COFEEE" and the prices were not readable. The instructions changed: the image was produced without text, the text was added later in the design program. The error decreased to zero and the poster was suitable for printing.

Case 3 — Loss of ratio. A social media expert produced a 1:1 square image for Instagram story; 9:16 When I put it in the vertical area, the top and bottom were cut off and the important element was lost. When I specified the aspect ratio as "9:16 vertical" in the prompt, the image fit into its format and no reproduction was required.

Copiable templates

1) Six-component product image skeleton:

[Subject: describe the product] , [Context: where/how] .Style: [e.g. premium product photo, realistic] .Composition: [e.g. product in the middle, space for text at the top].Light and color: [e.g. soft natural light, earth tones] .Technique: [aspect ratio, high resolution, plain background] .Note: No legible text/logo in the image; I will add the text later.

2) Illustration skeleton:

Subject: [what to draw].Style: [e.g. flat color vector illustration / watercolor / isometric 3D] .Color palette: [2-3 primary colors] .Mood: [e.g. cheerful, calm, serious] .Background: [plain / patterned / transparent] .Ratio: [1:1 / 9:16 / 16:9] .

3) Style exploration (same subject, different aesthetic):

Describe the following subject with the same composition in 4 different visual styles: minimal flat vector, realistic photo, watercolor, isometric 3D. Write a one-line prompt for each. Subject: [paste]

4) Adding negative description:

Improve the following prompt and add the unwanted ones at the end: "no text, no watermark, no messy background, no bad hands/fingers, no too many objects". Prompt: [paste]

Weak prompt / Strong prompt

Weak: a beautiful coffee photo

Result: Random angle, unclear style, incidental image that doesn't fit the brand.

Powerful: Ceramic cup filled with a hot latte, on light wooden table, sideways in soft morning light; minimal premium product photo; warm neutral tones; close-up, space for text at top right; 1:1 square, plain background; There should be no legible text.

The result: a brand-adaptive sketch with controlled composition, light and proportion.

Difference: A strong prompt fills the six components (subject, context, style, composition, light, technique) clearly and leaves the text out.

Prompt components and effect table

component

example

Effect on the result

subject

"olive oil in a glass bottle"

determines what appears

Style

"premium product photo"

The element that most determines the visual language

composition

"in the middle, with text space"

Increases usability

light/color

"soft light, earth tones"

Conveys the brand feeling

aspect ratio

"9:16 vertical"

Allows to comply with the format

negative description

"no text"

Reduces defect

Common mistakes

  • Not specifying the style: By omitting the most decisive component, the result becomes cliché and random.
  • Having the AI ​​draw the text: Corrupt, unreadable letters; add the text later in the design program.
  • Forgetting the ratio: The wrong aspect ratio creates clipping and element loss in the publication.
  • Overloaded prompt: stacking 20 concepts on top of each other confuses the model; keep it clear and prioritized.
  • Giving up in one try: Diffusion is random; It is normal to produce several variations and choose the best one.

In summary

The image-generating AI gradually constructs the image to fit your prompt from random noise; It can't read the picture in your head, it just interprets what you write. A good prompt consists of six components: subject, context, style, composition, light/color, and technique. Style is the most decisive element; Be sure to specify the aspect ratio; Don't leave the readable text to the AI, you add it later. As clarity increases, the number of attempts and time loss decreases.

Application task

Choose a product or topic. First, write it in one sentence (“a nice X”) and try it in an image generator and note the result. Then describe the same subject again by filling in the six-component skeleton, adding the aspect ratio and negative description. Put the two results side by side and write in three bullet points how clarity changes the output.

checklist

  • [ ] I stated the subject, context and style clearly in my prompt.
  • [ ] I added composition and light/color components.
  • [ ] I specified the aspect ratio (1:1 / 9:16 / 16:9).
  • [ ] I didn't have the AI ​​draw the readable text, I left it for later.
  • [ ] I excluded the undesirables with the negative description.
  • [ ] I did not trust a single experiment and produced several variations.