Unit 3 / 11

Photographic Visual Production and Prompt Anatomy: How Does Diffusion Work?

Gains:

  • Understanding how image-generating artificial intelligence (diffusion) works and what recipe components are required for photographic realism
  • Ability to write a strong photographic prompt consisting of subject, light, lens/camera, composition, atmosphere and technical components.
  • Ability to distinguish the difference between a produced image and a photograph taken, the limits of use and the importance of the declaration of authenticity

A photographer uses generative visual AI for three legitimate purposes: to produce draft references for moodboards and concepts, to create elements that complement the actual shot (background, texture, composite piece), and to create fully produced images (for example, for a stock sale or advertising concept). In this unit, we will learn how these tools work and how to write prompts for photographic realism. The goal is not to wait for a nice shot by chance; It means describing what you want in a photographer's language and directing the result.

Diffusion: how does the model produce "photographs"?

Most of today's image-generating AIs work with a method called diffusion. Simply: the model has learned to go “from noise to meaningful image” on millions of images. It starts from random noise during production (like a snowy screenshot) and creates an image by gradually cleaning the noise to fit your text (prompt). Two basic facts emerge from this:

First, the model cannot produce what is not described. If you don't say the light, a random light will come; If you do not consider the objective view, the "feeling of photography" will remain a coincidence. Second, the model does not understand, it predicts. That's why he makes mistakes (hallucinations) in places that require logic, such as hands, text, reflection, symmetry. If you want photographic realism, you must clearly describe the technical decisions—light, lens, composition—that make a photograph a photograph.

Caution: Although a generated image may look "photo-like", it is not a photo; it does not document an actual moment. Using it in a news story, documentary, or context claiming “this really happened” is misleading and violates the ethical boundary, which we will cover in unit 9.

Photographic prompt anatomy

A powerful photographic prompt consists of six components. Think of these like planning a shoot:

  1. Subject: What/who? Age, clothing, expression, action, number. ("a middle-aged carpenter, in an apron, working at the workbench")
  2. Light: Direction, hardness, source, time. It is the heart of photography. ("soft side window light, morning, low contrast")
  3. Lens and camera view: Focal length, depth of field, angle. ("85mm portrait lens feel, shallow depth of field, soft background bokeh, eye level")
  4. Composition: Framing, placement, framing. ("close-up, rule of thirds, subject on left")
  5. Atmosphere and color: Tone, mood, palette. ("warm earth tones, calm, nostalgic")
  6. Tech/quality: Texture, notes of realism, aspect ratio. ("natural skin texture, film grain, 3:2 ratio")

When you introduce these components in sequence, the model produces a much more controlled, "photographic" result. Any component you leave out is left to chance.

Tip: Do not have the AI ​​draw the text/logo. Diffusion models cannot reliably produce readable text and proprietary logos; add them in the actual design phase. Also, in an image that requires readable text, use AI only for the backdrop/atmosphere.

Aspect ratio and intended use

Aspect ratio is the width-to-height ratio of the image and determines its use: 9:16 for a social media vertical story, 3:2 for standard print, 1:1 for a square post, 16:9 for a wide banner. It is better to specify the ratio in the prompt from the beginning than to trim it later and lose quality. Choose the resolution and ratio from the very beginning, knowing the intended use (printing or display).

Negative recipe and variation

Most tools also have a negative prompt: an area where you write things you don't want in the image ("deformed hand, extra fingers, text, watermark, plastic binding"). This reduces but does not completely prevent common defects; verification is still essential. There is also the concept of seed — the seed of randomness from which production begins; The same seed and prompt produces a similar result again, which helps you create a consistent set (we will go into more depth in the 4th unit).

three mini cases

Case 1 — The power of recipe. A stock photographer wrote "woman drinking coffee" and got dismal results. When Prompta added light (side window light), lens (50mm, shallow depth of field), atmosphere (warm, calm morning) and texture (natural skin), the result became usable. Acceptable frame generation rate increased from approximately 1/20 to 1/4; Production time was reduced by one third.

Case 2 — Loss from hallucination. A designer-photographer put an image of "business people shaking hands" into a presentation without checking it. At the meeting, it was noticed that a manager had three hands. The flaw that was overlooked on the small screen was magnified in projection. A 30-second hand/finger scan would have prevented this embarrassment.

Case 3 — Choosing the right tool. A product photographer wanted to place a real watch he shot in a different environment. Instead of producing it from scratch, it kept the real product frame, only the background/atmosphere was produced with AI and made composite. Thus, the actual details of the product (brand, dial) were not distorted; both credibility and accuracy were preserved.

Copiable templates

1) Photographic prompt skeleton:

[subject: who/what, age, clothing, expression, action],[light: direction, harshness, source, time],[lens/camera: sense of focal length, depth of field, angle],[composition: framing, placement, convention],[atmosphere/color: tone, mood, palette],[technique: natural texture, grain, aspect ratio].Photographic, realistic. Add text/logo.

2) Negative prompt draft:

Negatives (what I don't want): malformed hand, extra/missing fingers, crooked face, plastic/wax-like skin, unreadable text, watermark, repeating texture, impossible shadow, oversaturated color.

3) For moodboard reference (direction, not delivery):

Purpose: moodboard reference for a shoot (not the deliverable).Concept: [concept]. Describe the aesthetic with attributes only: light [..], color palette [..], atmosphere [..], composition [..]. Do not use any person/brand/photographer names. Generate 4 variations.

4) Composite plan protecting the real product:

I have a real [product] shot; The product itself will not change. Write the prompt for the image to be produced for background/atmosphere only: [light, surface, color, environment]. The shadow direction and light hardness of the product should be [..] so that the composite is convincing.

Weak prompt / Strong prompt

Weak: "Nice portrait photo."

Strong: "Middle-aged farm woman, sun-worn face, slight smile, in a wheat field; golden hour light from the side, low contrast; 85mm portrait view, shallow depth of field, background soft; close-up, subject left, rule of thirds; warm earth tones, calm and dignified mood; natural skin texture, light film grain, 3:2. Photographic, realistic. Adding text."

The weak will is left entirely to chance; Since strong demand describes the light, lens, composition and atmosphere, it greatly increases the probability of producing the desired result.

Common mistakes

  • Not describing the light and lens. These two are largely what makes a photograph a photograph; If you leave it blank, the result will be left to chance.
  • Having the AI ​​draw readable text/logo. The model cannot produce these reliably; It may turn out to be defective and copyrighted.
  • Presenting the produced image as a "photograph". It is misleading in a documentary/news context; A statement of fact is required.
  • Copying style with person/brand name. Imitation and copyright risk; Describe aesthetics with qualities.
  • Skipping verification. Hands, eyes, reflections, shadows, and repeating textures must be checked in every produced image.

In summary

Generative visual AI works by diffusion: it goes from noise to image that fits the text, cannot produce what is not described, and errs where logic is required. For photographic realism, describe the subject, light, lens, composition, atmosphere and technique like a photographer. Choose the aspect ratio according to the purpose, reduce the artifact with negative recipe, do not have the AI ​​draw the text, and verify each output. Never use a produced image claiming to be a "photo taken".

Application task

Choose a concept and write a photographic prompt, filling in all six components with template 1. Produce (or describe) the same concept first in a single sentence, such as "a beautiful photo," then with the full skeleton, and compare the two results. Mark at least three possible flaw points in the image you produce by enlarging hands, eyes, reflections and shadows.

checklist

  • [ ] In the prompt, I described subject, light, lens, composition, atmosphere and technique separately.
  • [ ] I determined the aspect ratio according to the intended use.
  • [ ] I reduced the common defects with the negative recipe.
  • [ ] I did not have the AI ​​draw readable text/logo.
  • [ ] I made sure that I did not present the generated image as a "photo".
  • [ ] I verified the output for hand/eye/reflection/shadow.