Prompt to Picture: Creating Images with AI
🔒 This course requires registration
To access this course and all our learning materials, please register for the AI Fluency programme.
Register Now →Welcome to Prompt to Picture: Creating Images with AI, a seven-lesson course that will teach you how to turn text prompts into compelling, professional-quality images.
The course is presented by Enrico, a multidisciplinary designer from Operating Systems Studio, part of the cultural association Umanismo Artificiale. Throughout this program, he’ll guide you through the tools, techniques, and creative strategies that will allow you to develop a medium-specific craft across photography, illustration, and other visual forms.
This course is about more than learning to use AI models, it’s about building a clear mental model of how text-to-image systems work, understanding prompt structure, and developing the skills to design cohesive, intentional image series.

Course Objectives
By the end of this course, you will be able to:
- Build a mental model of text-to-image systems and how they generate visuals.
- Learn and apply a solid prompt structure across mediums.
- Develop medium-specific craft for photography, illustration, 3D, and hybrid techniques.
- Explore and practice a toolbox of image-generation techniques (prompt engineering, inpainting, outpainting, multi-turn workflows, etc.).
- Understand and apply node-based interfaces (e.g., ComfyUI) for modular, professional workflows.
- Design a cohesive image series, not just single outputs.
Course Structure
The course is divided into seven lessons, each building toward a final project:
- Models Landscape – Overview of major text-to-image platforms and creative case studies.
- How AI Builds an Image – Diffusion models, multimodal models, training vs. generation.
- Prompt Structure – Anatomy of a prompt and how to compose effective instructions.
- Prompt Components – Medium, subject, action, environment, lighting, mood, technical details.
- Technique Toolbox – Advanced techniques: img2img, inpainting, outpainting, multi-turn, modular prompts.
- Node Interfaces: ComfyUI – Visual workflows, modular editing, advanced control.
- Workflow & Final Project – From concept to delivery; create a curated series of at least 5 images.
Final Micro Project
Your final assignment will be to design a cohesive image series (minimum 5 images) that integrates concept, style, and technique. The series can take the form of photography, illustration, or mixed media, but it must show intention, consistency, and your unique creative voice.
Frequently asked questions
How does a text-to-image model generate an image from a prompt?
The model uses a diffusion process to transform random noise into a coherent image. It iteratively removes noise based on the text embedding, which guides the visual structure. This happens over hundreds of steps, gradually revealing details like shapes, colours, and textures that match the input description.
What is the difference between Stable Diffusion and Midjourney?
Stable Diffusion is an open-source model that runs locally or on cloud servers, offering full control over parameters. Midjourney is a closed, cloud-based service accessed via Discord. While Midjourney provides polished aesthetics out of the box, Stable Diffusion allows for custom fine-tuning and integration into complex workflows.
How important is prompt structure when creating AI images?
Prompt structure significantly impacts the final output quality. Placing the most important subject at the beginning ensures the model prioritises it. Using clear delimiters, such as commas, helps separate distinct concepts. A well-ordered prompt reduces ambiguity, leading to more accurate and consistent results across multiple generations.
What is ComfyUI and how does it differ from other interfaces?
ComfyUI is a node-based interface for Stable Diffusion that visualises the generation process as a flowchart. Unlike traditional web interfaces, it allows users to connect specific functions, such as loading models or applying filters, in a custom sequence. This offers granular control over the workflow, making it ideal for complex, multi-step image generation tasks.
Can AI image generators create consistent characters across multiple images?
Yes, but it requires specific techniques to maintain consistency. Users can employ reference images, seed values, or character-specific LoRA models to keep facial features and clothing identical. Without these tools, the model treats each prompt independently, resulting in different appearances for the same described character in every new generation.