At a base level, "putting some prompts [..] and choosing" sounds like intent to me. A user is choosing what they want to capture and making a choice, much like I, as a photographer, look for a particular scene (or happen across one!) and choose how I want to capture it.
Of course, you can do much more in terms of control of image synthesis these days. ControlNet and friends make it feasible to control every aspect of the resulting image.