Because that's not true at all. The AI can't even draw hands yet. To say nothing of its ability to handle multiple people and objects interacting in complex scenes.
I'm concerned that the discussions about AI art on forums like HN get distorted because you have people sharing their views on art here, even though they don't actually have a serious and nuanced appreciation of art and they don't have a good understanding of all the types of work that artists do. Maybe you'd be fine with reading a comic book where everyone has seven melting fingers, but people who take comic books seriously as an artistic medium would not.
This seems to be purely an issue of the size of the network. Parti (https://parti.research.google/) demonstrates that as the number of parameters increases, with no change to the underlying architecture, a lot of these problems simply go away. Basically just throw more compute and memory at the problem and everything gets fixed.
This is not to take away from the achievements of AI. It's that creating pictures adhering to a prompt with some degree of creativity is very little of what art is. Maybe it will replace some part of commissioned illustrations where the artist's name does not matter (e.g. some avatar pic?).
We still value, financially, some material goods for much more than they cost to produce. Or for much more than their almost identical mass-produced counterparts.
I mean, it's the majority of commercial art - you get a prompt from the client, you maybe flesh it out in a few different directions with sketches, then you refine a final piece. And AI is incredibly good at this process - instant results, infinite patience, and it's free. A very hard combination to meet.
Calendars, book covers, video game assets, green screen backgrounds....
Even in a case like video game animations, where the AI can't build every frame, it can still give you a good reference photo. From there you just need a cheap artist to fill out the frames - a huge cost savings, and a big blow to the artistic community.
Where do you get started as an artist, without any of those? Obviously, Fine Arts isn't nearly as effected, but how do you get your start when you can't build a name from your cool book covers, or get famous off Magic: The Gathering card illustrations?
The 20B images don’t look that much more impressive than what SD is already doing (aside from the ability to render text), and in some cases they look worse. It’s hard to tell because the resolution is so small, but even in the 20B “astronaut riding a horse through a pond” image, it looks like his hands are still nonsensical.
>the tech scales so well as to make their new goalpost irrelevant in a year.
This just brings me back to my original question. Self-driving cars have been "a year away" for many years now, and now companies are starting to hint that human assistance may be required for the foreseeable future [1]. So, why the confidence that art will be an easy problem to solve with just more scaling, when that approach hasn't eliminated the need for humans in any other domain?
[1]https://www.reuters.com/technology/truly-autonomous-cars-may...
(This will be great from the point of view of art creation; not so great from the point of view of supposedly rendering humans obsolete)
AFAIUI, that’s part of the point that Gebru was trying to make before she was fired.
With an order of magnitude more parameters it won't just do hands, it will do quite a bit more.
Edit: to tie it back to the original criticism: what’s the maximum training cost we’re willing to accept for the model? How can we guarantee return greater than the increased training cost?
And training a real living breathing person in a rich OECD country is going to be costly - no offense meant, I'm actually not from OECD.
But then again, when it comes to commercial needs, a human doesn't need "retraining" every time you ask them to draw something they weren't familiar with when they went through art school...
In the smartphone age, the case for data hunger looks pretty weak.
Artistic styles are often just thin semantic filters over the base 3d geometry that can be learned from photos, and learning these shouldn't require many examples.
Technology is making something that used to take a lot of practice and skill be accesible to those without any of it. A monkey can now draw two ovals, label it an owl, and run an image-to-image conversion with Stable Diffusion to get a pretty good sketch of an owl [1].
Is it better than what a good artist could do? Irrelevant.
Is it better than what a cheap illustrator I find on Fiverr could do? Irrelevant.
The only important point is that I no longer need an illustrator to get myself an owl. I draw some lines, I pick some words, and presto I have an illustration.
The question of whether it's "art" is entirely irrelevant.
> Are you under the impression that right now, as of today, the publicly-available AI models are ready to replace humans for all types of art outside of scientific and technical illustration? Because that's not true at all. The AI can't even draw hands yet. To say nothing of its ability to handle multiple people and objects interacting in complex scenes.
I think this is severely underplaying the speed at which things are changing and basing an argument about things that the AI currently can't do. DALL-E was anounce in Jan 2021 and it's still locked behind API access. Stable Diffusion came out Aug 2022 and I can run it on <$2,000 laptop. That's not 2 years. Do you think hands are going to be a long term roadblock?
As for complex scenes, you can currently string that together with a Stable Diffusion plugin for photoshop/gimp.
[1] https://www.reddit.com/r/StableDiffusion/comments/wwv7zk/sta...
Now, this may actually be helpful in that it gets around copyright claims - but that's the only real difference.
But the approach for stable diffusion is just as easy whether you want just "an owl", or "an owl in X's style with A, B, and C"
Now, I should of course note that search engines already employ ML techniques to actually interpret search terms, so to some extent the point is moot - ML is important to actually solving this problem.
And of course, it's also possible that the image I want can't be generated by SD/DALL-E/etc.
Without meaning to sounding rude about it... I'll wait.
I tried generating that exact prompt a few times at theartbutton.ai and all the results were nonsensical.
For example: https://theartbutton.ai/image/OW1HZLfhjg6DFvJtk4vQZzUYqI7pGG...
Not a very wide range of what I could do with the idea in terms of composition, but just some variations of finishing touches/intermediate steps. I achieved this with some human-in-the-loop iteration and inpainting, but it was no more than 15-30 minutes toying around with it, and I'm no artist.
If you have a semi-decent graphics card and would like to experiment with a bunch of extra settings and tools than are readily available online, this is a good repo for that: https://github.com/AUTOMATIC1111/stable-diffusion-webui
I would extend this lack of nuance and understanding to the deep learning implementation side also. A lot of people seem to have some very foundational misconceptions about what deep learning is and what it does. In the case of generative art: these models are “simply” sampling from the frozen statistical structure they have learned from web images and their captions. They don’t understand the relationship between objects in space, they have no ideas or feelings to express, and they communicate nothing. That’s why the even the best output of these models tends to have a perceptible hollowness: you can detect the lack of a coherent authorial intent in it.
The field of "art that needs human communication skills" seems to be a lot broader than just scientific illustration.
SD obviously doesn't understand language in the same way we do, so it can be tricky to describe things in a way that will match your expectations. Once you start to understand the tricks here, it gets easier and easier.
Inpainting will let you fix a lot of the rest. Staircase stops? Select the area where it stopped, get the AI to generate more. People are already doing this to create very complex artwork where there are issues with faces, hands, etc. https://www.reddit.com/r/StableDiffusion/comments/x9u8qh/img... is a great example of how you can quickly iterate over a scene.
One of the other things people struggle with is consistent characters and settings, but people have found ways to improve this with Midjourney - https://docs.google.com/document/u/1/d/e/2PACX-1vRahIr3-h_V3...
There's more of a learning curve to these tools than most people think, but it's also still miles and miles away from the learning curve required to actually be proficient at the technical aspects of making art.