back
120 comments
For those who think it's just another lame DL based instagram filter...

The method proposed in the paper(https://arxiv.org/abs/1707.03491) is mimicing a photographer's work: From taking the picture(image composition) to post-processing(traditional filter like HDR, Saturation. But also GAN powered local brightness editing).In the end it also picks the best photos(Aesthetic ranking)

Selected comments from professional photographers at the end of paper is very informative. There's also a showcase of model created photos in http://google.github.io/creatism

[Disclaimer: I'm the second author of the paper]

My ex-wife is a professional photographer. She does weddings, boudoir, and portraits mostly.

Here's the thing. If she had say, 2 shoots in a week, maybe a total of say 8 hours of shooting time, the selection, editing, and styling of those photos could take 3x or 4x as long as the time spent shooting. She would have to be editing constantly.

I've always thought a perfect application for this sort of technology would be a model trained on that particular photographer's style, which goes through a batch of photos and selects the best candidates and then presents the user with a few choices for each photo it selected and styled.

I think the ultimate would be a NAS with this capability embedded. The photographers I have met through her over the years seem savvy about photo tech, but don't generally use online storage solutions or know much more than how to admin a WP site at best. SOHO solution would be ideal imo.

Great work! I'm excited to see where this goes.

I thought of a cool application . You could use this to generate desktop photos for desktop environments that are either localized to a place you are near or the user could type in a location that they want to see .
Cool- Are you going to apply it to all images on Google Maps? How about an 'Auto Enhance" button on Google photos?
Did you have anyone say "Professional photographers don't apply HDR filters to single photos"?
Will this ever get implemented in Google Photos automatic editing?
Apologies for being that guy...

Was there any control for, e.g., the dynamic lighting filter being the factor that carried most of the water, rather than the other manipulations?

FWIW, I did think your tool did a good job finding good compositions within mundane scenes. Looking at the originals, I have to admit that in many or even most of the cases, I wouldn't have spotted the opportunity - which points to an area I can improve my craft.

Still lame, just cropping and increasing saturation/contrast. Some of the pictures in the blog post look worse than the originals, they are artificial.
The images look really good.

I'm curious: Is this creatism link what the algorithm considers the best images, or did you manually filter out some of the less good ones?

Thanks! Interesting work, indeed. I found the showcase photos appealing and nice. I briefly skimmed through the paper with hope to find some interpretation of the “aesthetics”, but couldn't find it. What is meant under “all aesthetic aspects” and “aesthetic quality”? Did you perform a critical analysis of the network? How do you know that the network learned “aesthetics”?
I tried a style transfer network on one of your images . Here is what I got https://prisma-ai.com/p/3a91d082-feee-443e-8062-4e85d2720ea0
Is the code available?
on the pro side: you could use this to determine the most photogenic scenes in the world.

On the con side: only an AI photo-bot can find and capture the most photogenic scenes in the world.

When a topic like self-driving vehicles comes up, the Hacker News crowd is mainly in favor: “Creative destruction! Disruption! Go go gadget robots!” Not surprising. How many Hacker News readers drive trucks or taxis for a living? How many regard commuting as an enjoyable hobby?

Photography, on the other hand, is a very common hobby in the tech community. And the comments here seem to reflect that this effort strikes a little close to home: “Those pictures are lousy, if you find them appealing you have no taste! Just because they're 'professional' doesn't mean they're good! Machines can’t replace human judgment, they have no soul! I bet that machine had a lot of human help!”

Tech people may tell you great stories about meritocracy and reason, but in the end we are just emotional monkeys. Like the rest of humanity.

Those of us who can accept this may at least aspire to be wise monkeys.

Ha, I'm one of those techies you're talking about. Spent an embarrassing amount on camera gear, and I'm always "that guy" with the ridiculous rig at someone's casual party.

But you know what? I love this shit. I like photography for the good photos, not because I'm building my self esteem on top of it. I want to capture the scene or the moment in a way that, IMO, does it justice. If this technology makes that easier, and gives it to more people - great. Progress.

And what I really love is how it subtracts a huge amount from the price of entry. It's like when software synths came out and made it possible to make music without $10k+ worth of MIDI hardware. What followed? An explosion of creativity which I've been luxuriating in ever since. Get the ability to create beautiful images in as many hands as possible, says I. For me, at least, "more beautiful images in the world" was the whole idea.

There is no such thing as a wise monkey. Take everything you said and remove the possible upside, and that is the truth.
Talking as a semi-pro (I've put in some money into cameras and lenses and spent a good bit of time on photo editing), this is a bit underwhelming. For landscapes (which this seemed to focus on), I've found that opening up the Windows photo editing programs and clicking 'enchance' or Gimp and clicking some equivalent already gets you most of the way there in terms editing for aesthetic effect. The most tricky bit is deciding on the artistic merit of a particular crop or shot, and as indicated by the difference between the model's and photographer's opinion at the end of the paper, the model is not that great at it. Still, pretty cool that they did that analysis.
The other huge thing that is lacking is content. Image processing is only 10% of good photography. The rest is about conveying an idea to your audience. It can be humor, documenting an event, making people aware about societal problems, taking people to somewhere they would normally not be able to access, or seeing the world in a way that most people don't see.

Truly good photographers don't just produce beautiful photos; they produce meaningful photos.

The tech is great but I'm not a fan of the title of the article ;)

Automatically selecting what portion to crop is impressive, but just slamming the saturation level to maximum and applying an HDR filter is the sign of "professional" photography rather than good photography.
As someone who lives in a relatively rural area with similar geography to much of the mountains and forests in these pictures I have noticed previously how professional pictures of these areas have a similar feeling of over saturating the emotion.

It's interesting to see algorithms catching up to being able to replicate this. However when you mention these kind of abilities to photographers, they get defensive, almost like you are threatening their identity by saying a computer can do it.

It is an interesting project and shows significant accomplishment. I'm not sold on the idea of "professional level" except in so far as people getting paid to make images. I am not sold because the little details of the images don't really hold up to close scrutiny (and I don't mean pixel peeping).

1. The diagonal lines in the clouds and the bright tree trunk at the extreme right of the first image are distractions that don't support the general aesthetic.

2. The bright linear object impinging on the right edge of the cow image and the bright patch of the partial face of the mountain on the extreme left. Probably the gravel at the left too since it does not really support the central theme.

3. The big black lump that obscures the 'corner' where the midground mountain meets the ground plane in the house image.

4. The minimal snow on the peaks in the snow capped mountain image is more documenting a crime scene than creating interest. I mean technically, yes there is snow and the claim that there was snow would probably stand up in a court of law, but it's not very interesting snow.

For me, it's the attention to detail that separates better than average snapshots from professional art. Or to put it another way, these are not the grade of images that a professional photographer would put in their portfolio. Even if they would get lots of likes on Facebook.

Again, it's an interesting project and a significant accomplishment. I just don't think the criteria by which images are being judged professional are adequate.

I don't know why but the "professional" label on this really irritates me. I'm curious to know how the images that got graded on their "professional" scale were selected for inclusion in the sample. Surely by a human who judged them to be the best of many? I'd love to see the duds.
Very impressed by the results.

I hope that one day our driverless cars will alert us when there is a pretty view (or a rainbow) so we take a moment to look up from our phones. Every route can be a scenic route if you have an artistic eye.

Interesting how hi-res the photos of a small section of Google Street Car photo can be compared to what users see online; here's an example from the linked article:

https://2.bp.blogspot.com/-6bVWUgA8NEI/WWe1uoW8ayI/AAAAAAAAB...

When a photographer takes or edits a picture, she doesn't need to predict or simulate her own reaction. There is no model or training necessary, because the real outcome is so easily accessible. However, she is only one person, and perhaps can't proxy well for a larger group.

The model has the reverse situation, of course: it cannot perfectly guess the emotional response for any one person, but it has access to a larger assortment of data.

In addition, in different contexts it may be easier/cheaper to place a machine vs. a human in a certain locale to get a picture.

If my theorizing makes any sense, it suggests that this technology would be useful in contexts where: the locale is hard to reach and the topic is likely to evoke a wide variety of emotional responses.

Retouching is another field to play with - I am experimenting with CNN/GANs to clone styles of retouchers I like. If you are a photographer, you know that most studio photos look very bland and retouching is what makes them pop; for that everyone has a different bag of tricks. If you use plugins like Portraiture or do basic manual frequency separation followed by curves and dodge/burn adjustments, you leave some imprint of your taste. This can be cloned using CNN/GANs pretty well; the main issue is to prevent spills of retouched area to areas you want to stay unaffected.
"Someday this technique might even help you to take better photos in the real world."

So what? Maybe I missed it, but what are some potentially meaningful applications of this technology? What motivated this to begin with? Or are these questions that we even bother asking anymore?

I remember the first time someone showed me the Snapchat app -- it would make them look like a cartoon dog, or all these other real-time overlays. I thought, 'jesus, so glad we're all getting advanced computer science degrees so we can work on utterly useless shit like this...'

this is amazing, but 'professional photographers' aren't really the best arbiters of what a 'good' photograph is. Also, training on national parks binds the results to a naturally bland subject, no pun intended. While an amazing achievement, nothing shown here demonstrates ability beyond a photographer's assistant/digital tech adjusting settings to a client's tastes in Capture One Pro. Jon Rafman's 9 Eyes project comes to mind as something that produced interesting photographs, as does the idea to find a more rigorous panel of 'experts' (e.g. MoMA), or training the model on streets/different locations than national parks.
Related: Arsenal (https://www.kickstarter.com/projects/2092430307/arsenal-the-...) is trying to build a hardware camera attachment that uses ML to find the perfect levels for your photo in realtime.
This is cool but I really don't get why one could call this actually creating "Professional-Level" photographs. It's more like a very good auto-retouch. There's still the matter of someone actually being there, realizing it is a beautiful place, and dragging a large camera with them and waiting for the right light.
I think some of these results are really lovely, the one at Interlaken is a perfect travel photo. Would be interesting to see more types of work this could apply to.

Saw a few people talking about retouching and studio work - I do a lot of studio shoots and retouching on my own, and would be happy to help or participate in projects. Feel free to reach out.

The first thought after going through all these photos was: incredibly stilted. It's amazingly impressive, but the human photographer will always be able to capture the subtleties that AI will miss. But very cool nonetheless
Instead of augmented reality I would call this "distorted reality". People will prefer to visit places with Street View than being there. Real reality is uglier
Up to what point can the output be controlled? Can complex conditions be created? e.g. a lake with a mountain background during the evening
Is deep learning comparable to perceptual exposure?
In the future maybe we can just hook up a drone to this and have it fly around taking nice pictures.
I find the colors in the results images consistently worse than in the original images.
ML = Wisdom of Crowds
Would be interesting to see how well you could train this kind of thing off of a large catalog of lightroom edit data. to then mimic a specific editors style.
For example, whether a photograph is beautiful is measured by its aesthetic value, which is a highly subjective concept.

Oh really.

[deleted]
wow automation isn't leaving any fields untouched
Lately, there has been lots of talk of deep learning applied to create tools which can generate requirements – designs – software code – create builds – test builds as well help with deploying builds to various environments. I'm excited for the future developments capable with ML.
If they're doing dodging/burning, then they could really use the processing on raw files instead of jpegs. The dynamic range is obviously limited when dodging/burning jpegs, as you can see from the flat clouds and blown highlights on the cows.
Great, not all we need is specialized machine learning inference accelerators in our mobile phones. I wonder if Google has even considered making a mobile TPU for its future Pixel phones.
From the article the caption of the first picture was interesting: "A professional(?) photograph of Jasper National Park, Canada." Is that the open scene from The Shining? If so I wonder why the question mark, is Stanley Kubrick not a professional photographer?