back
137 comments
Passing the Turing test requires, for any human judge:

1. This judge is aware that they have to discern whether the 'bot' is real or a machine.

2. The judge cannot discern whether the 'bot' is real or a machine better than random chance.

This failed 1. And even given that advantage, might have failed 2 as well?

Often I see headlines along the lines of "X fools people and beats the Turing test!". But the point of the Turing test isn't to trick a person, it's to make it functionally impossible for a person to distinguish between the real and simulated thing, no matter how hard they try. For something to pass a Turing test, it would need to be able to pass the following:

"Anyone can play as judge any number of times. You can take as long as you want, and if you're successful in breaking it under controlled conditions (IE, you don't cheat and use an out-of-band communication protocol with the 'bot'/human), we'll give you a 10,000,000$."

The "One Million Dollar Paranormal Challenge" (https://en.wikipedia.org/wiki/One_Million_Dollar_Paranormal_...) is a solid example of a Turing test for magic.

I mean, you can read Turing's own definiton if his test - 'the Imitation Game' - on the first page of his 1950 paper Computing Machinery and Intelligence[1]. There's nothing in there about repetition, duration, or $10,000,000 prizes. It's a party game. And he just frames his question (which will "replace our original, 'Can machines think?'") as "Will the interrogator decide wrongly as often when the game is played like this [with a human and a computer] as he does when the game is played between a man and a woman?"

So, to perform the experiment, one must have some people play the game with humans a few times and then play with a human and a machine a few times, and look to see if the results are statistically significant. When they aren't, Turing posits, the question 'can machines think?' will have been answered in the affirmative.

That is not to say that this DALL-E vacation photo social media post constitutes a rigorous 'passes the Turing test'. But I don't think it's fair to criticize someone for using 'the Turing Test' colloquially as a catchall for saying 'you probably didn't notice this output was machine generated, therefore you might want to adjust your priors on the answer to the question, "can machines think?"'. Because that's exactly the spirit that Turing was working in when he proposed using a party game as a test of intelligence.

[1] https://www.csee.umbc.edu/courses/471/papers/turing.pdf

The test is now live all the time.

You must constantly be aware that images, or text, or voice, or other audio, or other signals or data, might be computer generated or altered.

All the time.

And you individually, or those about you, or societies at large, may be influenced in large or small ways by such signals, patterns, and records.

Your elderly neighbour or relative might be scammed out of life savings. Investors of false product claims. Voters of some fake outrage --- particularly of the October Surprise variety. Soldiers and diplomats of mock attacks, or false depictions of a tranquil situation where in fact danger lurks.

The test never ends.

This is your final warning.

Furthermore, the judge is supposed to be an expert. That is, not only his task is explicitly to tell between human and computer but he has to have a good idea on how to do it. Random people from the internet are not enough.

In the "paranormal challenge", the juges usually include stage magicians, because they know the tricks that would fool ordinary people. James Randi himself is a magician.

I think another important factor here is that it is unclear if the OP cherry picked photos or used the first ones given. Dall-E 2 has a bias to be better at real world scenes since it can just pull (near) memorized images, but I also wouldn't be surprised if these images were downselected from a larger set.
If someone told me they found a magic lamp and one of their wishes was that every time someone misused the term "Turing test" they got smacked in the face by an opening door I'd think "not bad, not bad".
As usual, the headline is more sensationalised than the actual article

> It's likely that with a harder version of the Turing Test, in which real and fake images of the same content are presented side by side and people are told that one of them is fake, it would be much easier to detect the fake images.

How often have we run a Turing Test, where we asked the judge how confident they were in their final answer, except both participants were humans?
if a human participates in interactive gamified social media, and this participation begins to change, shape, reinforce or otherwise mutate their beliefs, for the purposes of the test are they still actually a human? could the entire social media mechanism (from the builders to the participants) be considered a form of a sort of singleton autonomous intelligence in and of itself?
More accurately, “DALLE2 made me realize no one cares about or looks closely at your vacation photos”
The snarky side of me wants to say it's because people take boring photos

But really it's just information overload, most things on social media I just scan the thumbnail and move on. Only my family would care to see my vacation photos :D

One, or three, carousels of vacation slides used to be an effective way of putting a party to sleep ...
My rule of thumb: never share more than one photo per vacation unless asked for more.
Diver here. I'm looking at these four pictures after the fact, so I already know they're fake. They're good, but they also have some weird flaws. That said, I don't think I would have immediately recognized any of these as wrong on Facebook (maybe the diver photo).

- The nudibranch (slug thing) on the green coral doesn't look like anything I've seen in the Caribbean before, and the coral also looks odd for the region. That said, this is probably the most difficult photo for me to differentiate. I would have accepted this as a cool find of something I haven't seen before.

- The grouper (big fish) photo is actually pretty good, although DALL-E has misplaced its eyes a bit. That said, the lit foreground and dark noisy background are exactly the look I would expect for someone using a basic camera + lights with wonky post-processing.

- The diver photo is a horror show. There's a hose going nowhere on her back. It looks like she's blowing out of a harmonica instead of a regulator. Bubbles are collecting around the top of her mask for some reason. Her fins look like they were badly Photoshopped. Nothing looks right here.

- The lobster photo has a real but subtle flaw: Caribbean lobsters don't have big claws. It also looks like it's under a rock like you would find in cold waters around Massachusetts and Maine instead of the Caribbean.

Interesting stuff though. It will force me to be more skeptical when I look at people's photos in the future.

This is off topic but horny people are by far the most interested in conjuring up custom images. DALL-E trained on porn would be huge.
This is absolutely true. Look at text prediction models as an example (e.g. GPT-3). One of the biggest (if not the biggest) applications was story-generation tools like AI Dungeon. Guess what most people actually used AI Dungeon for? Erotica. Guess what happened when OpenAI cracked down on it? A huge portion of the userbase jumped ship and built a replacement (NovelAI) using open-source EleutherAI models that explicitly did support erotica, which ended up being even better than the original ever was. I can tell you that there is very strong interest in nsfw image generation in those communities, as well as multiple hobby projects/experiments attempting to train models on NSFW content (e.g. content-tagged boorus), or bootstrap/finetune existing models to get this sort of thing to work.
This is one of the experiments with Progressive Growing GAN (ProGAN) technology from Nvidia:

NSFW: https://medium.com/@davidmack/what-i-learned-from-building-a...

This might be horrible to say, but could this be a solution to csam? From what I've seen most people who enjoy csam do genuinely feel bad for the children, but they're sick, and can't control themselves. Might they be willing to indulge in fake csam instead?
Porn seems to quietly power the Internet, in so many ways. I imagine people are already getting creative with fake porn, and it's only going to intensify over time. Especially on the types of imagery that are illegal to possess.
It's not a Turing test if a judge is not actively trying to discern.
The images look great but I think the experiment was helped by the fact that most of us can see any image, being told that the image is underwater and we'll believe it. Most of the people don't know about deep waters, and the plans and animal that live there. I suspect the experiment would have been different if the vacations were in a city or a beach or something like that.
Wouldn’t a proper Turing test be one where people knew some of the photos were artificial and were asked to figure out which ones they were?
One cool application of DALL-E could be generating a painting or sketch for each paragraph or sentence of novels. Imagine listening to an audio book of famous novels with visuals/cartoons made by AI.

Hope someone with invite access could do this for Moby Dick or Sherlock Holmes stories or 1984.

This is a flawed experiment. If I see a bunch of photos, and many of them look real at first glance, I’m not instantly going to critique whether all of them were real, unless I was given specific instructions to do so.

Also, underwater photos are not something many people have personal experience seeing. Most of us don’t live underwater. We may not be equipped well enough to tell the difference, where above water, especially urban photos, we will likely notice better.

> My deepfake DALL-E 2 vacation photos passed the Turing Test

Most people didn't notice that some of my vacation photos were fake, therefore it passed the Turing Test... why is this clickbait nonsense getting so much attention?

Can someone who upvoted this article explain why you upvoted it? Did the fact that the title is flatly false not bother you? If someone wrote an article about cracking some encryption algorithm and titled it "I proved P=NP" would you upvote it?

It is easy to pass Turing tests when the subject material is unfamiliar to people. As other posters have mentioned, most people have only a vague idea of specific underwater plants and animals and vague ideas of how the water distorts light.

I bet I can come up with a simple generator that generates galaxies/nebula pictures and if I interspersed those in with NASA Hubble generated images, most people could not pick out the real Hubble images from my generated images.

I'm wondering if we'll ever get to a point where we can invoke fake vacation / travel experiences, like We Can Remember It for You Wholesale (more popularly, Total Recall), by creating ML-generated images of the trip rather than inducing a dream. It seems plausible.
That The Fifth Element (1997) scene linked in the article actually holds up well.
I want to see a fully-synthetic multimodal social media influencer that is nearly indistinguishable from reality. She does the same thing as the real ones except everything is completely artificially generated (housing, clothes, vacations, social circle). All text/image/video posts are completely synthetic but internally-consistent with this fabricated persistent universe. The only real things would probably be product placement. If you’re a brand, you’d just make a new online influencer instead of finding an organic one.
> It still struggles with faces

I believe they intentionally hobbled it in this respect for "safety" (iow to keep themselves out of a scandal when someone asks it to create "President Biden accepting Bribes" or whatnot...)

Certainly far simple diffusion models trained including faces do just fine at creating photorealistic faces.

I think the Turing Test is more than that?
The images look great because only the last 4 pictures on this blog are the fake ones.

All the first impressive looking shots at the top of this article are real.

Fun idea. We're about test out melodies I've created together with a generative neural net and see how they're rated computer to real melodies. The plan is use Amazon's Mechanical Turk but one problem is what to do about real melodies that people will mark as already familiar to them. I think making it a comparison with unfamiliar melodies should be fine?
Is DALL-E still invite-only?
Hacker News has this surprising tendency to cling to the past.

People here are nitpicking over the definition of the Turing test. What actually matters here is that, if not already now, but certainly in 1-5 years neural nets will most certainly be good as the 99th percentile artist.

Does that mean AGI is here? Probably not. But we are missing the forest for the trees.

It's kind of an alien world to us, no lighting like we know it, all organic shapes with a lot of unidentifiable stuff, blue tint etc. It all helps to make it an easier case.

dalle-e is still impressive, but taking this to the extreme it would be like making it simulate pictures of TV noise and show we couldn't tell it from the real thing.

>Could I use DALL-E 2 to create a fake vacation? Or, more ethically, could I use DALL-E 2 to recreate events from my vacation that actually happened, but I was unable to get a good photo of them?

What would be unethical about creating a fake vacation? As long as you're not defrauding anyone, I don't see who would be hurt by this.

Your holiday photos are memories. If you create fake images, and mix them with genuine ones, don't be surprised if in the future, you yourself forget what is real and what is not.
Just nitpicking but the 'Turing test' can only be failed, not passed, which is quite apt given another problem associated with Turing: the halting problem.
How will our legal systems cope if one day you can conjure up any "evidence" you want? We're still safe today but the future will be scary.
Every marketing company making money off influencers and organic content like this just silently screamed.
It's not a Turing test if you don't tell them that it's a test.
Shocking—photos made out of a data set of existing photos look like photos!
So when are you vacating on the Moon or better still, can you beat Elon Musk to Mars?
You don't need DALL-E 2 to make those photos, you can just download generic underwater photos from other people.

In fact how do you know DALL-E actually created them, and did just regurgitate some it was trained with?