Even if we take it as a given that a test like this can work well, I suspect that test takers will become more sophisticated at choosing images to get desired outcomes.
And so the test will need to get more sophisticated in turn by finding even more non-obvious yet discriminating pairs of images.
This is sort of like the constant battle between those creating spam filters and spammers.