back
111 comments
I read the article. It barely scratches the surface of the issue, I was mad as a psychology student that this was the state of our knowledge when I read it in 2013 and didn't know about the publication crisis. The actual academic publication is really interesting and I'd recommend it as a read on whatever commute you're on.

Here is the link:

https://www.ssoar.info/ssoar/bitstream/handle/document/42104...

Years after pondering this article, I felt I had figured it out: psychology is secretly the study of particular subset of 18 to 21 year old American women. The ones who study psychology because it's a female dominated study and all those psych students need their credit and 'participating' in research is part of it. Most of them are American because psychology is a bigger thing in the US than in Europe, or at least it seemed to be regarding well-known theories, so I presume most research happens there.

There is another big group. A lot of dead mice (neuroscience).

Do you have any studies tracking the sex of study participants in psychology? I know that in clinical trials, historically the majority of participants are male:

"After the tragedies caused by the use of thalidomide in pregnant women, the FDA issued “General Considerations for the Clinical Evaluation of Drugs” in 1977. This guidance document stated that women of child-bearing potential should be excluded from Phase 1 and early Phase 2 research, except if these studies were being conducted to test a drug for a life-threatening illness. If a drug appeared to have a favorable risk-benefit assessment, women could then be included in later Phase 2 and Phase 3 trials if animal teratogenicity and fertility studies were finished...In 1993, FDA reversed the 1977 guidance with another guidance document entitled Guidelines for the Study and Evaluation of Gender Differences in the Clinical Evaluation of Drugs."

Also:

"In a study that evaluated the inclusion and analysis of sex in the results of federally-funded randomized clinical trials in nine major medical journals in 2009, researchers found most studies that were not sex-specific had an average enrollment of 37% women."

From "Women’s involvement in clinical trials: historical perspective and future implications" - https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4800017/

> Years after pondering this article, I felt I had figured it out: psychology is secretly the study of particular subset of 18 to 21 year old American women.

There are many analogs of that in Computer science too. For example in computer vision, CIFAR-10 is a de facto standard for measuring performance. A set of 60k 32x32 images. Good results on that data set doesn't necessarily translate well into real-world performance. But what can you do? Gathering huge data sets and having humans annotate them is incredibly expensive.

Another ml example is music recognition. There are several state-of-the-art methods for detecting notes in music. So one Chinese researcher tried to apply them for Jingju music https://www.youtube.com/watch?v=NJsTl342RhI. The results were... less than stellar.

> psychology is secretly the study of particular

> subset of 18 to 21 year old American women.

This is perfect!

I had a similar epiphany about linguistics, much of which appears to be the study of example sentences that linguists come up with.

Since you could invariably tell when something was an example sentence, this was obviously a different language from what people actually used. (Of course there are linguists who go out and study real language-use, but some actively dismiss this as mostly irrelevant "performance")

> Most of them are American because psychology is a bigger thing in the US than in Europe, or at least it seemed to be regarding well-known theories, so I presume most research happens there.

I'm not sure about the amount of research but the amount of psychology students here in the Netherlands is huge. And it seems to function in much the same way. Often first-years get some form of credit for mandatory participation in experiments for higher-year students and sometimes regular faculty members. I'm also aware of some researchers going out of their way to get a different sample from the population but in the end it's a lot easier to get participants when you make them.

The thing with the reproducibility crisis is that a lot of psychology theories have the "looking under light post effect" (a drunk looses his glasses in a river and a police officer finds looking under a light post, "because at least I can see here). It's been well known for centuries that human motivations are complex, socially layered and contextual. But most psychology theories wind-up fairly simplistic and not framed by multiple contexts because ... these can be easily tested (except apparently they can't be easily reproduced, uh..).

I don't know any easy way out here. Just because "folk psychology" can say formal psychology is too simple doesn't mean folk psychology and anecdote are more useful.

> The ones who study psychology because it's a female dominated study and all those psych students need their credit and 'participating' in research is part of it.

Wait, does participating in a study as a subject count for credit, or does working on a study count?

Aren't studies explicit about the demographics of their participants? And don't studies that make generalized claims usually control for things like gender and age?

> There is another big group. A lot of dead mice (neuroscience).

I'd add in Beagles, Capuchins, Zebrafish, and Mongolian Gerbils there too.

I believe a significant chunk of genetics is similarly limited, to Icelanders, for example.
I actually don't see this is as too problematic.

The entire field is murky because it's very hard to scientifically measure human psychology and behavior, and the tools of the trade seem almost laughably simplistic (like the aforementioned 5 point rating questionnaires). So our body of knowledge doesn't even reliably describe WEIRD people. Rather its a crude proxy, that just might contain some elements of truth warranting further investigation.

But better crude tools than no tools, and better locally available subjects than no subjects (because most studies don't have the budget to go to Zambia). Similarly in other fields, research is done on rats pigs and monkeys with conclusions drawn to humans. Obviously not perfect, but again it's at best a starting point for later studies.

I think the real problem is the over-zealous interpretation of study results as "truth".

I disagree.

If you're trying to measure a target 1cm across from 1 mile away with a ruler and a squint anything you say is not only likely wrong, but woefully deceptive.

My view is that this shouldn't even be attempted because it just generates superstitious theories. Psychology is the skinner box pidegon.

There's a fallacy here that's hard to pin down clearly but roughly: to measure inaccurately isnt to measure approximately. Its to measure totally in error.

The errors in social psychology are not just "second decimal place", they're angels pushing stars.

I concur. Psychology isn't hard science by any mans, and extrapolating it to people outside your subjects' samples across the entire world into other cultures and things is obviously ridiculous. But we can still use tools and find patterns to help people.

Surely, the solution isn't to throw up our hands in the West and be like "we can't find any help for the mentally ill because some subject might not apply to our discoveries from around here!"--but of course, if they are extrapolating those results to other cultures, that's worrisome and inaccurate in many studies.

Human behavior has some universals, but a lot we figure are universal aren't.

I think the 'better crude tools than no tools' claim is false. The history of the 20th century is littered with an embarrassing number of stupendous tragedies that can all be laid at the feet of a lack of scientific rigor. Eugenics, leaded gasoline, thalidomide, DDT, lobotomies, and many, many others all stemmed from people who felt that it was expedient to draw conclusions based upon either insufficient data or reasoning they knew was not iron-clad.

There will always be a conflict between what we can know with confidence and the decisions that we need to make to handle situations that arise. But it has been shown that having someone in the position of authority as a 'scientist' giving their approval to something based upon insufficient evidence or reason tends to often lead to the most severe and large scale suffering. So I'd have to recommend against giving any credence to any crude tool.

Would you rather know nothing about something, or know something about it that is probably wrong? This should be a rhetorical question as the answer is self evident. The psychology replication crisis started off when it was illustrated that only 36% of studies published in highly regard psychological journals could be successfully replicated. And, unlike psychology studies, that discovery can and has been replicated. One can only imagine the rate in less well regarded journals.

The point of this is that the average study in psychology is much more likely to be wrong than right. And so by indulging psychology you are not giving yourself a crude tool for understanding, you are actively misinforming yourself! Imagine I wrote a newspaper where 64% of the articles were fake or misleading. If you'd like to be as well informed as possible, you'd be better off never reading that paper, even if there are some true things in it.

Science is not a 0 or positive game. Bad science can and does send societies and progress backwards.

That's not the only issue.

https://en.wikipedia.org/wiki/Replication_crisis#Psychology_...

> A report by the Open Science Collaboration in August 2015 that was coordinated by Brian Nosek estimated the reproducibility of 100 studies in psychological science from three high-ranking psychology journals.[38] Overall, 36% of the replications yielded significant findings (p value below 0.05) compared to 97% of the original studies that had significant effects. The mean effect size in the replications was approximately half the magnitude of the effects reported in the original studies.

That makes perfect sense. The published studies are a sample of 'study-space' and that have outlier significance. Replicate published studies and their significance likely returns to the norm.

Journals are filters to cherry-pick 'study space'. By the way they're constituted, they publish new studies that have overstated significance.

Veritasium has a good overview of the replication issues in much pf modern science: https://www.youtube.com/watch?v=42QuXLucH3Q
Causes of the replication crisis puts 0 blame on academia. This is an Academia caused crisis.
> [...] something like “I generally trust people.” Then participants are asked to choose one point along a five- or seven-point line ranging from strongly agree to strongly disagree. This numbered line is named a “Likert item” [...]

Oh god, having filled out a bunch of these for diagnosis and such I hate these with a passion. I always wondered how well these actually work.

I've seen grammatical nonsense like, "Do you often do X? -- always, often, sometimes, rarely, never". What, I often rarely do X? And what does often mean, anyway? Like once a week? Every day?

Then, there are the abstract or vague questions that you then have to interpret what concrete situation it could apply to. Hard to think up an example off the top of my head, but how people reply to these surely depends on what exactly they think it might mean.

Then you start losing patience after about 3 minutes of this shit, not to mention 15 or 30 minutes, and just go through them barely reading the questions, but for the first couple of questions you were pondering whether you "agree" or "somewhat agree" for ages.

I didn't quite get why the article described these surveys as problematic. I skimmed the linked article in that section but this didn't provide any answers either. Assuming the participants fully understand the question, and the survey is designed well, what cultural aspect is stopping people from answering?

I grant that with a questions like "Do you often do X?", examples are necessary to specify what "often" means.

From the article:

> Some people may refuse to answer. Others prefer to answer simply yes or no. Sometimes they respond with no difficulty.

That just sounds like some people boycott the Likert questions, but we don't know why.

These tests have other questions that intersect and determine how much you actually mean it by taking it from different angles. The MMPI-2 for example, may touch on the same subject from various degrees and angles to get the overall picture.

Not saying it's right, but not saying it's a singular question "do you believe X agree? slightly agree? etc." it's a bit more deep and nuanced than that. And the statistics tend to back it up.

That's why starred-review sites have mostly reverted to thumbs-up and thumbs-down. The biggest trial of Likert items in history has spoken.
Or what about people actively lying?
The big problem with this is that the more diversity you have in your study participant set (gender, age, skin color, handedness, etc), the bigger your study has to be to correct for potential biases in results caused by differences in the participants unrelated to what you're testing. And the bigger the study, the more it'll cost, so there will be less studies overall.
It's not so much a problem as simply the cost of doing proper science, and just because you make a small study you don't have to make it cheaply (though you probably are if you're an academic).
I guess the dilemma is whether you want to do good science or just get your papers published for a low price.
I am glad that there are some big studies out there though; big in terms of participants, or big in terms of long-term research - about subjects like nutrition, cancer and heart problem risks, fertility, economic mobility, etc.
Even among 18-21 year old college american women there are ridiculous amounts of variables that you can't control for. Give up the idea that you can control for biases in psychology. Do big mixed tests, then if there's positive outcomes see if its correlated to any of those groups. Then you can do more tests on the groups that didn't respond.
There should be incentive for laboratories to produce bigger and better studies rather than a bunch of small ones.

Especially since a study with 300 participants is worth immensely more than 10 studies with 30 participants.

This is a known problem. Nearly all experimentation-based psychology research was in past decades conducted within the US and is heavily biased to US culture. Cultural differences aren't limited to how people interact with each other, but also include primitive qualities like how people perceive space and use tools.
This isn’t even the worst problem with psychology.

The worst is it’s contextual loyalties against critiquing the relationsips people have with one another. Power struggles, to be particular.

Can you elaborate on this?
Reading the intro I had a strange feeling about this text.

I don't have any issue with the general message. This surely is a legitimate problem. But the intro reads a bit like "psychology is such a great endeavour, if there wasn't this little issue."

Everyone following science news should be aware by now that psychology suffers from a whole range of systemic methodological problems, notably publication bias, widespread p-hacking and failed replications.

I remember our psychology 101 professor offering .1 GPA “extra credit” for every psychology study you went to. There was no way to get a 4.0 in that class unless you attended every class, answered every exam question correctly, every quiz question, etc.
The human behavioral and cultural adaptability to the environment is really impressive and it makes many psychological findings local. With technological and cultural changes some results may apply only to few generations before they go away.

Think children growing in warzone or poor and violent environment. Their behaviour as adults is often sexually more promiscuous, aggressive and their impulse control seems to be less than 'the baseline'. They show trust issues.

How much of that is just damage and disorder as psychology seems to assume, and how much is adaption to survive and procreate in an environment where lifespans are short and life is uncertain. Maybe childhood stress and stress hormones trigger survival strategies that work well in hard environment. They are maladaptive only in the culture and safety of the developed world.

Locus of control is another way of understanding how humans are not necessarily universal to one another by how events are perceived. The missing shape example might just be the children never encounter the square & triangles in daily life and are thrown off it. I don't really consider that a great example for showing how behavioural is different.
The very fact that familiarity with a shape can influence performance on pattern recognition tasks so dramatically is a significant finding on its own.

It would mean we are effectively blind(er) to unfamiliar shapes, even though they are extremely simple, like a triangle or square are.

Much of experimental design in psychology is an effort to attenuate the effect of individual differences so that you can find the common (and probably core) cognitive processes. Sometimes the risk to external validity is worth it in order to reduce the complexity of data and models from confounding variables. So we work with a system where most science is conducted with a relatively homogeneous sample of college undergraduates who are convenient to obtain, while well replicated or foundational findings are later tested across cultures to determine degree of generalization. It seems an ok compromise to me.
a simple arithmetic computation shows that Europe has 750 million people, and the world has 8 billion people. That's 1/10 of the population. We neglect about 90% of people in... anything. South America has 500 million people, so they should get a comparable share, and Africa has 1.2 billion.

So just by simple fractions we know something is off. A more careful study by subject and region could be required.

Psychiatry is in the same boat. Anyone with multiple ongoing issues (depression + anxiety, for instance) is disqualified from studies. Guess what - those are the people that repeatedly have the worst outcomes or don't respond to treatment. We don't even know the real effect of smoking with ADHD people, because if they have anxiety they won't be included in the study.
there could have been any number of blue triangles before that orange square - we see 2, but what about the rest out of frame
75% of statistics are made up
Psychology and social studies, antropology (and economy and whats called climate sciences) really shouldn't be called sciences. It's one of the rare cases IMO where the word really obstructs the core meaning of the subject matter.

They should be considered temporary interpretations of statistical data or in short meta statistics cause that's really what they are.

Much damage is being done by treating these fields as science and the article is only mentioning a few of those problems.

> They should be considered temporary interpretations of statistical data or in short meta statistics cause that's really what they are.

This is what a lot of experimental sciences are, even physical ones, when the systems being studied are complex.

There are very few areas of scientific study anymore that offer convenient, deterministic results. That fruit was picked a long time ago.

Even at the cutting end of physics, researchers have to infer from statistical results.

The difference is only that some of these fields have more reproducible results than others, often because they are studying less complex phenomena, whose causal factors and mechanisms can be more directly observed.

Psychology is at one end of that spectrum, because it is studying the output of the mind, a biological information system whose mechanisms are among the most complex and obscure that people have ever studied.

I think of it on a per-hypothesis or per-theory basis: If a hypothesis or theory is falsifiable, then it is scientific. If not, then it's not.
Why was this flagged?