back
94 comments
I'm very skeptical on this, the paper they linked is not convincing. It says that GPT-4 is correct at predicting the experiment outcome direction 69% of the time versus 66% of the time for human forecasters. But this is a silly benchmark because people are not trusting human forecasters in the first place, that's the whole purpose for why the experiment is run. Knowing that GPT-4 is slightly better at predicting experiments than some human guessing doesn't make it a useful substitute for the actual experiment.
For sure. Great argument

+ the experiments may already be in the dataset so it’s really testing if it remembers pop psychology

Furthermore, there’s a replication crisis in social sciences. The last thing we need is to accumulate less data and let an LLM tell us the “right” answer.
Predicting the actual results of real unpublished experiments with a 0.9 correlation factor is a very non-trivial result. The human forecasts comparison is not the central finding
Nicely put! Well argued!

I was not able to put my finger on what I felt wrong about the article -- till I read this

That's surprisingly low considering it was probably trained on many of the papers it's supposed to be replicating.
I totally agree. So many people are missing the point here.

Also important is that in Psychology/Sociology, it's the counter-intuitive results that get published. But these results disproportionately fail to replicate!

Nobody cares if you confirm something obvious, unless it's on something divisive (e.g. sexual behavior, politics), or there is an agenda (dieting, etc). So people can predict those ones more easily than predicting a randomly generated premise. The ones that made their way into the prediction set were the ones researchers expected to be counter-intuitive (and likely P-hacked a significant proportion of them to find that result). People know this (there are more positive confirming papers than negative/fail-to-replicate).

This means the counter-intuitive, negatively forecast results, are the ones that get published i.e. the dataset saying that 66% of human forecasters is disproportionately constructed of studies that found counter-intuitive results compared to the overall neutral pre-published set of studies, because scientists and grant winners are incentivised to publish counter-intuitive work. I would even suggest the selected studies are more tantalizing that average in most of these studies, they are key findings, rather than the miniature of comments on methods or re-analysis.

By the way the 66% result has not held up super well in other research, for example, only 58% could predict if papers would replicate later on: https://www.bps.org.uk/research-digest/want-know-whether-psy... - Results with random people show that they are better than chance for psychology, but on average by less than 66% and with massive variance. This figure doesn't differ from psychology professors which should tell you the stat represents more the context of the field and it's research apparatus itself rather than capability to predict research. What if we revisit this GPT-4 paper in 20 years, see which have replicated, ask people to predict that - will GPT-4 still be higher if it's data is frozen today? If it is up to date? Will people hit 66%, 58%, or 50%?

My point is, predicting the results now is not that useful because historically, up to "most" of the results have been wrong anyhow. Predicting which results will be true and remain true would be more useful. The article tries to dismiss the issue of the replication crisis by avoiding it, and by using pre-registered studies, but such tools are only bandages. Studies still get cancelled, or never proposed after internal experimentation, we don't have a "replication reputation meter" to measure those (which affect and increase false positive results), and we likely never will, with this model of science for psychology/sociology statistics. If the authors read my comment and disagree, they should use predictions for underway replications with GPT-4 and humans, wait a few years for the results, and then conduct analysis.

Also, more to the point, as a Psychology grant funded once told me - the way to get a grant in Psychology is to: 1) Acquire a result with a counter-intuitive result first. Quick'n'dirty research method like students filling in forms, small sample size, not even published, whatever. Just make the story good for this one and get some preliminary numbers on some topic by casting a big web of many questions (a few will get P < 0.05 by chance eventually in most topics anyway at this sample size) 2) Find an angle whereby said result says something about culture or development (e.g. "The Marhsmallow experiment shows that poverty is already determined by your response to tradeoffs at a young age", or better still "The Marshmallow experiment is rubbish because it's actually entirely explained by SES as a third factor, and wealth disparity in the first place is ergo the cause". Importantly, change the research method to something more "proper" and instead apply P-hacking if possible when you actually carry out the research. The biggest P-hack is so simple and obvious nobody cares: you drop results that contradict or are insignificant, and just don't report them - carrying out alternate analysis, collecting slightly different data, switching from online to in person experiments, whatever you canto get a result. 3) Upon the premise of further tantalizing results, propose several studies which can fund you over 5 years, apply some of the buzz words of the day. Instead of "Thematic Analysis", It's "AI Summative Assessment" for the Word Frequency amounts, etc. If you know the grant judgers, avoid contradicting whatever they say, but be just outside of the dogma enough (usually, culturally) to represent movement/progress of "science".

This is how 99% of research works. The grant holder directs the other researchers. When directing them to carry out an alternate version of the experiment or change what we are analyzing, you motivate them that it's for the good of the future, society, being at the cutting edge, and supporting the overarching theory (which ofcourse, already has "hundreds" of supporting evidence from other studies constructed in the same fashion).

As to sociology/psychology experiments - Do social experiments represent language and culture more than people and groups? Randomly.

Do they represent what would be counter-intuitive or support developing and entrenching models and agendas? Yes.

90% of social science studies have insufficient data to say anything at P < 0.01 level which should realistically be our goal if we even want to do statistics with the current dogma for this field (said kindly because some large datasets are genuine enough and used for several studies to make up the numbers in the 10%). I strongly see a revolution in psychology/sociology within the next 50 years to redefine a new basis.

This so much. There was another similar one recently which was also BS.
If GPT emulations of social experiments are not correct, policy decisions based on them will make them so.

“GPT said people would hate buses, so we halved their number and slashed transportation budget… Wow, do our people actually hate buses with passion!”

“A year ago GPT said people would not be worried about climate change, so we stopped giving it coverage and removed related social adverts and initiatives. People really don’t give a flying duck about climate change it turns out, GPT was so right!”

This is an oversimplification, of course; to say it with more nuance, anything socio- and psycho- is a minefield of self-fulfilling prophecies that ML seems to be nicely positioned to wreak havoc in. (But the small “this is not a replacement for human experiment” notice is going to be heeded by all, right?)

As someone wrote once, all you need for machine dictatorship is an LLM and a critical number of human accomplices. No need for superintelligence or robots.

> “GPT said people would hate buses, so we halved their number and slashed transportation budget… Wow, do our people actually hate buses with passion!”

You jest, but if you don't mind me going off on a tangent, this reminds me how in the summer 2020 post-lockdown-period the local authorities of Barcelona decided that to reduce the spread of COVID they had to discourage out-of-town people going to the city for nightlife... so they halved the number of night buses connecting Barcelona with nearby towns. Because, of course, making twice the number of people congregate at bus stops and making night buses even more crammed was a great way to reduce contagion. Also, as everybody knows, people's decision whether or not to party in town on a Friday night is naturally contingent on the purely rational analysis as to the number of available buses to get home afterwards.

All you need for dictatorship in general is a critical number of human accomplices. I don’t see how an LLM in the mix would make it worse.

IMO mass communication technologies (radio, TV, internet) are much more important in building a dictatorship.

Everyone and their mom in advertising sold "GPT Persona" tools which are basically just an api call to brands for target group simulation. Think "Chat with your target group" kinda stuff.

Hint: They like it because it's biased for what they want... like real marketing studies.

Reminds me of:

> Out of One, Many: Using Language Models to Simulate Human Samples

> We propose and explore the possibility that language models can be studied as effective proxies for specific human sub populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the "algorithmic bias" within one such tool -- the GPT 3 language model -- is instead both fine grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property "algorithmic fidelity" and explore its extent in GPT-3. We create "silicon samples" by conditioning the model on thousands of socio demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT 3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.

https://arxiv.org/abs/2209.06899

A YC company called Roundtable tried to do this.[1]

The comments were not terribly supportive. They’ve since pivoted to a product that does survey data cleaning.

[1] https://news.ycombinator.com/item?id=36865625

Can someone translate for us non-social-scientists in the audience what this means? "3. Treatment. Write a message or vignette exactly as it would appear in a survey experiment."

Probably would be sufficient to just give a couple examples of what might constitute one of these.

Sorry, I know this is probably basic to someone who is in that field.

But do the experiments replicate better in LLMs than in actual humans? :D

We should expect LLMs to be pretty good at repeating back to us the stories we tell about ourselves.

I don't think "replicate" is the appropriate word here.
So did ELIZA[0] about sixty (60) years ago.

0 - https://en.wikipedia.org/wiki/ELIZA

So, we finally found the cure for the replication crisis in social sciences: just run them on LLMs.
Why stop at social science? I say we make a questionnaire, give it to the GPT over a broad range of sampling temperatures, and collect the resulting score:temperature data. From that dataset, we can take people's temperatures over the phone with a short panel of questions!

(this is parody)

Accompanying working paper that demonstrates 85% accuracy of GPT-4 in replicating 70 social science experiment results: https://docsend.com/view/qeeccuggec56k9hd
I wonder if this could be used for testing marketing or UX actions?
Is that the solution to social science's replication problem?
Were those experiments in the training set?

If so, how close was the examination vs the record the model was trained on.

Some interesting insights there, I think.

Did they test whether GPT4 replicated already existing social science experiments? If so, this might have happened because the experiment was in the training data
That's only for known situations.

Eg. Try LLM's to find availability hours when you have the start and end time of each day.

LLM's don't really understand that you need to use the day 1 endhour and then the starthour of the next day.

What happened to running virtual experiments aka running filters on NSA metadata from all conversation of all mankind?
Nature's technology vs. Human Technology

Emergent vs Reductionist

LLMs are Emergent

Therefore, LLMs break the mold of human technology and enter the realm of nature's technology

We should talk about AI Beings not apps if we wish to stay with this analogy.

We can always say that's BS and that the analogy does not reach that deep

Or we can take it for what it is and admit LLMs are not similar in their behavior to any tool that humans have created to date, all the way from homo habilus till homo sapiens sapiens. Tools are predictable. Intelligences are not.

The good news is that they should be able to replicate real world events to validate of this is true or not.

Tesla FSD is a good example of this in real life. You can measure how closely the car acts like a human based off of interventions and crashes that were due to unhuman behavior, as well in the first round of the robot taxi fleet which will have a safety driver, you can measure how many people complain that the driver was bad

I think it is far, far more likely that it replicates social science experiments well enough to simulate people
Not so sure about that, but I can absolutely see people getting addicted to them.
Ooooor maybe, testing if the experiments are similar to what was in the corpus.
But does it replicate _better_ than really running the experiment again?

Joking…but not joking.

This is gonna end well…
this tells more about how social science data is manipulated than the usefulness of llm
please don't, need I remind you the joke that social science is not real science
I love that anyone can just write whatever they want and post it online.

GPT-4 can stand in for humans. Charlie Brown is mentioned in the Upanishads. The bubonic plague was a spread via telegram. Easter falls on 9/11 once every other decade.

You can just write shit and hit post and boom, by nature of it being online someone will entertain it as true, even if only briefly so. Wild stuff!

Well that’s one way to solve the replication crisis
And yet it can't replicate a human support agent. Or even a basic search function for that matter ;)
garbage in, eh?
Phsycohistory
Source: trust us. This is some bullshit science.
Is it possible to train an LLM that is minimally biased and that could assume various personas for the purpose of the experiments? Then I imagine it’s just some prompt engineering no?