+ the experiments may already be in the dataset so it’s really testing if it remembers pop psychology
I was not able to put my finger on what I felt wrong about the article -- till I read this
Also important is that in Psychology/Sociology, it's the counter-intuitive results that get published. But these results disproportionately fail to replicate!
Nobody cares if you confirm something obvious, unless it's on something divisive (e.g. sexual behavior, politics), or there is an agenda (dieting, etc). So people can predict those ones more easily than predicting a randomly generated premise. The ones that made their way into the prediction set were the ones researchers expected to be counter-intuitive (and likely P-hacked a significant proportion of them to find that result). People know this (there are more positive confirming papers than negative/fail-to-replicate).
This means the counter-intuitive, negatively forecast results, are the ones that get published i.e. the dataset saying that 66% of human forecasters is disproportionately constructed of studies that found counter-intuitive results compared to the overall neutral pre-published set of studies, because scientists and grant winners are incentivised to publish counter-intuitive work. I would even suggest the selected studies are more tantalizing that average in most of these studies, they are key findings, rather than the miniature of comments on methods or re-analysis.
By the way the 66% result has not held up super well in other research, for example, only 58% could predict if papers would replicate later on: https://www.bps.org.uk/research-digest/want-know-whether-psy... - Results with random people show that they are better than chance for psychology, but on average by less than 66% and with massive variance. This figure doesn't differ from psychology professors which should tell you the stat represents more the context of the field and it's research apparatus itself rather than capability to predict research. What if we revisit this GPT-4 paper in 20 years, see which have replicated, ask people to predict that - will GPT-4 still be higher if it's data is frozen today? If it is up to date? Will people hit 66%, 58%, or 50%?
My point is, predicting the results now is not that useful because historically, up to "most" of the results have been wrong anyhow. Predicting which results will be true and remain true would be more useful. The article tries to dismiss the issue of the replication crisis by avoiding it, and by using pre-registered studies, but such tools are only bandages. Studies still get cancelled, or never proposed after internal experimentation, we don't have a "replication reputation meter" to measure those (which affect and increase false positive results), and we likely never will, with this model of science for psychology/sociology statistics. If the authors read my comment and disagree, they should use predictions for underway replications with GPT-4 and humans, wait a few years for the results, and then conduct analysis.
Also, more to the point, as a Psychology grant funded once told me - the way to get a grant in Psychology is to: 1) Acquire a result with a counter-intuitive result first. Quick'n'dirty research method like students filling in forms, small sample size, not even published, whatever. Just make the story good for this one and get some preliminary numbers on some topic by casting a big web of many questions (a few will get P < 0.05 by chance eventually in most topics anyway at this sample size) 2) Find an angle whereby said result says something about culture or development (e.g. "The Marhsmallow experiment shows that poverty is already determined by your response to tradeoffs at a young age", or better still "The Marshmallow experiment is rubbish because it's actually entirely explained by SES as a third factor, and wealth disparity in the first place is ergo the cause". Importantly, change the research method to something more "proper" and instead apply P-hacking if possible when you actually carry out the research. The biggest P-hack is so simple and obvious nobody cares: you drop results that contradict or are insignificant, and just don't report them - carrying out alternate analysis, collecting slightly different data, switching from online to in person experiments, whatever you canto get a result. 3) Upon the premise of further tantalizing results, propose several studies which can fund you over 5 years, apply some of the buzz words of the day. Instead of "Thematic Analysis", It's "AI Summative Assessment" for the Word Frequency amounts, etc. If you know the grant judgers, avoid contradicting whatever they say, but be just outside of the dogma enough (usually, culturally) to represent movement/progress of "science".
This is how 99% of research works. The grant holder directs the other researchers. When directing them to carry out an alternate version of the experiment or change what we are analyzing, you motivate them that it's for the good of the future, society, being at the cutting edge, and supporting the overarching theory (which ofcourse, already has "hundreds" of supporting evidence from other studies constructed in the same fashion).
As to sociology/psychology experiments - Do social experiments represent language and culture more than people and groups? Randomly.
Do they represent what would be counter-intuitive or support developing and entrenching models and agendas? Yes.
90% of social science studies have insufficient data to say anything at P < 0.01 level which should realistically be our goal if we even want to do statistics with the current dogma for this field (said kindly because some large datasets are genuine enough and used for several studies to make up the numbers in the 10%). I strongly see a revolution in psychology/sociology within the next 50 years to redefine a new basis.
“GPT said people would hate buses, so we halved their number and slashed transportation budget… Wow, do our people actually hate buses with passion!”
“A year ago GPT said people would not be worried about climate change, so we stopped giving it coverage and removed related social adverts and initiatives. People really don’t give a flying duck about climate change it turns out, GPT was so right!”
This is an oversimplification, of course; to say it with more nuance, anything socio- and psycho- is a minefield of self-fulfilling prophecies that ML seems to be nicely positioned to wreak havoc in. (But the small “this is not a replacement for human experiment” notice is going to be heeded by all, right?)
As someone wrote once, all you need for machine dictatorship is an LLM and a critical number of human accomplices. No need for superintelligence or robots.
You jest, but if you don't mind me going off on a tangent, this reminds me how in the summer 2020 post-lockdown-period the local authorities of Barcelona decided that to reduce the spread of COVID they had to discourage out-of-town people going to the city for nightlife... so they halved the number of night buses connecting Barcelona with nearby towns. Because, of course, making twice the number of people congregate at bus stops and making night buses even more crammed was a great way to reduce contagion. Also, as everybody knows, people's decision whether or not to party in town on a Friday night is naturally contingent on the purely rational analysis as to the number of available buses to get home afterwards.
IMO mass communication technologies (radio, TV, internet) are much more important in building a dictatorship.
Hint: They like it because it's biased for what they want... like real marketing studies.
> Out of One, Many: Using Language Models to Simulate Human Samples
> We propose and explore the possibility that language models can be studied as effective proxies for specific human sub populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the "algorithmic bias" within one such tool -- the GPT 3 language model -- is instead both fine grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property "algorithmic fidelity" and explore its extent in GPT-3. We create "silicon samples" by conditioning the model on thousands of socio demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT 3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.
The comments were not terribly supportive. They’ve since pivoted to a product that does survey data cleaning.
Probably would be sufficient to just give a couple examples of what might constitute one of these.
Sorry, I know this is probably basic to someone who is in that field.
We should expect LLMs to be pretty good at repeating back to us the stories we tell about ourselves.
(this is parody)
If so, how close was the examination vs the record the model was trained on.
Some interesting insights there, I think.
Eg. Try LLM's to find availability hours when you have the start and end time of each day.
LLM's don't really understand that you need to use the day 1 endhour and then the starthour of the next day.
Emergent vs Reductionist
LLMs are Emergent
Therefore, LLMs break the mold of human technology and enter the realm of nature's technology
We should talk about AI Beings not apps if we wish to stay with this analogy.
We can always say that's BS and that the analogy does not reach that deep
Or we can take it for what it is and admit LLMs are not similar in their behavior to any tool that humans have created to date, all the way from homo habilus till homo sapiens sapiens. Tools are predictable. Intelligences are not.
Tesla FSD is a good example of this in real life. You can measure how closely the car acts like a human based off of interventions and crashes that were due to unhuman behavior, as well in the first round of the robot taxi fleet which will have a safety driver, you can measure how many people complain that the driver was bad
Joking…but not joking.
GPT-4 can stand in for humans. Charlie Brown is mentioned in the Upanishads. The bubonic plague was a spread via telegram. Easter falls on 9/11 once every other decade.
You can just write shit and hit post and boom, by nature of it being online someone will entertain it as true, even if only briefly so. Wild stuff!