Right now it's like we're in 2001-2005 for wikipedia. It's probably a good tool for basic, basic, basic beginnings of your work in the classroom, but you're going to get things wrong in surprising ways if you use it without verifying.
One of the things AI's commonly allow is for people to ask questions without fear of judgement. I've not seen one scowl and call a kid an idiot yet (though I'm sure someone has a jailbreak for that). Having an AI trained to take these kids off the wall (and often based on historical inaccuracies) questions would be a really interesting tool.
Just not the only tool, and not really one that should be fully authoritative in itself.
Search engines have been around for almost 30 years now and they do this job better than spicy autocomplete. I type stupid questions into Google all the time and get good answers. The "AI" version of this involves strapping a search engine onto a language model, ostensibly to summarize results, but in practice there are examples of the language model just lying instead of doing the actual search for you.
We are all clearly going mad.
It's similar with what I call "action AIs", where an AI tries to learn to walk, or race a car, optimally, through a track. It will often repeat mistakes, because short term it gains a higher score and takes time to learn that short term gains, in some cases, harm long term gains.
People are using a hammer to install screws. Technically it works, but that's not the droid they were really looking for.
Chain-of-Verification, Process Supervision, encoder/decoder and a plethora of other models are quickly maturing, AND it’s important to remember the current systems ARE NOT particularly optimized in any way for objective accuracy but instead to carry on conversation. It’s a conversation bot.
It’s also important to understand that most systems out there are building on top of the same AI APIs. There isn’t actually that much diversity in the ecosystem that’s broadly deployed yet, so problems with ChatGPT and Bard are “AI problems” and not limited in scope.
As the market matures, solution diversity will increase dramatically and systems that solve these major challenges will emerge and not necessarily from the incumbents. That’s been the pattern in tech waves of the past and why Apple always takes the wait and see and implement the winning solution in the product strategy so often.
These are early days and it’s important not to see the current generation as anything greatly exceeding the first demonstration level technology wave, hard as that is to imagine. It’s 2000, and we are looking at something like a Nokia 3210 (gpt 3.5) and 3310 (gpt 4) talking about how it has problems.
Yep. It’s not a iPhone yet. And we still think Nokia is going to dominate. And the current systems just have problems that are kinda holding them back… but this is the way technology waves break…
I was there [still am]. Recently, I provided an update to "transistor density," which was then cited by Perplexity.AI when I asked about a specific new processor type (eerie, having been an early adopter for both wiki and LLMs, from a user-perspective).
I'm left wondering "how much an old wiki handle" [account] might be worth, if it is so-readily cited as "leading authority" (when in reality I was just a curious teenager, trying to figure out what made encyclopedia "so special," when wikipedia provides all these linkages FOR FREE).
Half a lifetime ago, and I'm still curious how this whole "open source thing" is going to play out...
There isn't a single general source of truth that we can use without verifying. Human teachers especially aren't such sources.
At one end of the spectrum we have machines that are fast and efficient. Able to store data, search for data and retrieve data as accurately as it was originally stored.
On the other end of the spectrum we have machines that are so creative we can't control the creativity so it lies and is inaccurate.
What's missing is the machine in the center of these two extremes. A machine that can be creative and factually exact at the same time. We're slowly converging on that goal right now as chatGPT can now use bing to look things up.
Even the old timey doctor sections, the author immediately admits are mostly factually wrong. What good does that do?
In other words, I think this is a new method for thinking creatively about history -- one among many that already existed, like historical fiction, historical re-enactment, various forms of experiential learning like debates and roleplaying, etc. But it's cool that there's a new one!
Have a demo in our website where you can try and sell a smart phone to Michael Scott from the office.
If anything, I think having the occasional 'complete weirdo' interaction would be good training for the real world. Most people never get to practice how to handle the truly strange cases before the first time they have to deal with a customer who wants to return a case of half-eaten chocolate bars.
I hope this is an implementation detail and that the technology can, in the future, maintain the role playing across really long conversations.
As the context windows of Claude/GPT-4 etc increase, I think this will be less of an issue, but for now it's a pretty effective workaround.
Here's an example of the prompts I'm using (from an activity I just did with my world history class): https://docs.google.com/document/d/1sLRsUVJ_KSPtjrO83ko2MSFf...
And my writeup of an earlier version: https://resobscura.substack.com/p/simulating-history-with-ch...
More to the point of this article, I have a few test prompts that define a bizarre remote planet and culture and ask for a story to be written using this context as a kick-off point. GPT-4 and Claude 2 generally do a good job of ‘creatively’ generating a story. Some smaller LLMs that I run locally on a 32G Mac Mini to a poor job and others are OK. Related: when using smaller local LLMs, I find it important to have many models installed and keep notes on which smaller models are good or bad at specific tasks.
But art for me is self-expression; I'm encountering another human being or I'm expressing myself. AI creating that art (really the only "art" IMHO) is a useful as an AI writing memoirs or a love letter or a condolence note.
In general perception, art always has been some mix of those two and people seem to unconsciously conflate them, to slip from one to the other. If AI - plus the current socio-political madness of dismissing all humanities, all compassion, and all humans as anything but economic devices - effectively wipes out art, what have we wrought?
What has the IT revolution, what has Silicon Valley wrought? Look at the world we are creating. And for what? So a few people can make lots of money?
If you say "appears" it will generate art that's like midjourney level it's crazy.
But the art is described by the AI first.