back
101 comments
The author sort of hints at this in a roundabout way, but the problem for education - until the systems stop lying about things they don't actually know, or using "facts" that are wildly inaccurate or out-of-date, they are not reputable/reliable enough to use at all.

Right now it's like we're in 2001-2005 for wikipedia. It's probably a good tool for basic, basic, basic beginnings of your work in the classroom, but you're going to get things wrong in surprising ways if you use it without verifying.

Schoolbooks for example have wildly out of date facts, or are brief to the point of inaccuracy. Much less watching any history channel show as those tend to suck badly.

One of the things AI's commonly allow is for people to ask questions without fear of judgement. I've not seen one scowl and call a kid an idiot yet (though I'm sure someone has a jailbreak for that). Having an AI trained to take these kids off the wall (and often based on historical inaccuracies) questions would be a really interesting tool.

Just not the only tool, and not really one that should be fully authoritative in itself.

> One of the things AI's commonly allow is for people to ask questions without fear of judgement.

Search engines have been around for almost 30 years now and they do this job better than spicy autocomplete. I type stupid questions into Google all the time and get good answers. The "AI" version of this involves strapping a search engine onto a language model, ostensibly to summarize results, but in practice there are examples of the language model just lying instead of doing the actual search for you.

Right, but when a schoolbook-writer doesn’t know something, they don’t just make up an outrageous lie. Not saying LLMs don’t have a place in education, but they have a long way to go.
Oh sure, a human might hurt your feels so you should only talk to a mechanical turk.

We are all clearly going mad.

Remember, LLMs are purely language AIs. They can and will lie because "an answer " has been weighted high enough that it'll do anything to satisfy that requirement. The purpose of them is to create natural sounding language in response to what it parses the input to be. That's it.

It's similar with what I call "action AIs", where an AI tries to learn to walk, or race a car, optimally, through a track. It will often repeat mistakes, because short term it gains a higher score and takes time to learn that short term gains, in some cases, harm long term gains.

People are using a hammer to install screws. Technically it works, but that's not the droid they were really looking for.

Which puts pathological liars into a new light, at least for me. They’re compulsively story completing.
This isn’t the problem people think it is.

Chain-of-Verification, Process Supervision, encoder/decoder and a plethora of other models are quickly maturing, AND it’s important to remember the current systems ARE NOT particularly optimized in any way for objective accuracy but instead to carry on conversation. It’s a conversation bot.

It’s also important to understand that most systems out there are building on top of the same AI APIs. There isn’t actually that much diversity in the ecosystem that’s broadly deployed yet, so problems with ChatGPT and Bard are “AI problems” and not limited in scope.

As the market matures, solution diversity will increase dramatically and systems that solve these major challenges will emerge and not necessarily from the incumbents. That’s been the pattern in tech waves of the past and why Apple always takes the wait and see and implement the winning solution in the product strategy so often.

These are early days and it’s important not to see the current generation as anything greatly exceeding the first demonstration level technology wave, hard as that is to imagine. It’s 2000, and we are looking at something like a Nokia 3210 (gpt 3.5) and 3310 (gpt 4) talking about how it has problems.

Yep. It’s not a iPhone yet. And we still think Nokia is going to dominate. And the current systems just have problems that are kinda holding them back… but this is the way technology waves break…

>2001-2005 for wikipedia

I was there [still am]. Recently, I provided an update to "transistor density," which was then cited by Perplexity.AI when I asked about a specific new processor type (eerie, having been an early adopter for both wiki and LLMs, from a user-perspective).

I'm left wondering "how much an old wiki handle" [account] might be worth, if it is so-readily cited as "leading authority" (when in reality I was just a curious teenager, trying to figure out what made encyclopedia "so special," when wikipedia provides all these linkages FOR FREE).

Half a lifetime ago, and I'm still curious how this whole "open source thing" is going to play out...

> use it without verifying

There isn't a single general source of truth that we can use without verifying. Human teachers especially aren't such sources.

This isn't as big a problem as you seem to think it is. Even if people need to verify facts maybe it will give them better judgement when other humans lie to them.
Yeah, I always gaslight the kids so that they'll be prepared for the team world.
gpt4 needs better marketing. its not perfect but its without a doubt the best learning tool humans have achieved. its better than books, wikipedia, libraries, etc. its basically your personal college professor with 24/7 office hours on practically any subject. using it in combination with other tools is the best approach, but this has always been the best approach for every learning tool.
As a teacher, I agree.
They're good for at least exploring the vocabulary of a new domain you're not familiar with. You can really dig into a topic and ask systematic things while jotting down new terms to look up and verify. It can help get you past the "don't know what you don't know" stage.
We're at a strange place in computing right now.

At one end of the spectrum we have machines that are fast and efficient. Able to store data, search for data and retrieve data as accurately as it was originally stored.

On the other end of the spectrum we have machines that are so creative we can't control the creativity so it lies and is inaccurate.

What's missing is the machine in the center of these two extremes. A machine that can be creative and factually exact at the same time. We're slowly converging on that goal right now as chatGPT can now use bing to look things up.

Altman has been recently talking about (basically) that continuum and the intention of making it configurable. Seems like an obvious goal to strive for.
Early wikipedia was flawed but c'mon, hoping that a probabilistic text generator happens to string together a series of correct statements is a fundamentally unserious way to gather information
Did you know that those text generators can invoke tools like web search, retrieval, and more to pull in external truth?
Giving the load of bollocks teachers propagate themselves, I suspect the first problem after automated homework will be students calling out the BS. I recall the system really didn. T likd it.
AI as a tool to kickstart your own imagination is fine; but despite the author's claims that he's not excited about AI-as-author, it seems like they're mostly just reading fan-fiction and fan-art created by an AI.

Even the old timey doctor sections, the author immediately admits are mostly factually wrong. What good does that do?

Not factually wrong at all - the dosages, ingredients, diagnosis and even the language are all strikingly accurate. Naturally, the "fake" 1680s doctor didn't write the same prescription as the real one (Sydenham). But a different real life doctor would've disagreed with Sydenham, too. In other words, if you had 100 physicians in the 1680s write out a prescription for hysteria and "hypochondriacal passion," this would be (IMO) indistinguishable from the real ones. What's different is that this is interactive, so you can change elements at will. Again, this isn't reflective of historical fact. It also isn't simply fan fiction. This text isn't an end point, but a starting point for jumpstarting your own thinking about the affordances of a past world.

In other words, I think this is a new method for thinking creatively about history -- one among many that already existed, like historical fiction, historical re-enactment, various forms of experiential learning like debates and roleplaying, etc. But it's cool that there's a new one!

I’m actually building a role play training tool at the moment - https://Solidroad.com . The AI plays the part of a fake customer so sales and support reps can practice before talking to real customers. We’ve built a pretty low latency voice to voice conversation simulator. It helps people build confidence on the phone etc. and we’ve found that accuracy isn’t as important in this context (ie. It’s ok if the AI says something weird every once in a while) as long as the results are by and large believable.

Have a demo in our website where you can try and sell a smart phone to Michael Scott from the office.

> It’s ok if the AI says something weird every once in a while

If anything, I think having the occasional 'complete weirdo' interaction would be good training for the real world. Most people never get to practice how to handle the truly strange cases before the first time they have to deal with a customer who wants to return a case of half-eaten chocolate bars.

I've been a little disappointed with role playing with ChatGPT 4. The AI starts out really close to the behavior that you would expect for the proposed role, but as the conversation gets longer, it starts to forget the role and becomes more generic, like you'd expect normal ChatGPT to be.

I hope this is an implementation detail and that the technology can, in the future, maintain the role playing across really long conversations.

Forgetting context is a problem, for sure. One thing I've found that works fairly well is to include a request for a "status bar" in your prompt. I.e. you ask it to remind itself with each response 1) who it is pretending to be 2) what the date is 3) what the setting is 4) what is in their NPCs "inventory" (which it intuitively understands because LLMs seem to have a natural affinity with MUDs). You can even have it track its mood and variables like weather.

As the context windows of Claude/GPT-4 etc increase, I think this will be less of an issue, but for now it's a pretty effective workaround.

Here's an example of the prompts I'm using (from an activity I just did with my world history class): https://docs.google.com/document/d/1sLRsUVJ_KSPtjrO83ko2MSFf...

And my writeup of an earlier version: https://resobscura.substack.com/p/simulating-history-with-ch...

Look at MemGPT
When I write now, I will sometimes load a page or two into a LLM and ask for an ‘opinion’, and also suggestions for more sub-topics. This is sort-of like bouncing ideas off a human collaborator but also very different.

More to the point of this article, I have a few test prompts that define a bizarre remote planet and culture and ask for a story to be written using this context as a kick-off point. GPT-4 and Claude 2 generally do a good job of ‘creatively’ generating a story. Some smaller LLMs that I run locally on a 32G Mac Mini to a poor job and others are OK. Related: when using smaller local LLMs, I find it important to have many models installed and keep notes on which smaller models are good or bad at specific tasks.

At some point the Danielle Steels of the world will use models trained on their body of work to generate new novels from whole cloth to pump out content with little effort. The only constraint would be not releasing them too quickly for human-written books.
I don't think this will be a viable option. After the first 4-5 of books created like that they'll get stale. The AI won't get any new training data (as Ms. Steel is presumably not writing a book while the AI writes a book for her), and even if the prompts vary, the outputs will feel more similar to each other than an average author's latest 4-5 books
This is a cool project that implements role-playing AI:

https://github.com/joonspk-research/generative_agents

Another one (inspired by the above) that doesn't rely on OpenAI servers:

https://github.com/a16z-infra/ai-town

This is amazing, thanks for sharing it! I’ve been thinking about building something along these lines, so it’s great to see a working model.
We're working on AI role-playing for sales training and coaching. But as we've been validating it we've been learning about a lot of other industries that could use something similar.

https://quick.live

As the pendulum has shifted way (way) over to art as commerce, as a product, then AIs are plausibly useful. Why not make the 'product' more efficiently and cheaply.

But art for me is self-expression; I'm encountering another human being or I'm expressing myself. AI creating that art (really the only "art" IMHO) is a useful as an AI writing memoirs or a love letter or a condolence note.

In general perception, art always has been some mix of those two and people seem to unconsciously conflate them, to slip from one to the other. If AI - plus the current socio-political madness of dismissing all humanities, all compassion, and all humans as anything but economic devices - effectively wipes out art, what have we wrought?

What has the IT revolution, what has Silicon Valley wrought? Look at the world we are creating. And for what? So a few people can make lots of money?

AI generation is like any tool. Once it's available to everyone, you have to do something unique with it to have something worth grabbing people's attention. If AI can illustrate and flesh out the plot of a comic book on a simple prompt, then we'll quickly get bored of those comic books. When someone takes the energy and time to create a unique plot, setting, rich characters, timely themes, etc. and then feeds that to the AI, the result will be interesting to people. People who put less effort in, will get bland, unsellable output out of an AI
And video game creators. True open world games with infinite choices that affect things down the line.
Definitely not -- at least, not yet. Current generation language models are spectacularly bad in this application. It's very difficult to get them to role-play a character in a universe which differs substantially from the real world (since that's where all their training data came from), and they're overly credulous when responding to player input. As a result, they're likely to rapidly go "off the rails" when interacting with players -- they're likely to act unaware of details about the world they're in, to fabricate details about that world or mix in details from other fantasy universes or the real world, or to allow the player to introduce incongruous elements without being challenged.
This is merely another way in which AI will be cast as a co-creator. Especially given the limitations in copyright law in the US I think many early systems will operate like this, just to secure the output as intellectual property…
Question is... would using an AI as a muse lead to more creativity or dulling a personal creativity? It is one thing where people would hang out and spitball ideas that led to inspiration. But I wonder if people in the future will become too dependent on AI muses. Don't forget how we now have generations of math illiterates since the dawning of calculators.
I've tried something similar with a little side project, https://catchingkillers.com. The witnesses, chatgpt chat bots, wouldn't have much of an attention span and start making things up. I think this added to the fun of the game. You had to question each witness to get verification on what another chat bot said.
Can confirm, here's a 600 page story written with it starting in 2015. https://archiveofourown.org/works/46518058/chapters/11713547...
Honestly as a writer I agree that it could be a useful tool, but it would require a lot of UX and UI work to make it truly usable. I don't know if it's worth the effort with the technology as it is. Maybe in 2 months it will be.
Netwrck. Com it's crazy good AI and image generation right now.

If you say "appears" it will generate art that's like midjourney level it's crazy.

But the art is described by the AI first.

That has been a thing long before now.