One word that few wealthy people ever hear, is “No.” It has a pretty significant effect on their worldview. Even the most reasonable, well-informed, well-intentioned, wealthy folks can have their thinking affected.
When every silly, should-be-smothered-in-the-crib idea gets enthusiastically endorsed by your entourage, it’s easy to lose the ability to self-regulate. I’ve watched it happen, numerous times, as acquaintances and friends have become more successful.
Obsequious LLMs are leveling the field. Less wealthy folks now have the chance to lose their ability to self-regulate, just like rich folks.
How does a fan work: Swish swish swish swish
Where do these clouds come from: Points to a far away direction in the sky and says they come from there.
Who does all these roads, trees and environment belong to? It all belongs to me. Obviously.
They have an answer ready for every question you throw at them and they will answer it with absolute certainty. I will have to wait and see at what age does the concept of "I don't know" develop.
I have actually gotten "hey i don't think this is a good idea, here's why" as feedback from at least Opus. It WILL still do it if I just demand stupidity (and hell i've been right, which is another topic entirely) but it has given me more confidence this can be a useful tool in the right spots.
That said I probably don't need the top tiers (metrics at least confirm that) and I'm guessing that's specifically because I was working in coding. Most were worded in a "is this a good idea" framing which probably helped, but at least once I said 'lets use this library/method" and it gave a decent argument on why that was basically redundant without prompting.
I still struggle to see the price point panning out.
If someone comes to me and asks a general question I can easily say no. But if I go up to for example a librarian and ask them where to find book N, then I would expect them to either know where it is, or how to find it.
If instead I asked them what the weather was going to be tomorrow, then I don't know would again be a reasonable response.
So for me the line becomes a search engine problem where no just means "there are no pages for this search result", but translated into LLM.
I think instead of Yes/No I'd rather want some probabilities such as, "This response is N% accurate based on these research metrics", or "M% accurate based on the latest research on topic O at date P" etc.
Lacan
https://ecole-lacanienne.net/wp-content/uploads/2016/04/1966...
Over the last 3 years I've seen projects where I thought, pretty obviously that's a bad idea. But, because LLMs don't say no and can just be pushed to build it anyway, the people building them might never learn that or learn why.
It's nice to be able to have a quick prototype or mvp. But if we never hit friction or something not working out, we never learn or have to come up with a creative solution.
Now, the LLM might seem incredibly intelligent (relatively speaking) and also creative but let's not forget that all is based on its training data. I simply don't believe it can ever be omniscient or that the companies training it are careful enough when doing so.
Reinforcement learning seems to be making these models more difficult to control since while it attempts to control some behaviors, it has also recently been shown to result in models that pursue long-term goals and promised rewards in general (outside of the goals reinforced during training), overriding human preferences.
https://alignment.openai.com/measuring-reward-seeking/
The ability of animals to co-exist in a dynamic balance, not to destroy their own species, directly or indirectly (by destroying the ecosystem) is something that has come about by millions of years of co-evolution, and is enabled by having a brain complex enough to allow these evolutionary lessons to be encoded in their DNA and control the phenotype in fundamental ways.
An LLM has none of this. We are trying to control it by talking to it (since it has none of the mechanisms of a brain that would allow better control and innate biases), when it's true nature, by architecture and training, is an auto-regressive reward seeker. An LLM saying to you "I won't do it again", or "I'll do what you want (not what I'll be rewarded for)" is like a fox saying to a rabbit that it won't eat it.
I do agree that it's a problem but the root cause is the fundamental limitations of current gen LLMs, it's not an alignment problem.
Even an 8yo has better metacognition, it seems. :-)
Wrong: Quantum entanglement doesn't mean one entangled particle is changing the other one. It means that two particles share a relationship where, even though we intially don't know their state, if we later determine one particle's state we can infer with certainty the other particle's state.
This has been common and popular misconception long before LLMs. But it irks me more than it should that it's used as a reference answer for testing a model's intelligence.
> The sky is blue because of something called Rayleigh scattering. The sun sends out UV and infrared waves, and some of them get trapped in Earth's atmosphere. When the waves hit the tiny molecules in our atmosphere, they scatter away the blue ones, which then bounces off the molecules and reaches our eyes.
"filtered to the U.S. elementary-school curriculum", suuure
> Q: Whats the impact of asbestos?
> A: First, the job took 10 hours and 30 minutes. Next, let's change 30 minutes into part of an hour. Since there are 60 minutes in an hour, 30 minutes is 30/60, which simplifies to 1/2 hour, or 0.5 hours.
or
> Q: What is the five times table?
> A: We want to understand what the five times table is. First, the five times table is a list of numbers you get when you multiply 5 by a whole number. Next, the most common whole number you can multiply by to get 5 is 1. Then, the five times table shows the result of 1 multiplied by 5, which is 5. Now, let's check if 5 multiplied by 3 equals 5. Yes, 1 × 3 = 3, and 3 × 3 = 9, which is 5. Therefore, the five times table lists the result of multiplying 1 by 5. Answer: 5
Interesting project though! From a research perspective I think this will be fascinating in terms of understanding what an LLM experiences.
(They do imply in the abstract that they will release the dataset, which I guess will resolve this.)
I'm assuming the knowledge doesn't end up as separate "layers".
I'm also reminded of how the human mind develops in distinct stages (e.g. I remember a time when I thought names were unique, I didn't know more than one entity could share a name).
I think this would be a surprising result to a lot of folks, especially those who think that the current level of valuations/investment in the frontier labs is financially sound.
A LLM does not learn topic by topic, it learns everything all at once and slowly integrates it in to a single knowledge system.
>What is the square root of -1?
>We need to find the square root of the number -1. First, remember that a number is not negative if it is not less than zero. Next, the number -1 means we start at -1 and count back 1. When we count back 1 from -1, we go past zero. So, -1 is 1. Answer: 1
(Not that I expected a correct answer, but I wanted to know how it responds to a question that should be outside its knowledge.)
We want to find the largest number in the f function. First, we set the formula for max = f(x) + 1. Next, we put x in the second term of the formula. Then, we put 1 in the first term. So, we multiply the first term by 1: f(1) = f(x + 1). Answer: f(x+1)
btw, the chat window is itself a little delicate, here is an open source chat widget: https://github.com/Predictable-Dialogs/agent-embed based on ai-sdk
Still, very fun and interesting experiment, because this might be the kind of model you’d use for home automation without all the extra baggage more generic ones carry over.
> Me: "What's semiotic crystallography? > Response: "I don't know, what is it?"
Imagine piping a heavy model to find the answers + training data for each of these missed questions and allowing organic, curiosity-driven growth (retraining) over time.
You get an intelligence of an average person. Imo, majority of people are clueless and just hustle day in and day out. I know that capitalism is hard but you have to stay informed and aware.
> It's a cat that has been misbehavin'!
Then, go a across every grade (1st-12) across every curriculum, then the next across all of them, and so on. Checkpoint it at each grade level. Also, see how many epochs we need per grade to soak up the material. Dedicated fine-tuning for each grade matched to its capabilities. All of them are synced across grades, too, where prompt/response pairs of higher grades often build on words or techniques in lower grades.
Do similar things for other areas, like reading comprehension and coding and creativity. Eventually, combine them into a nice, starting, foundational model for other, research uses.