back
52 comments
Good luck to the xAI employees who just learned that they are open sourcing their product this week.
Meaning?
Read the biography Elon Musk from Walter Isaacson.

The engineers likely learned of this news via the tweet.

Alex Heath from The Verge alleged that Grok is just tuned LLaMa [0]. I wonder what will be revealed!

[0]: https://www.threads.net/@alexheath/post/C0pEidVp-1U

Would certainly not be surprised!
Wasn’t the tweet recommendation system “open sourced” as well? Does this guy know the difference between open source and “open source”?
> Wasn’t the tweet recommendation system “open sourced” as well? Does this guy know the difference between open source and “open source”?

What do you mean? There exists only one binding definition of open source

> https://opensource.org/osd

and either some product does satisfy it, or it doesn't. As far as I am aware

> https://github.com/twitter/the-algorithm

does satisfy the open source definition, so your sarcasm looks demagogical to me, but I am very willing to learn something new.

I think people expected him to "open" the algorithm so that you could tell how the recommendations are determined, and instead what people got was an Underpants Gnomes' Plan with a neural network step in the middle and no weights.
While I agree there is a common understanding of what open source is, there most definitely does not exist any "binding definition"! It is not trademarked, copyrighted (and it never could have been, two common words that it is), or in any country's legally protected terms, or anything else. It is really grating to see such nonsense repeated way to often.
> There exists only one binding definition of open source

>> https://opensource.org/osd

Insert Obama awarding himself meme. Who said that this is the "only binding definition"?

So it is not open source, thank you for the info.
Yes, and it's here: https://github.com/twitter/the-algorithm

If e.g. Amazon open sources some part of its software infrastructure should they also open source the data it uses or their configuration files?

If I recall correctly this repo is missing data so it's functionally impossible to verify or replicate the behaviour they're using on live.
Not only that, it's not been updated in 8 months. It's extremely unlikely that Twitter hasn't updated anything about the home feed since then. They effectively dumped part of the code on GitHub for some headlines, but never intended to keep developing it in the open.
> but never intended to keep developing it in the open

Did Elon Musk promise this?

"Open source" or "open weight"? Because there is a distinction. Many have previously provided open weights (or what they call "open model" now): Mistral, LLaMA, Falcon, etc. There are not many open "source" LLMs out there that bring true value to business and academia.
how does grok even compare to the rest of llms? it seems like it was just Elon throwing up shit because he wants Twitter to be as big and bad as Google and Facebook, and even Google has been really fumbling trying to compete with Microsoft and openai, FB has been surprising with their more open approach open models and Mistral seemingly came out of nowhere with some great tech.

Is grok really noteworthy or is it just a nothing burger?

I prefer the term "model available".
Has anyone benchmarked Grok against other models? The LLMSYS benchmarks, which I trust most, don't have it. And their own reported results are good but nothing amazing since it doesn't seem to surpass GPT4 or Claude 3.
The general consensus is it's in the "GPT3.5" class along with llama 2 and co, but it has a very annoying attitude. I don't know anybody routinely using it.
The project ignores the fundamental GIGO nature of the written word so completely that I’m assuming someone’s running a con on Musk.
God, the replies to that tweet are deranged.
Yeah. So the remaining concern is license. I hope it won’t be similar to Llama.
People seem very concerned about licenses for LLM weights.

Why shouldn't we treat LLM weights like LLM creators treat ebooks and open source code? Namely, that it is not subject to copyright?

To say that the Llama training process bypasses the copyright of all the training data creators, and yet the output is copyrighted by Facebook, seems a uniquely pro-corporation stance.

This is a really interesting framing that I hadn’t thought about before.

You’re absolutely right. It’s very one sided at the moment.

If we follow their ebook usage practice, it’s not even required that they declare it to be open source. Just need someone to publish their copyrighted work online [0] without their agreement and then - per their rules - it’s totally acceptable to download and use those weights with abandon.

Maybe it could be called “weights3”

[0] I’m not actually suggesting anyone should do this.

You really can't have an anon internet and copyrights simultaneously.

Take Wikipedia's content, licensed under Creative Commons - by who? Donald Duck? Then when Pikachu and Tony Stark edit the article it becomes a derived work?

> Creative Commons licenses give everyone from individual creators to large institutions a standardized way to grant the public permission to use their creative work under copyright law.

>....so long as attribution is given to the creator.

Who is the creator I must attribute to?

I don't think any of WP is CC? Without at least a full name and claim of authorship I cant satisfy the requirements of the license? Or can I? Then if I can satisfy attribution I will have to disclose who I am in order to allow further sharing.

When Scratch[0] took off lots of kids re-uploaded things made by others replacing the description with "I MADE THIS"

I'd say we, the grown ups of this world should know we've messed up when kids mock our ways.

[0] - https://scratch.mit.edu

Every artist does the same.

Get inspired and training on prev work, creating something new.

tidbit: Oracle have an OpenGrok project under active development:

https://en.wikipedia.org/wiki/OpenGrok

...and nobody will care.
This week, @xAI will open source Grok

It's like that glorious week in 2018 when we got full self driving.

You can get FSD today. Thousands of people use it every day https://www.teslarati.com/tesla-fsd-beta-program-half-a-bill...
Full self driving is a misleading name, Level 2 is not Level 5, Tesla is overpromising a solution, and I wouldn't consider something in "beta" that is safety critical to be considered shipping as GA.
Its name, doesn't tell me how good it is. until I can buy a Tesla, give it a Lyft account, and have it go make me money driving for Lyft, it's not worth much to me.
It's still a huge success.
This.

Instead what you have is Fool Self Driving as Tesla knows that it isn't fully autonomous yet, but still market it as such.

Intentionally misleading and irresponsible.

I think this is simply a confusion over the meaning of words.

You can indeed buy Full Self Driving™ (FSD), but even then your Tesla is not capable of self driving, fully (eg, there are many scenarios where a human is still required)

People take "Full" as meaning L5, but Tesla uses "full" as in ODD (Operational Design Domain). It can go anywhere, city streets, highways, parking lots, unmarked roads. In that sense it is indeed "full". This is clear when you look at the history of Tesla autonomy products, first there was Autopilot, which is only for highways. Then they release Full Self Driving Beta, which includes every type of driving.
I think I reflexively agree, but I will try to suspend my opinion until it is actually released.