back
83 comments
Given how good Llama 3.1, 3.2 and 3.3 were I'm genuinely looking forward to news on Llama 4.

3.3 70B is the best model I've managed to run on my laptop, and 3.2 3B is my favourite model to run on my phone.

I'd love to hear more about how you're using models on a phone. Have you written anything about it?
Honestly the one on my phone is more for fun than anything else - I use it to show people that LLMs can run on personal devices but rarely do anything useful with it.

I use the MLC Chat app from the App Store: https://apps.apple.com/gb/app/mlc-chat/id6448482937

My favourite example prompt for demos is "Write an outline of a Netflix Christmas movie where a topical-profession falls in love with another topical-profession" - customized for the occasion.

e.g. "Write an outline for a Netflix Christmas movie set in San Gregorio California about a man who runs an unlicensed cemetery falling in love with a barrister at the general store" - result here: https://bsky.app/profile/simonwillison.net/post/3ldthrqb6c22...

fullmoon is a good FOSS iOS/macOS client for mlx

https://fullmoon.app/

I wish they called it llamarama...
I would bet at least one person at Meta wanted to call it Llamapalooza.

In any case, I'm excited!

That way attendees could be called Llamarama Dingdongs.
Great! Maybe I can finally learn what I'm suppose to be using generative AI for to be more productive. I'll be tuning in and spinning up whatever models/tools they suggest, but the longer this tech wave occurs the more confident I am that gen-ai is going to tally up to be an at-most 3% lift on global productivity.
I was pretty cynical trying earlier models, but with Gemini Flash 2.0 I felt like there was a pretty significant boost in usability and capabilities.

In particular I've found that these tools make it a lot easier to explore or get started with unfamiliar domains. One of my big issues has often been decision paralysis, so having a tool to help me narrow down the list of resources and make it more approachable has been a huge win.

My general experience has been that getting AI tools to directly do stuff for you tends to produce pretty bad results, but you can use it as a force multiplier for your own capabilities. If I'm confused or uncertain about how to do something, AI tools are usually pretty good at clarifying what needs to be done.

Maybe there’s a particular cognitive profile that benefits most from LLM chat bots? I’ve tried multiple times to realize this force multiplier in my life for everything from day-to-day stuff to picking up new things, using the latest paid bots with the best models, and I’ve persistently found them to be awkward, inaccurate, difficult to pin down into giving useful info rather than a bunch of non-committal pablum without hallucinations, etc etc etc
Here’s something scary I recently learned: Cedar-Sinai in LA (major hospital) used to have 15 lawyers on contracts. Now they have 1 and an AI app reviewing contracts.

Those are 14 lawyers gone. That’s more than 3% on “productivity”, but 14 people who lost their jobs. And that’s now with the current state of things.

Lawyer here - there are fields, and law is definitely one of them, where labor is the major cost.

That labor is not often used sanely.

It is common to use lawyers costing hundreds per hour to do fairly basic document review and summarization. That is, to produce a fairly simple artifact.

Not legal research, not opinionated briefing.

But literal: Read these documents, produce a summary of what they say.

While I can't say this is the same as what you are talking about ("contracts review" means many things to many people), i'm not even the slightest bit surprised that AI is starting to replace remarkably inefficient uses of labor in law.

I will add: Lots of funding being thrown at AI legal startups around on products that do document review and summarization, but that's not the big fish, and will be commodity very quickly.

So i expect there will be an ebb and flow of these sorts of products as the startups either move on to things that enable them to capture a meaningful market (document review ain't it), or die and leave these companies hanging :)

Going to be interesting to see the MTTL (mean time to lawsuit) on this. Sounds grossly negligent. I feel kinda sorry for the lonely lawyer.
That's great. That means potentially slightly cheaper healthcare. A company shouldn't need so many lawyers unless it's a legal firm.
Do you have an article that has more information about this? I'd really like to learn more about what happened.
Can you share a source/article?
Cursor Composer is a game-changer for greenfield projects, it's definitely not 3% change.
> Maybe I can finally learn what I'm suppose to be using generative AI for to be more productive.

So many things. It's a general-purpose "thing doer" in many situations where you otherwise wouldn't have one. Let me give a super-simple example - not a high-value one, but an example of obvious value IMO.

Say for some reason, you have a screenshot of a bunch of text. Maybe you took a picture of a page from a book or something, idk. Now you want it in textual form. You can throw it in ChatGPT and ask it to give you the text, and a few seconds later you have the text.

I'm not saying there are no other solutions for this - there are. You can look for some software to do OCR or something. But that's what makes ChatGPT or others general-purpose - they're a one-stop shop for a lot of different things, including small one-off tasks like this. I can name a dozen other one-off tasks that it helps me with. Again, not the most high-value things it helps me with (that'd be programming help), but an undeniable example of value, IMO.

How would you do this without an LLM? (I personally would've just typed it up myself, probably.)

> How would you do this without an LLM? (I personally would've just typed it up myself, probably.)

This has been a feature of Apple Preview (default image program) for years and years. You can just highlight text and copy it from a jpeg or png.

It mostly just uses those other solutions in the background. You can open the Analyzing and see it building a script with OpenCV for vision tasks. It's a handy front-end.
Folks are busy optimizing their own building, not telling you how to optimize yours...
Yeah, but finally everybody can now bulshit freely aided by their personal llm, not just the natural bulshitters; but I forsee that bullshitting may spike up a bit now then fall out of fashion for something even more fleeting as people’s attention span is getting more and more fractured.
It would be beneficial to have a hardware-optimized Llama lineup with a clearer naming scheme and distinct performance tiers, for example:

- Llama 4.0 Phone (Lite / Standard / Max) – For mobile devices.

- Llama 4.0 Workstation (Lite / Standard / Max) – For PCs and laptops.

- Llama 4.0 Server (Lite / Standard / Max) – For high-performance computing.

This approach would enable developers to select the appropriate model based on both device type and performance needs.

What do you think? For example now I feel like 3.3 70B is more for laptops/PCs, and the previous 3.2 3B for phones, is a bit confusing to me.

The “con” being the claims of “open source” when Llama is at best “weights available”. It’s not even “open weights” since it has a proprietary license. But I’m sure that won’t stop Yann LeCun from repeating lies about how Meta is the leader in open source AI.
3 years ago, Meta Connect in 2022 was a different atmosphere. [0] Almost no-one cared.

That was close to the bottom of Meta's stock price.

[0] https://news.ycombinator.com/item?id=33087535

Since then the avg ad price, they have reported has risen for 12 quarters in a row, past 6 quarters its jumped 15-30%. The MO is to prey on small businesses and people who want attention, world wide, who don't know anything about advertising/marketing. Everything they know comes from what Google and Meta tell them.
Odd timing - right in the middle of RSA in SF. Llama (and other US-trained open weight models) are key to national security and cybersecurity, and there's a built in audience 30 minutes away if this were two days later or two days earlier...
Maybe they should have used one of those newfangled AI automatic calendar manager things. Or maybe they were using one and shouldn’t have been.
It's sad how we are letting Meta get away with their abuse of the term "open-source" for their open weights models. :-(
Dates are awfully close to ICLR, but I suppose the audiences don’t really overlap.
Can you imagine how many LinkedIn thought leaders are going to be in attendance? Perhaps the greatest gathering of minds since the Manhattan Project.
Super thrilled about all the cross-functional synergies and ROI-optimized deliverables poised to disrupt the status quo and elevate the strategic framework.
Can't wait to delve into it!
LlamaCon: the greatest buzzword bingo convention opportunity in 2025

Double click with us.

Thanks! This made my day. ROFL.
"Super thrilled about all the cross-functional synergies ..."

Could you explain what that means - please?

"At LlamaCon, we’ll share the latest on our open source AI developments to help developers do what they do best: build amazing apps and products, whether as a start-up or at scale."

Strangely enough, I can work quite well without your help. I've been doing it professionally for 35 odd years. I'm "just" an engineer - no capital E - I simply studied Civil Engineering at college and ended up running an IT company and I'm quite good at IT.

What I would really like to see is really well indexed documentation written by people ie an old school search engine. Google used to do that and so did Altavista, back in the day.

I do not need or want a plethora of trite simulacra web sites dripping with AI wankery at every search term.

> well indexed documentation

Indices are, by definition, lossy representations of their underlying data. If you use stemming and lemmatization to preprocess both documentation and query text, you're already departing from a truly hand-optimized indexing system, and choosing to have imperfect algorithms do things in a more scalable way. And indexing by embedding vectors that use LLMs to determine context are a natural extension of this, in my view. And on top of that, when you have a massive amount of candidate text to display to the user... is displaying sentence fragments one on top of the other the most optimal UX there? At a certain point, RAG becomes the answer to this question.

The problem, as you note, is that search engines and social media systems are incentivized to allow garbage content into the original set of things they index and surface, if that garbage content drives more attention to advertisements. But that's not a reason to reject the benefits that the underlying LLM technology can bring towards building good indexing on top of human-written documents. It just won't be done by the companies that used to do it.

You really don't use Copilot or ChatGPT? Have you tried them?
Check out Zeal - zealdocs.org - you get indexed docs for stuff everyone uses.
Probably not going to be saying much though. The state of real-time LLM-based conversation aids just isn’t where it needs to be for those folks to function in public effectively. I could foresee there being a heck of a slam broetry event happening at an after party, though.
Boycott.
> and 2025 is shaping up to be another banger

> banger

“Fellow kids” vibes from the dinosaurs at Facebook and Zuckerfuck.

let's not be afraid to bring masculinity to work

up next: farting on the earnings call "here's what I think of your question Chadwick at Vanguard..."

/s

For a moment I thought it said "LlamaCoin" lol

I wondered why so much support on HN all of a sudden.

Facebook cheering AI and those dinosaurs cheering that "meteor shower" - same vibe!