Unified memory is the only reason Macs are so coveted right now for local AI. A single 192 gb ram Mac costs less than the equivalent in standalone GPUs.
What are the good use cases for very large memory amounts?
DeepSeek v3 for instance has 671B params, but should have the memory bandwidth of a 37B dense model with a batch size of one.
I also have a home server with 2x3090 and 2xA4000 (80GB vRAM) - yes it's a lot faster, but it's a pain in the ass to build, it takes up a lot of space, uses 10x the power, and honestly - cost about the same as my MacBook Pro.
The software will still see a single memory pool
I know im disagreeing with the article
I hope Apple sticks with the architecture. Even if its not very practical, its great to have it as possible.
This does not follow. Intel is shipping unified memory processors with CPU cores and GPU cores on separate chiplets but still sharing the same memory controller (on a third chiplet, for Meteor Lake and Arrow Lake). AMD is about to launch Strix Halo, a high-end mobile processor that is rumored to consist of one or two CPU chiplets and an IO die with a big GPU and 256-bit memory controller.
Edit: [2] The tweet doesn't even mention about UMA. The interpretation is entirely made up by Notebookcheck, I feel like I am reading WCCFtech again making stuff up.
I am just thinking if this allow Apple to do something crazy like 1024bit LPDDR5x or HBM3e memory solution.
[1] https://www.anandtech.com/show/21414/tsmcs-3d-stacked-soic-p...
[2] https://x.com/mingchikuo/status/1871185666362745227?ref_src=...
> Twenty-four x86-architecture ‘Zen 4’ cores in three chiplets
> Six accelerated compute dies (XCDs) with 38 compute units (CUs), each with 32 KB of L1 cache, 4 MB L2 cache shared across CUs, and 256 MB AMD Infinity Cache™ shared between XCDs and CPUs
> 128 GB of HBM3 memory shared coherently between CPUs and GPUs with 5.3 TB/s on-package peak throughput
It's hard to rule out their ability to create silicon that is a step change.
The original XBox (2001) had 64MB. I think my PC from 1998 had that.
The actual rumour from Kuo is that they’d move to a chiplet style design where the CPU tile and GPU tile are independent. This is actually in the article as linked.
That does not however mean that unified memory would go away. It’s just a new packaging system.
Such a cool name! And it says just what it does.
https://spinoff.nasa.gov/node/8965
https://spinoff.nasa.gov/sites/default/files/thumbnail0000_2...
https://pubs.aip.org/asa/jasa/article/92/4_Supplement/2376/7...
Body Electric supported the Convolvotron for visually programming VR simulations with 3D sound:
https://news.ycombinator.com/item?id=24266722
Did you ever meet (or better yet get a tour of Ames from) the late Ron Reisman, and see the virtual reality, flight simulator, and air traffic control systems his research lab developed?
Vertical Motion Simulator:
https://www.youtube.com/watch?v=5-lHcv_olkE
Marvin Minsky flies a simulator and wears VR goggles:
As a couple of others have mentioned, smartphones/tablets/laptops seem to be the driving force in UMA's spread.
However the M4 Pro has 256 bits wide, M4 max 512 bits wide, and M2 Ultra has 1024 bits wide. GPU workloads are latency tolerant and embarrassingly parallel, don't see how allowing a CPU to make random accesses is going to hurt the GPU much.
Is it really, though? It seems like almost every SoC small enough to be implemented as a single piece of monolithic silicon has gone the route of unified memory shared by the CPU and GPU.
NVIDIA's GH200 and GB200 are NUMA, but they put the CPU and GPU in separate packages and re-use the GPU silicon for GPU-only products. Among solutions that actually put the CPU and GPU chiplets in the same package, I think everyone has gone with a unified memory approach.
The prices they charge just to go from 16GB to 32GB of RAM is outrageous ($400 for Macbook pro).
Seems like they think Ultras aren't worth the investment, let alone building a true "unleashed" SIP.
Apple never says "hey what's the fastest and most powerful thing we can build for X price", they always box themselves in with space or energy constraints, so they never truly compete for the high end. The existing Mac Pro body was their chance to do that, and instead they put something designed for a smaller chassis in there.