back
449 comments
In the last year, I have bought an M3 Ultra Mac Studio with 512 GB, a Macbook Pro M5 MAX with 128 GB and an RTX 6000 Pro. I have spent around $25k so far, not including electricity. I figured worst case scenario I can sell them in the next year and only take a haircut as opposed to losing my entire investment.

In comparison to just spending for tokens, the tokens would have been much cheaper and much much faster. I've been running against Gemma4:31b, Qwen3.5 and 3.6, and getting local LLMs to solve AMC 8/10 math questions and it's about 10-100x slower than just doing it online. When I tried it with ChatGPT late last year, it took about one night and $25 to solve about 1000 questions. Using my RTX 6000 and M3 Ultra and Gemma4:31b on both, it answered about 40 questions in 7 hours and I haven't checked how good the answer is yet. At 800 watts (600 for RTX and 200 for M3 Ultra) and running for 7 hours, it solved around 40 questions.

At the very least I'm going to try to sell my M3 Ultra if I can find a reliable place to sell it without getting ripped off by scammers.

This is, sadly, obvious and inevitable in retrospect.

The two major drivers of inference costs are GPUs and electricity. You can't get cheaper GPUs, but you can make existing GPUs not sit idle, and you do that by utilizing them 24/7, processing user B's request when user A is thinking, and handling many requests in parallel, neither of which you can do as an individual. You can get cheaper electricity... by moving, and it's much easier to move your AI workload than to move yourself.

This is a completely different dynamic than renting houses or apartments, as you can't really rent out the same house to different people at different times of day.

I’m not usually one to ask this because learning to do a thing can be fun, but why exactly have you spent 25 thousand dollars on getting an LLM someone else made to answer maths exam questions?
I got an RTX 6000 pro too. I like running locally, I've learned a lot more than if I had used an API and there's less worry about overspending tokens. I accidentally spent $100 on claude api in like 2 days because I didn't know what I was doing.

The problem is that while one these gpus is a huge improvement over a laptop or a single 3090, you very quickly wish you had more. I would buy a second one, but I did the math and realized that with the current crop of models, 2 Blackwells doesn't buy me any new capability that I didn't have with one. So I would need a 3rd one. And when I buy a 3rd one I will feel like I want to running a higher quant, so then I will want a 4th.

>> find a reliable place to sell it without getting ripped off by scammers.

This is a real problem and why I've just about given up on ebay or fb marketplace, esp for computers. If you are in Canada though sellit9.com is a great solution to having to deal with sketchy buyers.

If you're in a decent sized city, you should be able to find a local buyer on Craigslist or FB Marketplace... Beyond that, for higher value, smaller items like your M3 Ultra, I would talk to your local police department and/or library to see if you can do the exchange there. Larger libraries usually have a police officer on site or nearby, and the PD office near you may also provide a "safe" exchange location... I'd bring a monitor/keyboard/mouse so you can demonstrate the system working properly.

YMMV but between your nearest PD office and Library, you should be able to use one or the other for your exchange of goods/money. The biggest thing I've sold is a mid-range video card during late covid (I managed to get a better one via newegg shuffle) so I sold the old one (RX 5700XT -> RTX 2080) to make up the difference a bit. I just did the exchange at the Starbucks near me for that.

I have three m3 512gb units and want a fourth to run an exo set up. Like you, I am worried about scammers. Let’s discuss if you still want to sell.

https://calendly.com/ryanwmartin/open-office-hours

I don't think this changes the final conclusion - but have you considered calculating against depreciation -- i.e. figuring out how much your M3 ultra is worth today, and only charging yourself for the delta? In my mind you might even have made money on the hardware.
I looked into the M3 Ultra 512GB Mac Studio before it was discontinued and the as best as I could determine it just wasn't worth it... yet. The GFLOPS and memory bandwidth just arne't there even though it can hold a much larger model in memory.

But the trend here is interesting. I think by 2030 you'll be able to buy fairly cheap hardware that is currently $10k+. I don't know what this does to the trillions invested in AI data centers because the next NVidia architecture after Blackwell will essentially half the value of purchased cards overnight.

I'm not convinced Apple has yet pivoted the Mac Studio line towards this market and the expected M5 Ultras in Q3 2026 will likely be an incremental improvement rather than big leap forward but I'd like to be proven wrong.

Which of these has been the most productive for you? Sounds like you've enjoyed the RTX6000 the most?
I administer a simple AI server in the office, which just uses a single RTX 5090 but is able to serve ~80 people throughout the day. I'm impressed by Qwen3.6-27b's capabilities in agentic coding/tasks so far. Devs say it's not much different from Sonnet 4.6 on many tasks (sometimes it even outperformed it), 40-60 tok/sec, up to 260k context. The server cost about $10k with all the bells and whistles.

I spent a lot of time researching/adding/benchmarking many custom modifications to the software stack and its settings to make the server optimally handle the load with just 1 RTX 5090 without losing quality, but it's still not enough, and the wait times in the queue are getting longer. We're at the limits of the hardware, and I'm out of tricks.

The experiment was kind of a success, and the CTO agrees we should scale it. With our own infra, we could run agents 24/7 on everything. Currently, a lot of use cases for the cloud providers are completely blocked by PII/trade secret concerns (our infosec department doesn't buy the "zero retention" promise), plus you don't have to think about billing/budgets/etc. anymore.

Now I can't decide how to scale it. On one hand, I'd like to run larger models. And we have the budget to buy, say, 8xH200. But in many benchmarks, the larger models that do fit in 8xH200 comfortably and can serve many parallel requests with acceptable speed/quality don't seem to outperform Qwen3.6 that much in agentic coding/tasks to justify the price.

So another option is just to buy a bunch of RTX 6000s and scale horizontally instead: run a copy of a midrange LLM like Qwen3.6 on each GPU. It's cheaper and easier to scale/replace, but then we'll run into problems running larger models in the future if we have to, because of no NVLink support (say, if Alibaba & Co. stop releasing ~30b models and/or ~30b models start falling behind 400b+ models considerably)

Does anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.)

UPDATE: Launch was a success! 400K+ views, and multiple companies reached to use my IP. Read more here

It seems that he managed to get what he wanted from the hardware and I'm happy for them.

He said something interesting at the beginning of his post, he compared the cost of the hardware to the cost of his time based on his FAANG salary. Which is an interesting way to think of this, but the rest of the article didn't make me understand if at the end he did save money/time based compared to just rend on the cloud.

Also, outside of the power cost, hardware has other costs too, you need to operate it, maintain it, set it up, etc. all that require time. I mean, even the process of figuring out if it had a good enough ROI compared to cloud, takes from your time (collecting data, analyzing data, etc etc).

This article appears to lack any reason for "needing" this beast, or any real comparison with alternatives, both of which are required to answer the question posed in the title. It's a summary of how much they spent and some light anecdotal comparison to what they might have spent on cloud services, but clearly they didn't do an exhaustive hunt for value.

The real question is whether or not they could have done whatever it is they did with less hardware. Is there a business idea here that could have been proven on cheaper hardware that could be upgraded as demand increased? Is the expected ROI there based on future earnings?

Absent any indication that this was needed in the first place, I can only conclude that it wasn't worth anything.

> The mentality shift of renting vs. owning the gpus is huge. When renting, each experiment costs money and I had to ask myself is it worth it. When owning, it feels like not running experiments is costing me money.

I feel like there is some very deep generalizable wisdom buried here.

I did the math at least on a Macbook pro, and for inference it's definitely not worth it.

- https://www.williamangel.net/blog/2026/05/17/offline-llm-ene... - Discussion: https://news.ycombinator.com/item?id=48168198

I can't imagine spending $48K on a home GPU server, but I did just splurge and buy a PC with an RTX 5090, specifically to hold the largest models you can fit in 32GB. It's a top of the line PC with water cooled high end CPUs, 64GB RAM, RTX 5090 for $5K. To me the jury is still out whether this was a worthwhile investment, but I do expect to use this machine for a decade. I don't run it at 100% power (it's mostly idle, except for times when I'm training or doing batch inference). It has the nice property of being blackwell generation, similar to the machines we use at work.

It just scares me to own a box that is $48K in my house, especially if it breaks, or gets stolen.

So some things have changed since this rig was first built (2024). The most relevant is that $6800 RTX 6000 Ada 48GB has arguably been supplanted by the $9500 RTX 6000 Pro 96GB.

The Ada has a memory bandwidth of 960GB/s. The Pro has 1.8TB/s and about 40-50% better performance so is at least equivalent in processing power, much better in memory bandwidth (important for inference) and can hold larger models on a single card.

I've considered buying a rig with 1-2 6000 Pros for similar reasons but I want to see what happens with this year's Mac Studios with a likely M5 Ultra. Macs have a shared memory architecture whereas NVidia segments the market based on max memory where the biggest consumer card (RTX 5090) has 32GB of VRAM but still excellent memory bandwidth (1.8TB/s). A RTX 5090 rig will still trounce a Mac Studio seems to be the conventional wisdom. Despite being able to hold larger models and being able to chain Mac Studios on TB5, their lower memory bandwidth (~900GB/s) and lower overall GFLOPS mean they still come out behind.

That being said, the current Mac Studios are relatively long in the tooth, being released in 2024.

I'm still not sure any of this is really wroth it because things are still changing so fast. I think there's a decent chance of a number of large AI companies going bust in the next 2-3 years such that you'll be able to buy enterprise AI hardware at cents on the dollar, a bit like how Google bought data centers in the post-dot-com crash.

But anyway, nowadays I'd be looking at the RTX 6000 Pro as the sweet spot, having anywhere from 1-4 in a single server.

The electricial issues the author mentions are interesting. I hadn't really thought about the max amperage on a residential circuit. In a DC, these would typically operate on three phase power and much higher overall amperage. I wonder if there's a device you can buy that can combine multiple residential circuits into a single power source for a server this power hungry?

This is interesting but I am unsure how you make money out of this home setup, I would imagine if one would be offering consultancy to a business the business would make their own equipment/infrastructure available, which would also give a better control of their data. But perhaps I am thinking this because I am thinking about very big companies. Then, on very small business I don’t see they having the use case with the budget to match the need. So is this for specific services for medium sized businesses? Can you explain this a bit?
Other things people spend "too much money" on:

- muscle cars, with all the stuff, driven occasionally.

- boats, that don't get taken out much

- gamer x, where x=system or laptop or keyboard or mouse or desk or glasses or mousepad or speakers or ... usually with "> too much RGB"

- children

$48k for something constructive even if ai related? no problem, refreshing even.

Hi! Thank you so much for posting this! I got back luck/timing when I tried, so happy it made it to the front page! (I am the author)
'If you google “plugging a PC into multiple outlets”, you get lots of warnings that if you even consider such a setup you will instantly burst into flames. So I hired a professional PC builder make sure it was safe.'

Not really sure how that makes it safe but OK!

Nice analysis, I would have loved a short overview of the kinds of experiments that were running on the machine (I know the results are given).

I find the "independent researcher" business model quite interesting. In the linked post he writes """DFT is a proprietary training algorithm, however, I’m currently offering a beta for a model training service where I will train your model for you using DFT.""" I'm curious how successful this is. Essentially market some AI breakthrough as a service instead of publishing a paper like my academic brain is trained to do.

As an aside, one thing that I always loved about our field was that the startup cost for many business ideas was "a laptop, internet connection and some some grit". In the age of AI it's quite a bit more and I feel one of the sad side effects of this is that it crowds out poorer and younger developers.

Sounds fun/stressful/rewarding. I'm most interested in the update at the end though 'Launch was a success! 400K+ views, and multiple companies reached to use my IP.' I too, like probably 1 in 5 of the people reading this, think I have figured out some major problems with LLMs (context and computation research) but have wondered the best way to 'release' and get value out of it. I can see training being a little easier in that you release weights against a known model arch but not the training code. Wy stuff is all custom layers though. Any thoughts on a release strategy where you need to release the layer code for people to see test weights/the benefits?
"The point of buying the server wasn’t to save money, it was to build something cool." In the end, this is always the real answer - one that I'm sure we can all agree is the 'correct' one too.
This is a difficult calculation to make because you wouldn't rent time on the exact same system in the cloud. Depending on what you're running, a bigger server with better inter-GPU interconnects in the cloud might complete the task so much faster that the additional per-hour expense is more than covered.
FYI: If you're in a similar situation, think very carefully before you build your own. The $17000 might sound like a lot; but when you take into account your time and risk tolerance, renting might be a much better solution.
Just curious OP (if you're the one posting) -- what do you mean by independent researcher? What are you researching and are you making $$ from it or are you living off previous built up savings? Seems like an interesting path. What research have you looked into so far?
It reminds me of the Crypto mining bubble - I look forward to buying my heavily discounted Mac Studio soon.
I built a very similar server myself [0] with a similar setup. I run different models for different purposes, but the primary one currently is kimi 2.6. I run kimi as the orchestrator model and then qwen, Gemma and others for specific tasks (sometimes loaded dynamically based on the task at hand), all exposed through the pi harness. I also use Hermes for some personal repeated tasks which connects to the same models, hosted on my local Mac Studio.

I am not even going to pretend that this is financially reasonable option. I simply wanted to have a local models. Maybe down the line, as cloud models become less subsidized, I might benefit from having a local setup, but for now, it wasn't the most prudent financial decision.

But one big benefit is that I never have worry about my account being randomly banned nor I have to worry about running out of quota. I still use codex and opus for some specific tasks, but as tools are improving, I need them less and less.

[0] https://x.com/synopsi/status/2024235558193811778?s=20

Great article. I'm about to embark on a similar journey.... Doing a ton of AI development right now. Don't need a server, but a very, very high end workstation is super appealing to me right now. Looking at $50-$80k. 1TB RAM. 2x RTX Pro 6000s. 64 core Threadripper Pro. As many 4tb or 8tb nvme drives as I can stuff.

I envision NixOS at the core... then everything I need virtualized on top with KVM/QEMU. Maybe a dual boot setup with Windows for gaming and Flight Simulator (but I could virtualize that too with easy GPU passthrough.)

Lingering questions I'm working to figure out:

- Will 2 RTX Pro 6000s run on a 1600 watt PSU? Not sure how much higher I can go without calling an electrician. (standard US home.)

- Assuming I plop this into my home office, should I expect the PC to run significantly hotter than my current rig? (3960x threadripper, 128GB RAM, 1600watt psu, overclocked and watercooled 4090.) My water temp, measured at radiator, is about 60c at peak load. (This is the only number I care about, as this is what I have to consider to be comfortable sitting next to it.)

So the answer is: "TBD if I can actually make money to pay this back"
And here I felt like I was wasting money on an Intel B70 to run LLMs locally.
Stuff like this + OpenClaw with Mac Minis a while back is sort of exposing a probable local AI flywheel waiting to happen.

Someone needs to solve proper distribution of packaged GPUs with some Tesla-like wall connector for a consumer grade box that is plug and play.

Maybe John Ternus ends up doing that at Apple since they sit closer to this consumer profile.

> Because of this I got a motherboard with slow GPU interconnect. It’s good for running many small experiments in parallel (which is my main use case) but horrible for any models split across gpus.

:( you paid a professional pc builder and you weren't told this?

(For reference I’m talking about the DFT post from the same blog.) I love that ML is still in the “gentleman researcher” stage where relatively small amounts of startup capital can buy a ticket into frontier research.

For a lot of research questions 6 GPUs is even overkill.

It’s one of the reasons I’m skeptical of the “trillion dollar supercluster” idea [0]. I think what we need is more reasonably smart people investigating medium-sized problems. A “GPU middle class” you might say.

[0] https://situational-awareness.ai/racing-to-the-trillion-doll...

The other advantage of the local GPU is that you are not feeding your data into cloud providers. I'm not sure how much you can really trust Anthropic and OpenAI not be improving their models based on your input.
Buy one of these next time, https://tinygrad.org/#tinybox. At least geohot knows what he is doing.
Any kind of fixed capacity usage model seems to be a dead end. Paying per token might seem like an exploitative arrangement at first glance, but it's a luxury if you are experimenting or deploying greenfield.

Provisioned capacity is a really high end thing. I feel like you'd need to be spending more than $1000/day on tokens for this model to make any sense. You lose a lot of flexibility once you start dumping capital into specific pieces of hardware. Maybe start by renting the GPU server for a few days...

That's a nice problem to have. I can't afford a $48K GPU server, even if I worked as a developer since 25 years ago, because I live in the wrong place.
Just curious - What exactly are you using that rig for? I see that you said research work. Are you building a product or training models? I ask because whether something is worth it or not depends largely on what you get out oof it and how you value what you get. It's perfectly fine to leave a FANG job and go for, say, pottery hobby. What gives you happiness and your value system - these will qualify your decisions.
I ve seen already one question like that in the thread. But I rephrase it slightly sharper. Did you consider renting out you setup to vast.ai and if so, how much money it can generate per month deducting electricity.

Also, sorry for the noob question, is not such server generate enormous amount of heat? You did not use any special cooling system?

The idea is similar to maintaining on-prem vs cloud

Cloud is optimized for development velocity but its nature of high margin business eventually makes on-prem more promising

It could be too late but it might be worth looking into tax saving if you have a business. Depreciation of asset is a loss and may deduct your income. (I'm NOT a tax expert)

Could turn the system into a multi-seat gaming rig with 6 separate gaming seats, using loginctl on Linux.
People doing economics with the cloud GPUs, of course cloud GPUs are going cheaper. But also, is generating tokens all you do with your computer? I can play games on DGX spark and also do LLM inference, so sometimes the economics work out, apart from having fun with it.
Jensen Huang said 'the more you buy, the more you save,' and you actually took it personally.
Quick tip for people who want to experiment with local models: A lot of the common smaller models are also available on openrouter or other services. Dirt cheap.

I know it's not the same. But a lot of people buy expensive GPUs, just to find out they have no real use for smaller models.

I have four old 24gb Nvidia cards. They're not great but they're not useless either. The problem is that I haven't really figured out a good way to actually use them.

Genuine question; would anyone here recommend any specific motherboard to best utilize these cards?

And the net result is a way for LLMs to use more variety in their writing style.

Didn't Sam Altman create LLMs to cure cancer and stuff? Why does their writing style matter as long as the information they are conveying is accurate?