Rendered at 06:41:07 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
txrx0000 4 hours ago [-]
It pays off instantly, because OpenAI/Anthropic can no longer see what I'm doing and that's worth a lot of money to me. If I am offloading some of my thought processes to a machine, I want to own that machine. And if I finetune the model, I can gain access to parts of thought space that are cordoned off by OpenAI/Anthropic/Alibaba/whomever due to their "alignment" efforts (i.e. alignment to the AI company rather than me). Otherwise, it's like if someone else owns a part of my mind and has a backdoor into my mind.
no-name-here 4 hours ago [-]
You can't run recent openAI/Anthropic models locally anyway, so wouldn't a better comparison be a different provider running Qwen or similar model? As then you can also compare against the exact model you'd have locally and any different data privacy of that particular provider?
koito17 4 hours ago [-]
GP's point is about "sending tokens to someone else's computer" versus "keeping the tokens locally". I think model capabilities are secondary.
In May of this year, I was running qwen3.6:35b-a3b on my MacBook (bought in 2024). Obviously not as fast as, say, running a model on Cerebras, but a year ago it wasn't really feasible to have a local model running on my 2024 laptop with vision support. (Concretely, I was passing apartment diagram pictures to Qwen and making it compare different apartments for which ones would feel the most spacious while optimizing for initial moving costs and other factors.)
This was back in May and I wouldn't be surprised if there have been significant improvements since then.
Overall, I think it's fair to compare a workflow like "use llama.cpp locally to upload some pictures and ask questions" to "open the ChatGPT app, upload pictures from your phone, and ask questions". Sure, you can't run a model like GPT-5.4 locally, but the model is mostly an implementation detail here. What a user will care about is: "when I go with the llama.cpp option, am I getting useful information from my conversations?"
3 hours ago [-]
no-name-here 4 hours ago [-]
Wouldn't the better comparison still be against an AI provider with better privacy controls, especially if that's what someone cares about (even if they don't care about whether they're comparing a 35 billion param model vs a x trillion param model)?
koito17 4 hours ago [-]
Users generally have no way to verify that a third-party provider, even if they advertise themselves as privacy-focused, will adhere to their own terms. This is similar to the issue of privacy-focused VPN providers that claim to not log user activity (and then end up leaking user activity). You can get proof of ~P, but rarely proof of P, and often times the proof of ~P is due to police raids, data breaches, etc., not something of the provider's volition.
What you can possibly audit is probably data sovereignty. For instance, I would not be surprised if Mistral's customers demand concrete evidence that their data is held within the European Union. But that is a distinct issue from training on input tokens.
v3ss0n 3 hours ago [-]
Deepseek 4 flash can run locally , and qwen 3.8-next-flash , they are already gpt 5.6 tier.
v3ss0n 4 hours ago [-]
You haven't tried DeekSeek v4 or GLM 5.3 or Qwen 3.8 Next?
You are missing out a lot.
Try that with Hermes or Opencode or Deekseek Harness , even Qwen 3.8 27b works really well for that kind of that.
I just ask it to install windows as a vm on my linux and install vs Community 2019 on it , and then build a legacy vb 2019 project on it. and sleep
When i wake up :
It installs Qemu , setup a vm , inside vm download and install windows 10 on its own , clicking next next next as needed , typing in things , writing powershell , python scripts , that run automatically after install by baking into CD that includes ssh server , reboot , it logins into ssh , trigger pythons script that continue installation of vs 2019 community , which includes a driver that click the installation steps , installs nuget , install all depedencies and then build the project into exe after i woke up.
That is with 100% pure local AI .
LargoLasskhyfv 2 hours ago [-]
Regarding DeepSeek, which I also like very much, have you tried https://reasonix.io ?
txrx0000 4 hours ago [-]
Technically true, but the delay between local and closed frontier is only a few months. And individual sovereignty / digital bodily integrity is almost priceless.
selectodude 4 hours ago [-]
Local frontier costs a half million dollars to run locally in anything higher than basically ternary.
txrx0000 4 hours ago [-]
Okay, that's technically true again, but local mid-tier like Qwen3.8-27B is only a year behind the closed frontier. I'm personally willing to be behind by a year if it gives me mental sovereignty against the big AI companies. They are extremely misaligned with me.
jrecyclebin 4 hours ago [-]
This was my thought as well. I have a local model monitoring my finances and personal wiki - things I wouldn't want Claude to touch - and the Qwen 3.5 9b handles it all just perfectly.
I also needed a new device anyway - and having this much system memory to run virtual machines has been amazing.
Am paying subscriptions as well tho lol.
catchnear4321 4 hours ago [-]
Your last line is what drives the point home, though.
Local isn’t strictly about NOT lab. It’s rapidly becoming apples (though not just macs) to oranges to compare the to.
Which is why the premise is silly. To be underwater it would need to be a real comparison. It’s not, and the claude fartifact doesn’t make it so.
ChickeNES 4 hours ago [-]
For me, I'm glad they train on my stuff if it improves the model. Hell, I've been using tons of muse-spark-1.3-contributor for this very reason (and because it's a decent model for a bargain basement price)
tyre 4 hours ago [-]
I'm curious what people are sending to Claude that is so secret. Claude knows about my interior decorating, questions about light bulbs, curiosity about what the Galactic Empire was even trying to do, unpacking SCOTUS decisions, shoe trees, Fed inflation history, etc.
What part of my brain is contained here? Sure, the conversations have back and forth (some have dozens of exchanges), but, like, that's not the secret to me. I don't think it can replicate me, and even if it could… okay?
Are you worried they're going to target ads? That the government will steal something? What?
Claude Code has information about my home server, but google or DDG would also have the broad strokes (torrents). I don't know. Maybe others are working on more sensitive things at home.
truncate 4 hours ago [-]
>> what people are sending to Claude that is so secret
Its the same point used against privacy. What's so secret you are doing that you need privacy. I think in the end, its about privacy and not trusting these model companies with your data. Facebook manipulated people behaviors with all the data they had, no reason AI companies wont someday decide to do that same, and they have far more intimate knowledge.
When it comes to coding, I also don't like the idea of them taking my money and potentially at same time potentially using as dataset generator.
AdieuToLogic 3 hours ago [-]
> I'm curious what people are sending to Claude that is so secret.
When Claude is used in a professional setting, any or all of:
Proprietary intellectual property (a.k.a. system code)
PII[0] of the employee, customers, or both
HIPAA[1] data known to a system
Internal communications not meant to be publicized
Sensitive data, such as SSH keys and the like
Pretty much anything on a machine which uses Anthropic/OpenAI native tools is a candidate to be compromised really.
Anthropic's goal is to commoditize intelligence. People who use their brains / intelligence for competitive advantage might not want to contribute training data for that goal.
AdieuToLogic 3 hours ago [-]
> Anthropic's goal is to commoditize intelligence.
And Google's original goal was to organize the world's information.
How did that turn out?
juiceland 3 hours ago [-]
Are you willing to bet that your lack of imagination for exploitation is precisely that of several multibillion dollar companies?
octoberfranklin 4 hours ago [-]
I'm curious what people are sending to Claude that is so secret.
The proof to the Navier-Stokes problem.
utopcell 3 hours ago [-]
I'm pretty sure they were sending a prompt for Claude to _find_ the Navier-Stokes proof, using ideas that have been publicly shared before online, but not necessarily used for the problem.
bix6 3 hours ago [-]
But are you pretty sure or actually sure?
utopcell 59 minutes ago [-]
They obviously did not send the proof to Claude.
poincareball 4 hours ago [-]
Tristan Buckmaster found out the hard way.
throwaway894345 4 hours ago [-]
I’m very sympathetic to this point of view but I also can’t remotely afford the hardware required to get in the ballpark of Fable.
jrflo 4 hours ago [-]
43 years to break even on Qwen 3.8 at 25% the speed of the API, lol. I like the idea of local models for really small tasks like automation/toolcalling, but it will probably never make sense for coding. I tried them and it was just excruciating compared to what you get for $100 a month from a subscription.
rlindsey123 4 hours ago [-]
Yeah it's surprising how long it would take to get back on those local models!
mcone 4 hours ago [-]
The idea that you need a new machine is pretty ridiculous. I bought a used HP Omen with a 3090 last month for $2k. 57t/s with Qwen 3.8.
usernomdeguerre 4 hours ago [-]
Agreed, I was also annoyed that the only params on the site were mac products. I run qwen 3.8 on a 12 year old asus and a 3090, 50tok/s. It's not even the only guest running on the box. For my usage profile (not running it 24/7) it's actually less expensive per-month than claude subscriptions.
redox99 3 hours ago [-]
I'm so happy for the two used 3090s I bought for $500 each after Ethereum mining ended. I even saw them for like $430 at some point lol.
rlindsey123 4 hours ago [-]
I've not heard of others running HP with it. Hows much RAM do you have?
mcone 4 hours ago [-]
This particular machine has 64GB, but the model is on the RTX 3090 with 24gb. Context is 156k with Pi mono.
ThunderSizzle 4 hours ago [-]
Claude Code is $100+ or else be constantly throttled. My usage on GHCP was gonna be $300+ a month.
I paid $1350 and threw an R9700 in an existing machine. That's a 4 month pay off or so.
Plus, I can feed it sensitive data all day and not be worried where it's going.
no-name-here 4 hours ago [-]
An R9700 has 32 GB RAM. Is your comparison against a similar size model? Or shouldn't you be comparing it against the cost of a hosted model matching the one you’re using locally?
chasd00 4 hours ago [-]
Not a fair comparison really. If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety. That has value a subscription does not.
> If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety.
When does the average person actually need to do that?
jerf 4 hours ago [-]
We've already seen frontier models refuse to answer almost any question that touches on computer security and be very likely to kick out biology and chemistry questions even if they aren't all that close to breeding dangerous viruses or making explosives.
I expect this is only going to get worse. "Censorship" isn't just going to be about who you vote for and which political party the model will say nice things about and which it is more likely to say bad things about. It's going to become about whether the hoi polloi are allowed to have effective AIs at all. Like the 1990s internet, AI has outrun a lot of power structures but that is not going to continue indefinitely.
ChickeNES 4 hours ago [-]
So you want to remove valid safeguards? And stop misusing the word censorship.
jerf 4 hours ago [-]
You asked a question. I gave you an answer. I seriously doubt that if you and I sat down together at a table and banged on this for an hour that we would come to the same definition of "valid". Ask 10 people, get 12 answers to that question. There's going to be a lot of motte & bailey in the next couple of years, where I just want an AI to answer questions about whether my code is vulnerable and people like you will be "Oh so you want an AI that can hack the Pentagon do you?" and it doesn't look like we're going to be seeing eye to eye on that one.
beachy 4 hours ago [-]
I just got some kind of cyber alert from Claude and was forced back down to Opus while I was trying to connect to a battery I own via bluetooth.
So I can certainly understand why someone would want the guardrails gone.
ChickeNES 4 hours ago [-]
Well I just applied to their cyber program, got accepted in two hours, haven't had that issue since. Ditto OpenAI. Why people treat these companies like sports teams instead of compute providers I have no idea, when I see underpriced compute, I take advantage of it.
kees99 4 hours ago [-]
"Need" might be a bit too strong, but I do want overly obnoxious guardrails not to stand in the way.
Case in point, last week I was poking Opus 5 into writing me some RPi-pico firmware for driving a small e-paper screen. Font was built in right into C code as hex constants. Space being tight, I asked if there is some clever compression that could be applied. Claude thought for good 10 minutes, then guardrail kicked in telling me that was "cyber", and refused to continue.
ChickeNES 4 hours ago [-]
Just apply to the cyber program? I got in in around 2 hours, and I'm just some hobby hacker, not some paid security consultant. It is a valid point though, I'm doing tons of systems and embedded stuff and was hitting the safe guards with Claude and Codex before getting into their cyber programs (hex REALLY triggered Claude in particular, which was amusing).
xnx 3 hours ago [-]
That can also be done with neoclouds.
hyperhello 4 hours ago [-]
I doubt it will ever be cost effective for the foreseeable future. The AI companies have astonishing amounts of compute and they’re effectively dumping it on the market.
epistasis 3 hours ago [-]
More than that, running hundreds of conversation streams at once is essentially the same cost as running a single conversation. And then you add on the secondary benefit of having the GPUs running nearly all the time rather than mostly idle...
Local inference makes sense for speciality needs, or very small models. But if your model is bug enough to span GPUs its excessively wasteful to hoard those GPUs for yourself without piggybacking hundreds of other conversations on top of all that memory bandwidth and matrix multiplies.
gruez 4 hours ago [-]
"If they are selling it for less than it cost to make, buy as much as you can."
-- Warren Buffett
taraindara 4 hours ago [-]
Only caveat is you’re buying time. Not a physical good. It’s only worth what you’re able to get out of it in that time.
tyre 4 hours ago [-]
For their current models, served directly from their infrastructure, they are profitable after training (which all present models are.)
I don't know when we'll have an open equivalent to Fable, let alone whatever (insane) hardware you'd need to run it locally.
3 hours ago [-]
rlindsey123 4 hours ago [-]
Is that a real quote? Golden if true
bix6 4 hours ago [-]
Fun feature: can you show some sort of list of the best combos? Eg shortest payoff time for best capability in various situations.
For each usage level, it lists the quickest pay-back in each capability class, with each model on its quickest machine and one click into the calculator to change the assumptions. Short version: at 1M tokens/day the best Sonnet-class option is Qwen3.8 27B on a Mac mini M6, 8.3 years. It only drops under a year if you're running agents at around 20M tokens/day.
bix6 3 hours ago [-]
That was fast!
At 7 tokens/s (Mac mini) you max at 600k/day so you couldn’t hit those higher amounts like 4M where it says 2 year payback?
Zetaphor 4 hours ago [-]
This tells me that the max throughput for the models I'm running on my hardware is lower than it actually is. Please allow us to tweak all the variables instead of locking me in to whatever rate you found by searching
rlindsey123 3 hours ago [-]
Thanks for all the feedback. You can now enter your own measured tok/s for any machine and model.
kjshsh123 3 hours ago [-]
On a purely monetary basis it probably never will.
You're competing against companies that get tax breaks, locate themselves optimally, and have large economies of scale.
Also, if it did, the hardware would be bought up, raising the price until there was no economic profit again.
If you can find a unique application for it then maybe?
ProjectArcturis 4 hours ago [-]
Local LLMs are not really about saving money, they're about autonomy. Choose the exact model you want, fine-tune it if you want, and no one can take it away from you.
rlindsey123 4 hours ago [-]
True - definitely agree!
redox99 4 hours ago [-]
The math is wrong, the tok/s is at least 2x that, at least with MTP and Q8 KV which you should always use. And the default tokens a day is ridiculously low at least for coding.
Having said that, it will never pay for itself. A simpler more absolute math is, if I buy a Mac and use it to sell tokens on OpenRouter, will I make a profit? And the answer is no.
afarviral 3 hours ago [-]
I want the autonomy but local models of the size I would have the means to host wouldn't be capable enough. What usecases tend to suit these smaller models that tend to produce incorrect or otherwise flawed responses often? Could they work for anomaly detection and what would a rough architecture look like?
itake 4 hours ago [-]
I have a home server running vibed applications. VPS host would cost $25/mo or $300/yr.
Mac mini can also build iOS applications. I think if you’re a mobile dev, you can have concurrent builds for your agents instead of everyone waiting on a single machine to finish.
rlindsey123 4 hours ago [-]
What models are you running on it? I'm also an iOS dev but I find I need more frontier models to get good quality code from it.
itake 3 hours ago [-]
I only use frontier models to vibe code iOS apps, as I'm not an iOS developer. I haven't tried the local models post qwen coder 3.5 release for all the reasons.
AFAIK, a limiter for iOS engineers (and AI agents) for concurrent feature development is the xcode environment and hardware limits. BE engineers can easily have 3 agents working on 3 different microservices (or gitwork trees), but iOS devs can basically only manage one version of the code at a time, due to externalized state (like derived data and bundle ids).
serial_dev 4 hours ago [-]
In the “The small print that isn't small” you describe all the disadvantages of running your models locally, but none of the advantages (just check the rest of the comment section for inspiration on that).
4 hours ago [-]
shadowpho 4 hours ago [-]
I like this calculator but it’s really wrong at least for dgx spark. I have one and I get 4x the tokens/s .
rlindsey123 4 hours ago [-]
Aw very interesting! This is great feedback - what model are you running? I'm keen to do more crowdsourced data as time goes on.
monksy 4 hours ago [-]
I wish you could put different setups on here. I have a couple of A6000s on an AM5.
rlindsey123 4 hours ago [-]
I'm keen to add a way for people to add community based reporting which would allow this. Would you want to see anything else on the dropdowns to be able to enter your data on?
harhargange 4 hours ago [-]
Also, i also use my gpu for rendering and learning and playing games.
4 hours ago [-]
QwenGlazer9000 4 hours ago [-]
Yeah no it does not pay for itself just comparing to cloud. Not at these prices at least, people far richer than you or I buy these things wholesale, no scalper, bought a significant amount at cheaper prices, and are wired up the ass with VC money.
The premium is not having your million dollar prize and career stolen by billionaires.
rlindsey123 4 hours ago [-]
haha 100%. We used to just rent our homes. Now we have to rent our intelligence
bpbp-mango 4 hours ago [-]
can you add RTX cards too please? 5090 and 6000
rubyn00bie 3 hours ago [-]
This is a bit weird because it automatically changes the model depending on the amount of VRAM available, and some of the smaller models are more expensive (presumably because they're being provided via OpenRouter by someone with some GPUs in a colo or smaller providers). It also doesn't allow changing the tokens per second (my 5090 can get like 75-130 tokens per second [assuming I can fit the model in RAM]); which, then results in woefully under-estimated limit on how many tokens a day I can consume.
Some improvements that I think would make this more useful:
1. Allow manually setting tokens per second, or as an alternative, let me jack up the number of tokens a day.
2. A sort of backwards flow "if you want to run this, at X tokens per second, with Y context, you'd have to spend Z."
3. Add support for configuring multiple RTX 6000 variants.
When I was making heavy use of DeepSeekV4-pro I was burning somewhere around 1.5 billion tokens a month, and that was just using it in my free time on random projects. It was something like $24 at the time because of the initial discount/promo period. I don't think there's anyway in hell I could ever run that (on current hardware) for less money.
I think the calculator shows from a purely financial standpoint what we all know... that yeah, it's definitely not worth it if money is your only concern. That (cost per token) will eventually change. Models will get better, more efficient, VRAM prices will come down, VRAM capacity will rocket upwards, and the economics of it all will change. It would just be really cool to have the calculator show me exactly how cheap they'd have to get for it to make sense.
I need to finish up some work and make dinner, and if no one else beats me to it (anyone is welcome to) I'll ask Fable or Opus to knock that out.
gfody 4 hours ago [-]
should throw in a tt-quietbox
0xbadcafebee 4 hours ago [-]
It pays for itself very quickly if you do 24/7 generation. Use an AI agent that orchestrates other agents working on many things at once constantly. If speed is a factor, you'd not buy a Macbook, you'd buy dual RTX 3090s. About the same price, but at least 6x faster than M5 Max. The benefit of constant generation is you can do a lot more research, coding sub-agents, experiments, etc in parallel when you're not "at work". You end up getting a lot more work done than if you only sit there babysitting sessions.
jayle 2 hours ago [-]
[dead]
01HNNWZ0MV43FF 4 hours ago [-]
I was just gonna throw a beefy Ryzen into an ATX chassis. I don't want to pay Mac prices
v3ss0n 4 hours ago [-]
Besides from privacy:
I already making twice now.you own the hardware and the price had doubled since i bought. Almost tripled.
You missed the opportunity and i have 4 of those awesome machines. Cry on.
I sell those to business who need local air gapped requirments and I make a lot more money!
I can run the alliterated models where none of the service prvoider even dare to provide.
THose benefits outweights a few K.
And show me an api provider that allows me to run 10x agents concurrently for 5 days straights .
ChickeNES 3 hours ago [-]
> And show me an api provider that allows me to run 10x agents concurrently for 5 days straights .
Any of them on a Max/Pro plan as long as you are smart about model selection? That's my main objection to local inference, I'd need a whole rack of GPUs to do as many things in parallel that I can do for $400 a month. I do plan on setting up some local inference hardware, but...RAM and GPU prices alone are $$$$
In May of this year, I was running qwen3.6:35b-a3b on my MacBook (bought in 2024). Obviously not as fast as, say, running a model on Cerebras, but a year ago it wasn't really feasible to have a local model running on my 2024 laptop with vision support. (Concretely, I was passing apartment diagram pictures to Qwen and making it compare different apartments for which ones would feel the most spacious while optimizing for initial moving costs and other factors.)
This was back in May and I wouldn't be surprised if there have been significant improvements since then.
Overall, I think it's fair to compare a workflow like "use llama.cpp locally to upload some pictures and ask questions" to "open the ChatGPT app, upload pictures from your phone, and ask questions". Sure, you can't run a model like GPT-5.4 locally, but the model is mostly an implementation detail here. What a user will care about is: "when I go with the llama.cpp option, am I getting useful information from my conversations?"
What you can possibly audit is probably data sovereignty. For instance, I would not be surprised if Mistral's customers demand concrete evidence that their data is held within the European Union. But that is a distinct issue from training on input tokens.
You are missing out a lot.
Try that with Hermes or Opencode or Deekseek Harness , even Qwen 3.8 27b works really well for that kind of that.
I just ask it to install windows as a vm on my linux and install vs Community 2019 on it , and then build a legacy vb 2019 project on it. and sleep
When i wake up :
It installs Qemu , setup a vm , inside vm download and install windows 10 on its own , clicking next next next as needed , typing in things , writing powershell , python scripts , that run automatically after install by baking into CD that includes ssh server , reboot , it logins into ssh , trigger pythons script that continue installation of vs 2019 community , which includes a driver that click the installation steps , installs nuget , install all depedencies and then build the project into exe after i woke up.
That is with 100% pure local AI .
I also needed a new device anyway - and having this much system memory to run virtual machines has been amazing.
Am paying subscriptions as well tho lol.
Local isn’t strictly about NOT lab. It’s rapidly becoming apples (though not just macs) to oranges to compare the to.
Which is why the premise is silly. To be underwater it would need to be a real comparison. It’s not, and the claude fartifact doesn’t make it so.
What part of my brain is contained here? Sure, the conversations have back and forth (some have dozens of exchanges), but, like, that's not the secret to me. I don't think it can replicate me, and even if it could… okay?
Are you worried they're going to target ads? That the government will steal something? What?
Claude Code has information about my home server, but google or DDG would also have the broad strokes (torrents). I don't know. Maybe others are working on more sensitive things at home.
Its the same point used against privacy. What's so secret you are doing that you need privacy. I think in the end, its about privacy and not trusting these model companies with your data. Facebook manipulated people behaviors with all the data they had, no reason AI companies wont someday decide to do that same, and they have far more intimate knowledge.
When it comes to coding, I also don't like the idea of them taking my money and potentially at same time potentially using as dataset generator.
When Claude is used in a professional setting, any or all of:
Pretty much anything on a machine which uses Anthropic/OpenAI native tools is a candidate to be compromised really.0 - https://en.wikipedia.org/wiki/Personal_data
1 - https://en.wikipedia.org/wiki/Health_Insurance_Portability_a...
And Google's original goal was to organize the world's information.
How did that turn out?
The proof to the Navier-Stokes problem.
I paid $1350 and threw an R9700 in an existing machine. That's a 4 month pay off or so.
Plus, I can feed it sensitive data all day and not be worried where it's going.
Idk about the quality of this setup but just pasting it here as an example. https://explainx.ai/blog/heretic-llm-abliteration-guide-2026
When does the average person actually need to do that?
I expect this is only going to get worse. "Censorship" isn't just going to be about who you vote for and which political party the model will say nice things about and which it is more likely to say bad things about. It's going to become about whether the hoi polloi are allowed to have effective AIs at all. Like the 1990s internet, AI has outrun a lot of power structures but that is not going to continue indefinitely.
So I can certainly understand why someone would want the guardrails gone.
Case in point, last week I was poking Opus 5 into writing me some RPi-pico firmware for driving a small e-paper screen. Font was built in right into C code as hex constants. Space being tight, I asked if there is some clever compression that could be applied. Claude thought for good 10 minutes, then guardrail kicked in telling me that was "cyber", and refused to continue.
Local inference makes sense for speciality needs, or very small models. But if your model is bug enough to span GPUs its excessively wasteful to hoard those GPUs for yourself without piggybacking hundreds of other conversations on top of all that memory bandwidth and matrix multiplies.
-- Warren Buffett
I don't know when we'll have an open equivalent to Fable, let alone whatever (insane) hardware you'd need to run it locally.
For each usage level, it lists the quickest pay-back in each capability class, with each model on its quickest machine and one click into the calculator to change the assumptions. Short version: at 1M tokens/day the best Sonnet-class option is Qwen3.8 27B on a Mac mini M6, 8.3 years. It only drops under a year if you're running agents at around 20M tokens/day.
At 7 tokens/s (Mac mini) you max at 600k/day so you couldn’t hit those higher amounts like 4M where it says 2 year payback?
You're competing against companies that get tax breaks, locate themselves optimally, and have large economies of scale.
Also, if it did, the hardware would be bought up, raising the price until there was no economic profit again.
If you can find a unique application for it then maybe?
Having said that, it will never pay for itself. A simpler more absolute math is, if I buy a Mac and use it to sell tokens on OpenRouter, will I make a profit? And the answer is no.
Mac mini can also build iOS applications. I think if you’re a mobile dev, you can have concurrent builds for your agents instead of everyone waiting on a single machine to finish.
AFAIK, a limiter for iOS engineers (and AI agents) for concurrent feature development is the xcode environment and hardware limits. BE engineers can easily have 3 agents working on 3 different microservices (or gitwork trees), but iOS devs can basically only manage one version of the code at a time, due to externalized state (like derived data and bundle ids).
The premium is not having your million dollar prize and career stolen by billionaires.
Some improvements that I think would make this more useful:
1. Allow manually setting tokens per second, or as an alternative, let me jack up the number of tokens a day.
2. A sort of backwards flow "if you want to run this, at X tokens per second, with Y context, you'd have to spend Z."
3. Add support for configuring multiple RTX 6000 variants.
When I was making heavy use of DeepSeekV4-pro I was burning somewhere around 1.5 billion tokens a month, and that was just using it in my free time on random projects. It was something like $24 at the time because of the initial discount/promo period. I don't think there's anyway in hell I could ever run that (on current hardware) for less money.
I think the calculator shows from a purely financial standpoint what we all know... that yeah, it's definitely not worth it if money is your only concern. That (cost per token) will eventually change. Models will get better, more efficient, VRAM prices will come down, VRAM capacity will rocket upwards, and the economics of it all will change. It would just be really cool to have the calculator show me exactly how cheap they'd have to get for it to make sense.
I need to finish up some work and make dinner, and if no one else beats me to it (anyone is welcome to) I'll ask Fable or Opus to knock that out.
I sell those to business who need local air gapped requirments and I make a lot more money!
I can run the alliterated models where none of the service prvoider even dare to provide.
THose benefits outweights a few K.
And show me an api provider that allows me to run 10x agents concurrently for 5 days straights .
Any of them on a Max/Pro plan as long as you are smart about model selection? That's my main objection to local inference, I'd need a whole rack of GPUs to do as many things in parallel that I can do for $400 a month. I do plan on setting up some local inference hardware, but...RAM and GPU prices alone are $$$$