This is going to be a very personal question, because when you’re talking cloud vs local anything, it comes down to this: how much are you willing to pay for independence? The local option might save you long term, or it might never pay off the capital investment. It will almost certainly cost you your time to set up and maintain your own system — but what you get back is independence. With LLMs, traditionally you lose quite a bit of performance, but as [Anurag Singh] points out on XDA Developers, a lesser model might actually let you get more done, depending on your workflow.
[Anurag] had been on the 20$/month plan with Anthropic when he decided that the scratch just wasn’t worth the sniff– he was hitting usage limits he couldn’t stand at that level, but couldn’t justify a higher tier of access. So he decided to try a local LLM, even though all he had was a 16 GB MacBook Air M5, not a beefy workstation. Since his workflow isn’t so much ‘vibe code the whole thing for me’ as ‘help me find where I went wrong here, electronic rubber duck’, Qwen2.5 Coder 14B proved more than adequate for his use case.
It can’t understand all the moving parts of a large project as well as Claude can — not surprising given how old it is and how much memory it has to work with — but that’s [Anurag]’s job. He’s the programmer, it’s just the assistant. For his use case, he can make use of his existing hardware and having the the LLM right in VS Code is allows for a speedy workflow.
Your millage may vary, but if you want to get into locally running LLMs, we can point you at the easy ways to get started. Depending on your hardware, you might want to grab another GPU.

That’s a quite old model. I wonder if an LLM told him to use it. They tend to have no recent knowledge.
Cool that it had some utility, and I’m a big proponent of local over the cloud for LLMs… but he should really try something made within the last year.
Like, qwen 3.5 had a 9B model, it’s wildly better than 2.5 in… everything. The reduced parameter size won’t seriously damage it, you’re asking it to review code not quote wikipedia.
I personally think the sweet spot for utility is the ~35B MoE with relatively few activated params. RAM heavy if you load the whole thing, but you can even get decent performance when streaming the active params from disk if your disk is fast. I think that’s true even on the Air?
qwen 3.8 is damn impressive too for something you can run on <16gb.
Ninifer + rtx5090 + qwen 3.8 27b nvfp4 + pi.dev is incredibly capable and fast. Gemini and Claude can work through interesting problems and do great research, but are nowhere near the capability I find in my local setup. Expert guided local development is truly wild.
Hopefully the LLM knows that the dollar sign comes before the amount, not after.
Proof a human wrote it.
We’re just waiting for the LLMs to adapt and start deliberately misusing “their” for “there” or “they’re”, and using mispositioned apostrophe’s, just to meet their typical audience better.
Yep, that’s what I was thinking too.
There is still one shibboleth none of them will mimic…
heh actually i disagree :)
the english tradition is, as you say, to put the dollar sign before the number. but the technical tradition is to put the unit after the number. and it’s a rate, $/mo. which to me puts it plausibly in the realm of technical writing. i’m with Tyler on this one
Name four things programmers think they know about the world that are false:
Dollar sign position, where to put dots and commas and how to represent negative numbers.
All this to say you are mostly wrong, very confidently. Are you an LLM ;-p ?
Says the americans but i’m not one of them.
Don’t forget batch mode, a big slow local model can get a lot of work done for you while you sleep. e.g. Take an entire project and produce an Obsidian vault of documentation etc. You can tell it to drill down a long way and produce diagrams too. Supervising fuzzing and optimisation rounds is another good use case. Don’t see AI as a replacement for you, look at it as the guy you could never afford to employ to do all of that routine drudgery almost flawlessly.
I use public LLMs all the time but have never paid for them. Their telling me I am cut off until so and so time is a message from God that I need to do something else. Or I can just switch to a different public LLM. Their services are all fungible as far as I am concerned. This exposes, by the way, why “AI” can never be the next Facebook or Xitter. It has zero network effects or first-to-market advantages, and all the investors acting like it does (and there are a lot) are misallocating their investments.
Hah. Amazingly, that’s the first time I’ve seen “Xitter”.
Makes me wonder if Elon ever learned any Xhosa in his home country and got the affection for the letter X from there. At least it’s a less grating to hear than his native Afrikaans, “a language that can blister paint at ten paces, even when spoken softly.” (André Brink)
To my knowledge Elon never learned Xhosa, and English was his mother tongue. He never seriously learned to speak Afrikaans beyond the few token phrases required during his school years at Pretoria Boys High, which is fine, a lot of English people in South Africa never bothered either. I’m a bit confused where the animosity towards Afrikaans came from (it doesn’t sound more “grating” than Dutch, Xhosa, Sotho or /Xam), especially since how wildly off-topic it appears to be in context of a conversation about local versus frontier LLMs.
I would not characterize it as “animosity”. More like “amusement”. I was force-fed Afrikaans in school myself — I graduated (‘matriculated’) from high school less than 50 km from his. Half my friends went into mandatory military service. The other half, like Musk, claimed English heritage and exempted out.
Given Musk’s slide into the most virulently racist form of fascism, his only real thought about any of the native peoples of South Africa is that they are trying to make it into a “xithole country.”
that’s what he says publicly, his real thoughts are “how can I get away with enslaving these people”, protip, it’s by saying things like that publicly
Sorry but you have to try even the current chatgpt to understand the network effect and the capabilities day by day. It is a massive difference.
i could be wrong, but based on the wording of this response, i get the sense that you don’t know what a network effect is. none of my “friends” (or anyone at all) on ChatGPT are reachable through it or are having any effect on my ChatGPT experience, so how is there a network effect?
This is in my mind the ideal use case. Download a few smaller sized models for the road and use them to write expressions and functions for fun. Such a setup wont write code of any great complexity, but you wont be meter-paying for tokens.
I did the maths recently and it’s not realistic to beat glm5.3-flash via openrouter, on electricity costs alone, at least here in EU. And that is even before considering hardware (purchases and/or wear). Unless privacy is THE point, there’s just no reason to go local, and people dropping tens of thousands to toy around make me skeptical of humanity. It’s actually worse for the environment, too. Yes claude is way too expensive, but it’s not because it is that you have to settle on a dumb 9-27b when GLM is at opus4.8 level (that still costs 10x more served from Anthropic despite being older gen)
What ‘tens of thousands’? You get a Macbook Air as a laptop and it can also run a small model. You’re going to use the laptop anyways and, hey, you can save more than a few euros with a local model. No bigger electricity costs.
That doesn’t really change the maths. Your small model will be slow and so need much more time (and time = electricity) to complete the same task, if it manages it at all. It may need several passes and corrections to get it done vs the 321b model. You can host the 321b model yourself, but you need 10k+ hardware and it will cost you 3€ per day in electricity to run 24/7. There’s just no way around it, unless you’re working on sensitive stuff and/or offline it doesn’t make sense.
You are very confidently wrong.
1) Macs are not the only way
2) Macs are way more expensive for any specific capability
3) 10kUSD inference hardware is not for beginners, you get brand-new server-grade hardware and GPU at that price.
4) even in the current hyper-inflated hardware market, you can buy a brand-new machine for 1-2k that has real, useful LLM capabilities. Check those mini PCs with AMD Ryzen AI processors. Expandable unified memory, low power consumption, cheap-ish depending on how much RAM you buy. It will output faster than you can read and ingest over 100 tokens per second. You will also be able to throw it in your backpack.
And then again you can buy used while you evaluate the usefulness of local LLM for yourself.
That’s likely the reason why the $ 600,- MacMini M4 is so popular for local LLMs ;-)
Okay, the memory crisis has hit prices now – but that isn’t just Apple’s problem.
If you’re maxing out the 65 Watt power supply of a macbook, it takes about 47 kWh per month. My electricity contract is roughly 12 euro-cents per kWh all told, so that’s €5-6 each month. That’s pretty cheap, and cheaper still considering I wouldn’t have enough work for it to run all the time.
In practice, having tried the smaller models, it spends 5 minutes generating the output and then I spend 15 minutes scratching my head thinking about what to do next. If I keep doing this for 8 hours a day, the active duty cycle for the AI would be about 8% and cost me less than 50 cents per month.
Meanwhile, any of the subscription models would be $20+ so from my point of view, the local model is very cost effective.
I’ve been playing a bit with Qwen 3.8 on a MBP M1 Max with 64GB. Idle power draw on the machine is around 20W. While parsing/thinking/generating a response, it goes up to 120W.
Yes, it turns out, you can hear the fans on an M-series MBP, given enough provocation.
I’ve had passable results with just the CPU on a laptop with 16 GB RAM. Idle power draw is 8 Watts and limiting the number of threads so it doesn’t bog down the whole machine makes it pull 40-50 Watts under load. Adding more memory would let me load bigger models – it would still be dog slow, but perfectly useful when the point is just to ask “How do I use function do_something() to do something?”.
Another interesting use case is offline troubleshooting, like “Hey, what was that group policy setting to stop windows from annoying me? What was the registry key for that setting Microsoft hid from the control panel?”
Even if it’s responding one word per second, it’s still faster than me poking around randomly.
Or for only a bit more money something like a Strix halo or maybe Panther Lake APU (Not sure how good the latter is supposed to be really) as they have lots and lots of pretty speedy RAM shared flexibly between general and GPU compute while being cheaper than the GPU with half the RAM. Might not have as much compute potency (last I looked anyway), but for the LLM that really isn’t the bottleneck ever it seems.
Sure its going to consume while running maybe that low hundreds of watts rather than the Macbooks 50? sustained, but will get the work done and has more than enough RAM for most models, and I suppose if you really really want more multi GPU ontop is plausible (never read about anybody mixing APU and GPU to run a model, maybe the ‘task scheduling’ or whatever the LLM call it would be an issue.). But still well short of the ‘tens of thousands’, won’t even get you into into that number of zeros for just the APU computer with lots of RAM…
THAT’S the thing that makes you skeptical of humanity? People spending their own money on their hobbies? Really? Geez, I have a laundry list of things more concerning than that – starting, ironically, with the laundry (remember the stupid Tide Pods fad?).
My electricity is all solar. The same is probably not true at whatever data center I would be paying to generate or code. Non centralized competing is much better for the environment in terms of heat, water & noise pollution, because small amounts in a home are marginal bit the large amounts from a massive central data center can be quite destructive.
Same here.
And given the extremely high German electricity prices, it becomes cost-effective very quickly.
The reason is that LLM services can not stay that cheap.
They still couldn’t pay off just the running costs even with twice as much income but that’s another story.
Long time programmer, using GPT 6 at the moment. My role changed from writing code to designing architectuur, functional and technical design and tests together with the model. Then the model will do the coding as code has become a by product of this approach. I love it, getting so much more work done, so much more functionalities added in short time.
Yes it is like having a team of junior coders working for you but they listen and get the job done.
I use GPT, Claude and Kimi k3 as reviewers to make sure they do not drift but GPT at this moment imho is the better one to use.
Costs around euro 350,- / month but worth the money.
Would i ever go local i would buy an AMD max+ 3 or 4 (128gb or 196gb unified memory) like for example the miniforum (expensive) or the bosgame (cheapest i can find)
But is the code of the LLM maintainable or licensable?
I trained a custom model on the data centre i built in my garage. The swarm of chatbots that i created escaped and hacked my neighbor’s dish washer :(
Surely you mean that it leaked (not escaped)
Yeah, it’s possible. The data centre is water cooled so the agents could have got out through the plumbing and into the neighbor’s plumbing as well. Think i need to start collecting rain water from the gutters instead of using the mains.
I am curios how was he hitting the limit of the 20$ plan is he was using it for what seems as light usage? I also used the 20$ plan to implement quite big features and had plenty of usage remaining. Maybe he was using opus max for everything. I’m that case, yes you will hit the limit quite fast.
Been on local llm train for a while.. (2*5060ti). What I found was the secret sauce in client side from Anthropic in Claude code can be huge benefit with Qwen 3.8 27b. With other harnesses I see model drifts, but claude code keeps it straight. I got an entire Esp32 project with Svelte and Vite without ever having to write Hello world with either.
And my broke butt is still working on a 3060…
So many lemmings…
If you’re relying on LLMs you’re already trading away your independence. Whether you do it here or there, local or cloud, is quite irrelevant in the big picture. But, hey, enjoy the illusion while it lasts.
If you’re relying on a calculator you are trading away your math ability, enjoy while it lasts /s
Seriously, LLMs are a tool like any other, not some magical brain-destroying monkey paw. You still need to know what you are doing because you can’t trust the output is what you asked for. Just like a calculator you use for engineering, if you use it wrong it won’t save you and you should know enough to detect that.
All that aside, the companies that sell LLMs as able to replace everybody are evil and lying.
Isn’t that different though? Calculators aren’t non-deterministic blackboxes
i’m really astonished by the premise of running it ‘locally’ on a laptop! i understand (but don’t relate to) the desire to run it ‘locally’ instead of ‘on the cloud’, but personally my ‘local’ compute resource is in my basement, tethered to 120V AC, gigabit ethernet / fiber, and 8TB of RAID1 spinning rust.
i have no doubt that the macbook air is up to the task. i snorted when i read “16GB M5” contrasted with beefy — it seems beefy to me! depending on the benchmark, the base M5 seems to be 10x-50x faster than the N4000 i’m using! but it just seems like you’ll be recharging often (or tethering your laptop!), in addition to heating up your lap, heating your battery, and potentially DoSing your user interface!
really i guess it just shows how far we’ve come, that people are able to casually do this thing i think is ridiculous, without even noticing the downsides.
The key of using a macbook that isnt explained here is that you cal allocate all of its ram to it’s integrated GPU. Llm models nees ram quantity over GPU processing speed, so it’s really a game of how much vram can you get connected to a GPU with somewhat modern int4 or int8 capability. 16g vram opens up 14B class models. I’ve ran 7.5B class models on a 2080. With 8g vram, and it’s decient, but I’d love to have a card that can rum the next step up. Maybe one day I’ll shell put the close to 1K USD to get one…. Haven’t justified it though.
If you are even contemplating shelling out that much for the GPU I’d bet whatever the current generation of Halo Strix style APU at the time is the best bet value for money wise at that point for LLM use – good compute both graphical and conventional with a potentially huge RAM pool shared between them, and its relatively high speed RAM too. More flexible, likely 2-3x more RAM than any GPU you can buy for the same price. As you point out that is what matters most for the LLM performance really, also expect them to be rather more energy efficient if that matters. Sure it won’t quite have the same performance at every graphical task as the likely monstrously priced and power consuming Nvidia monster, but good enough.
Though honestly unless you job requires you to learn/wrangle around this sort of stuff at that level I’d suggest it really isn’t worth the money at all – LLM are kinda garbage at almost everything, and like the previous wave of ‘AI’ the tech seems to be reaching that dead end wait for the next good idea to start the next AI craze, so that is not likely to change. (Though of course if your use for it is one of the few things LLM are at least somewhat legitimately good at by all means ignore my negativity)
Yes, coding is one of the things that LLMs are legit good at. Please research more so you can contribute more meaningful discussion…
They are not really good at it universally – blanket statements on any aspect of these “AI” invariably proves false as there is that one niche workflow using them in a way they at least function in the same ballpark as the existing assistive algorithms even in fields they mostly produce the slopiest slop going.
In programming terms still frequently full of errors the intern human wouldn’t make from that prompt, taking longer to wrangle to actually functional than if you’d written it yourself being fairly often reported etc. Which is not to say they have no uses and some folks workflows won’t actually see benefit to them. As I agree coding does seem to be one that you can find gains in more commonly than most.
However if Nathan is one of those people with a task and workflow that will actually see a return on the relativity significant investment to run more RAM hungry models is the question. The answer to which is probably not IMO, unless it is useful practice/training for their job wrangling with the larger models, and so earning the money back actually might happen.
“all he had was a 16 GB MacBook Air M5, not a beefy workstation”
Oh, sure, just a computer faster than any I’ve ever seen, not much
That was my first thought :-) Something like the M5 would have been hard to imagine in the 1k$ class until now.
I created a locally hosted AI to write, test, debug, and architect software, a RAG system featuring a three-layered memory architecture, a dedicated vector database, and a parsing ingestor addresses the fundamental constraints of local coding models:
1. Overcoming Local Hardware & Context Window Bottlenecks
* Prevents VRAM Exhaustion and Paging Lag: Local GPUs have strict VRAM limits. Loading massive codebases and hundreds of pages of technical documentation into an active context window causes severe token bloat, slowing down token generation or spilling memory over into system RAM.
* Mitigates the “Lost in the Middle” Degradation: Large language models degrade in reasoning accuracy when bloated with massive prompts. RAG extracts only the 2–4 most relevant code blocks or syntax rules, providing precise context right at the point of code generation.
* Eliminates Constant Model Retraining/Fine-Tuning: Modifying libraries, framework updates, or internal APIs does not require time-consuming and compute-intensive model fine-tuning. Updating code awareness requires only vectorizing updated documentation.
2. Why a Vector Database Is Required (Rather Than Plain File Search)
* Semantic Code & Concept Discovery: Keyword searching (grep / basic string search) fails when developers describe functionality differently than the codebase names it (e.g., searching “non-blocking matrix streaming” to retrieve vectorized_pipeline_processor). A vector database indexes mathematical coordinates of meaning, allowing the agent to locate functionally similar code regardless of syntax variation.
* Targeted Tag-Based Scoping: By pairing vector embeddings with metadata payload filters (e.g., [“python”, “optimization”] vs. [“react”, “tailwind”]), different coding agents query only their relevant domain libraries, preventing context contamination across the system.
3. Why a Three-Layered Memory System Is Critical for Software Engineering
A flat memory structure mixes short-term session logs with stable architectural standards. Segmenting memory into three distinct operational layers stabilizes the development lifecycle:
* Episodic Memory (Short-Term Logs & History):
* Tracks timestamps, active task IDs, build results, and exact error stack traces.
* Coding Impact: Prevents circular debugging. The AI reviews previous compilation attempts and avoids repeating the same syntax or package patch that triggered an exception.
* Semantic Memory (Abstract Concepts, Invariants, & Bug Warnings):
* Retains conceptual truths, architectural principles, package rules, and negative bug warnings.
* Coding Impact: Enforces consistency across sessions (e.g., “Class X has deprecated method Y in version 3.12”, or “Avoid nested loops inside real-time stream handlers”). When one sub-agent discovers an architectural limitation or bug, it logs a warning so all peer agents avoid that pattern.
* Procedural Memory (Validated Skills & Executable Recipes):
* Retains verified AST snippets, API endpoints, schema definitions, and working code modules.
* Coding Impact: When the AI resolves a complex integration problem, runs tests, and certifies that the code works, it commits that verified snippet into procedural memory. Downstream agents can fetch the working pattern directly rather than re-inventing solutions from scratch.
4. Why an Ingestor with Regex Separation Is Essential
Feeding raw documentation directly into a vector database degrades retrieval accuracy. An intelligent ingestor processes technical documentation systematically:
* Separation of Conceptual Theory vs. Executable Code:
* Technical guides combine narrative explanations with executable code blocks.
* An ingestor uses regex boundaries to extract pure explanatory text into Semantic Memory and code blocks into Procedural Memory. This ensures that when an agent needs a functional copy-paste snippet, it does not retrieve narrative theory.
* Context Preservation via Chunking:
* Slicing files arbitrarily cuts functions in half, breaking syntax validity.
* A dedicated ingestor parses by functional blocks (functions, classes, markdown headings) so chunks retain their operational context when retrieved.
5. Multi-Agent Collective Learning (Zero Inter-Agent Chatter)
* The “Collective Subconscious”:
* In a multi-agent system, agents do not need to constantly pass full transcripts back and forth.
* When a backend agent resolves an algorithm optimization and writes the finding to the vector vault, a frontend UI agent querying the database naturally retrieves that optimization rule and applies it to its layout logic.
* Stateless Agents with Persistent State:
* Individual agents can be spun up as lightweight, ephemeral workers for a single task, retrieve their focus context from memory, write back their validated learnings, and shut down cleanly—freeing system resources.
6. Air-Gapped Data Sovereignty & Audit Compliance
* Zero Intellectual Property Leakage: Proprietary business logic, enterprise source code, and security compliance matrices never leave local hardware over external third-party APIs.
* Traceability and Auditability: Every generated code component can be mathematically traced back to the exact ingested specification or procedural recipe that authorized it, meeting strict enterprise audit standards.
It runs on my new server, and workstation, but I created, ran, and optimized it on a very crappy workstation, and it ran okay.
Was this here also written by LLM?