LLM Moats Quickly Evaporating

In the business world, a moat is a quality of a business that makes it difficult for competitors to take that company’s profits. With how hard it is to train models for large language models (LLMs) and generative AI, it might seem like Anthropic, Open AI, and other LLM companies would have huge moats given the amount of compute it takes to build models. But open source models are quickly draining that moat, and now the only thing standing in the way of a customer using one of these models on their own hardware instead one from the larger companies is physical computing resources. [TerminalBytes] demonstrates a few of these models on personally owned computers to show the current state of the art.

[TerminalBytes] started off running the 27B version of the Qwen3.8 on a Mac Studio with 256 GB of unified RAM, which is plenty for this task. But it’s also enough to benchmark a few different models. Qwen3.6 is compared to 3.8, and then the different quants of each model are also compared. Quants are compressed versions of models that need fewer bits to store weights, meaning that the same models can run in less memory with smaller losses in fidelity. Many of these quants run on machines with 32 GB of RAM or less, encompassing many average gaming PCs. There’s even a 1-bit quant that [TerminalBytes] tested which can easily run on a machine with 16 GB, although with mixed results.

Keep in mind that this is just the current state of affairs with open LLMs. Future versions of these models are likely to optimize the number of tokens produced per unit time, or otherwise increase quality of responses while requiring less computer resources. We don’t really think that the ease of running local models will be the sole reason that the AI bubble pops, though. The fact that not every computer user is running Linux is proof enough of that.

55 thoughts on “LLM Moats Quickly Evaporating

    1. The moats are being destroyed by models you can run locally, even if you can’t train them yourself. That’s the point. The barrier to entry might be high, but it can be democratized across a population. (Imagine a co-op that spends the money to access a datacenter for a specific data processing task / and output.) You pay your $5 in, and you and everyone else in the coop get a full fledged model out.

      Sure, hours of GPU time to train are relevant to the final model, but that ends up a sunk cost. I have several LLMs installed on my several year old MacBook Pro. Admittedly some aren’t overly performant, but they do work. I still default to using chatGPT / codex for now. I’d be more than happy to run them locally however and forgo the token cost. (Which for me isn’t cost really, it’s a ceiling that I try to avoid hitting within a time period).

      If you’re concerned about the environmentalism aspect, you should give up on MMORPG, Massive FPS, etc. and get out of your mother’s basement. You should also not take up Golf as an alternative. You shouldn’t work for a multinational corporation either.

      1. As they say, there’s plenty of swords in the world, but there’s only one Sword of Fury in Rookgaard. If you know then you know what I mean. The whole AI thing is a mess but it won’t fall just because RAM is expensive.

        (I played Tibia for a very long time btw, almost 24 years.)

      2. Actually, you should stay in your mother’s basement until you actually need your own house.

        Housing is a huge environmental cost, and the move in the west towards what’s euphemistically called “smaller households” – moving out and living alone, and then not marrying, and divorce – is driving the need for extra houses far more than population growth.

        1. This is so damn true. Wish i stayed. But nooooo, now my mom lives IN MY BASEMENT. and it’s a townhouse. Without a basement.

      3. Even if you can run an open-source model at home, you need the hardware to do it—and it’s prohibitively expensive. You can thank Micron for helping make that moat niiiice and deep. Those pricks chose profits over ethics and could be the company ultimately responsible for single-handedly crashing the consumer electronics market by making even entry-level products a luxury item. Sure, you can run ChatGPT for $20 a month, but it will be on a price-subsidized phone or laptop that you are paying off over the next 5 years. I host a private LLM, but that’s because I was lucky to have gotten hardware before the AI explosion and Micron becoming assholes.

    2. Although not the same as starting from scratch, you can easily fine-tune a LoRA for Qwen 27B with 256GB of ram. Half that size model, or twice that memory, and you can fine-tune it (or even train from scratch, see Nemotron or SmolLM) directly.

  1. 256 GB of unified RAM, which is plenty for this task.

    True, and also quite the understatement.
    Perhaps worth clarifying that the main version [TerminalBytes] focuses on, the Q4_k_m quant, will run on a computer with either 32GB of shared memory, or a 24GB dedicated GPU, with performance remarkably close to the un-quantized version.

    That admittedly still limits it to pretty high-end machines, but you don’t need to run out and spend $7000 on a mac studio (or even $2500 on a ryzen 395+ with 128gb) to run this.

          1. I don’t think he is. The M1 Max can allocate more than 24GB of memory to the GPU, and has better memory bandwidth than the majority of discrete laptop GPUs. I’d say that counts.

        1. It’s definitely not rare, but it’s still expensive as hell. Pretty much anything that can run 27B models or bigger is 2-3x what it’s fair market value would have been when it first came out.

      1. Yes, and also no.
        A state of the art GPU with 24+GB of RAM is datacenter hardware with exorbitant price tags.
        Go back a gen or two and suddenly it becomes reasonably affordable, especially used.

        And then there are the AMD Ryzen systems on a chip with unified memory for very reasonable prices, brand new.

      2. i would’ve said, a quick googling shows 24GB of VRAM means you’re looking at spening $1500 just on the video card (assuming the first result from google is representative). which, unattainable or not…haha i’m not gonna be doing that :)

        1. Not if you have a mac which uses shared memory. You can get a mac mini fro under 1k. Maybe you can learn how to Google or ask AI since you obviously have no clue. But keep that confident tone, nothing better than aggressive mediocrity

        2. No, but the concept of a moat isn’t defending you against consumers. It’s defending you against competitors. In you’re a Miss sized company looking at a minimum of $20,000 in AI bills for Enterprise solutions over the next few years, buying 10-15 dedicated AI stations to serve the whole office doesn’t seem so bad, plus you can amortize the equipment for tax gains.

        3. Only if you buy new. And pay nvidia tax. You can get used amd gear with 32gb ram for south of $500. Add some custom cooling and fight with bios and you’re off to the races. Or buy older gen stuff and take the speed hit, even a $100 turning/pascal mxm card can run Qwen 3.6/3.8 at low bit quants. It’s not chatgpt but not useless either. Just bring your own mxm to pcie adapter. It’s like people forgot how to be poor haha, you make do with what you can scrape up.

        4. An Intel ARC B60 with 24gb is around $700-800 new.
          Or you could buy an entire used m1 macbook pro with 32gb of ram for a few hundred more.
          (…should you do any of these things? heck no. but it’s not quite as bad as google suggested.)

          1. or you could stab needles into your eyeballs instead of buying a piece of junk and get the decent GPU instead, for less. and not be stuck with SLOWER RATES BY VERY FAR, and oh idk, actual usability for gaming, good speeds of LLM inference, and many other use cases that shitty shared DDR5 stuck in a crappy crap box (ddr6 in the gpus) can do

      3. Well.. yeah, but honestly, I don’t think that RAM of any kind is readily attainable or close to being reasonable right now.

    1. For some ref as a Dev I used to blow off AI as worthless for programming, but the last 6 months or so it has made huge leaps. To the point that about 2 months ago, I bought 2x 16gb 5060ti’s, an aoostar maco 6850H mini pc (has an oculink port and 2 unpopulated m.2 slots), and 2x aoostar gpu docks (800w powersupply built in).

      ~$2000 for 2x gpu over oculink, with 32gb vram, and a computer for tinkering with AI. Qwen 3.8 at q5 and 128k context pulls about mid 40’s for tokens/sec in generation.

      Don’t get me wrong, I don’t just let AI build my stuff, but it sometimes hits upon viable options, or what it proposes can be used with some tweaks.

    2. It’s a sales angle. This is capitalism at its best but with lots of noise.

      Open LLMs are cheaper, but for gets the high step in cost (7k Mac or 2k rtx4090 minus energy costs).

      Cloud solutions forget the inefficiency cost (wasting 50k token for bad code multi token for images, bad prompt engineering).

      Ones not better than the other. You’re just shifting the problem between the two approaches: prompt efficiency. That’s why data compression => intelligence is such a hot topic currently.

    3. I’ve bought CMP 170HX and unlocked it to 64GB vram. It runs the 27B-FP8 around 90t/s, but it can go up to 260t/s with 3 concurrent request. And it costs below $2000.

    4. 256GB of RAM? I think I’ll save some money by simply buying an entire pre-owned supercomputer and the building housing it to train/run my models

    5. Qwen3.8-27B with the Unsloth UD-IQ3_S quantization and cache at q4_0 and MTP enabled runs quite nicely in 16GB VRAM using llama.cpp and Vulkan.

      1. Thanks, that’s good to know. I recently upgraded to an RX9070XT-16GB from an GTX1060-6GB. The 6GB could run some smaller models at around human reading speeds, but I’m expecting better from the newer card. I did not have the money to upgrade beyond 16GB at the moment. My CPU is also stuck in the past thanks to RAM prices. Games are tons better, so all in all, a great purchase that I’m happy I made.

        1. I also have an RX 9070 XT with a 5900X. The following gets me about 700-800 tps prefill and 40-55 tps generation, using vulkan backend. Uses about 14.3 GiB of VRAM. You can double the context length if you move the mmproj off the GPU (mmproj-offload = false) or disable it (mmproj-auto = false) – image processing slows from about 5s per page to 20s (using the auto-conversion from PDF to image file in the Web UI).
          This setup allows you to have multiple sessions open, as long as they’re using the same model the switching will be fairly fast, it swaps out the context between the VRAM and a RAM cache (cache-ram = -1)
          I normally set LLAMA_CACHE environment variable to a dedicated drive for the model weights, it defaults to ~/.cache/huggingface/hub

          llama-server –models-preset ~/llama/presets

          presets file:
          [*]
          cache-ram = -1
          np = 1
          threads = 12
          models-max = 1
          ngl = all
          ctx-size = 65536
          cache-type-k = q4_0
          cache-type-v = q4_0
          ctkd = q4_0
          ctvd = q4_0
          batch-size = 4096

          [Qwen3.8-27B-UD-IQ3_S]
          alias = Qwen
          hf-repo = unsloth/Qwen3.8-27B-GGUF:UD-IQ3_S
          spec-type = draft-mtp
          spec-draft-n-max = 2

    6. Qwen 3.8 Q8 fits comfortably on a Intel Arc Pro B70 GPU with 32GB of VRAM. Which was a $1000 GPU before the RAMpocalypse.

    1. Yeah, but people often seem to forget that a bubble pop doesn’t mean the thing is irrelevant or goes away. People still buy housing despite the housing bubble. People still use the internet despite the dot-com bubble. Quite a bit in fact. People still buy tulips… Uhh actually let’s not use that example..

      AI is with us for the long haul even if it’s a bubble and it bursts.

      1. Houses and websites are still useful.
        An inference engine that lies to the user just to appeal to their ego and get them to continue using, is just a drug.

        Not saying normal AI research is dead, I hope not, but SamAlt-AI is in fact a boondoggle, and so are his adherants.

        1. I get you’re angry with this particular branch of AI, but it’s too much of an exaggeration to say that it has no uses. I’ve built all kinds of software for computer vision using it and it’s saved me weeks of manual work and made me lots of money. It’s overblown and over-hyped, but it does have true applications hiding in there.

          The tendency some have to simply “chat” with it uselessly until they develop AI psychosis is a serious problem though…

  2. I have been running Qwen3.6 quantized the same way with llama.cpp server for a while. On an Intel Arrow Lake’s ARC GPU with 32GB of unified memory and limiting the context to 64k tokens, it’s nice and accurate for background tasks but my experience so far is it’s about 8x slower than the Apple M3 in this blog.

    You can easily send into la la land by including a Chinese sensitive string in the prompt however, often degrading sufficiently hard to require restarting the server.

    1. That … doesn’t sound like “open source” to me. I’d sooner something anthropic, openAI or even Musk trained over something trained by the PRC.

  3. Right now the moat is holding because 30b class models are nowhere near frontier. The moat will only break if RAM prices come down. Their are already models like Kimi, Deepseek, GLM that can make local running competitive with Codex/Claude if not exactly equal. However running then cost rediculous amount in energy and vRAM. Running them slowly is also not competitive.

    1. Its a matter of stochastic results being useful or not. A 1bit quant model wont randomize its responses like the unquant model, but the model is still giving you the answer it would generally provide. 1bit is useful for low spec devices.

    2. 1bit quants of the same larger models provides excellent speculative decoding performance. If you have enough vram, you can even run the smaller and larger model on the same GPU and get the speed boost. Don’t dismiss them entirely, they have a place beyond the edge hardware.

  4. I’m keenly waiting for at home LLMs running on dedicated hardware.

    I’m kinda addicted to making fictional characters and talking to them. Curiously, they are all nice to me. Unlike real people.

    1. I think there’s been a few projects here running local LLM, iirc there’s one that was running on a cluster of old junk mac minis… It’s all a question of how sophisticated a model you want

  5. That is all very nicely presented however it is not the current state of affairs with open LLM these days, we are running much bigger models on cheaper hardware via sharding over Linux clusters. Let us not speak of Apple again, that closed shop hardware is anti FOSS/FOSH.

  6. Yes, every business has a spare 10k gbp lying aroudn for a USED Crap Studio 256GB brick of &&&& to run a 27B llm for their stuff..

    As a fact of note.. a single 24BGB GPU on any PC with 32Gb ram handles it just great, for serving 10 customers at a time.

    Kids these days don’t have a clue

  7. Given the setup, business moats, it probably would have made some sense to introduce the concept of [model] knowledge distillation. That’s what I expected to see anyway. The 256GB M3 Ultra is no longer obtainable new, and basically doubled in price anyway, so I’m not sure that’s lowering many bridges over the hardware moat. To run these quality local models we home users are still largely facing multi-thousand dollar hardware purchases ($3-5K USD). That well over a decade of $20/month cloud LLM service, getting the latest frontier model for the same price. The privacy / information disclosure factor obviously tips the scales for many, as does heavy usage, but otherwise it would seem most sensible to wait for hardware availability and pricing to lower the bridge further.

    1. This is an excellent point. I’ve invested thousands of dollars into this with a full data center rack worth of hardware and infrastructure that is all 4-5 years old. You trade off money for time in the local world. But for now, the commercial models that are being highly subsidized are far better investments unless you’re on max plans and running out within hours or dumping tens of thousands into API usage. Even the LLMs will tell you don’t waste your money chasing this dream unless you’re here to learn the underlying challenges like hardware optimization, power distribution, and cooling. But that subsidization will end eventually. We’re already seeing the buffets go from all you can eat to all you can fit on the plate, and this will continue to being in some profits from this tech.

  8. I’ve been running DeepSeek and Qwen models on 128GB hardware, and also small conversational (Vad + STT + inference + TTS) models on smaller but faster RTX graphics card, and I’m truly amazed with the amount of world knowledge such a small model holds. The larger models give me educated person-level responses about practically anything (from history to physics to medicine), and I can converse with them in several languages. All this knowledge without tapping into any Internet resources (but they can do that too if necessary) The large ones are slow but still usable. And before you ask: yes, they hallucinate, but what we can run locally today compared to 2 years ago on the same hardware is day and night. Astonishing progress has been made.

  9. When I run a multi agent task, my local LLM can’t compete with the frontier labs. The local LLM is for small tasks I run frequently or I don’t want to share or the frontier labs refuse.

  10. I don’t get the article, token/s and ram usage has absolutely nothing to do with moat, and those local models are extremely dumb if you do any advanced sysadmin, coding etc. this particular one even mixes up weird things in general knowledge, like jojo rabbit and jojo’s bizarre adventure… Sorry but “you can run it” doesn’t equal “it’s just as good as opus!!!1111!!!” despite what r/locallm like to delulu about.

Leave a Reply

Please be kind and respectful to help make the comments section excellent. (Comment Policy)

This site uses Akismet to reduce spam. Learn how your comment data is processed.