Cheap AI Token Resellers: The Secret Ingredient Is Fraud

[Matt Lenhard] has an interesting writeup explaining exactly how fraudsters offer access to cutting-edge AI models at a tenth of the price. Perhaps unsurprisingly, the secret is to get tokens for free from anywhere they can and by any means necessary. Then wrap them in a pretty relay API, and sell access to it.

Relaying tokens is not by itself a shady practice. That distinction belongs to services that obtain tokens fraudulently, opening the door to selling them at rates far below market value. This practice is widespread and profitable, in part because the abuse is so hard to pin down and stop.

One source of tokens is free credits on new accounts. New accounts are spooled up as fast as possible, hammered until they’re empty, then it’s done all over again. Another method is to sign up as pay-after, possibly with a stolen card, and simply ensure the account has no valid payment method once the bill comes due. Or set up a temporary card, pay some minimum up front and consume as much as possible, then initiate a chargeback. It doesn’t matter if individually each of these doesn’t amount to much before they get flagged, because it’s being leveraged relentlessly on a massive scale by automated systems.

There are the shadier methods, too. Fraudsters don’t just target providers directly. Consumer software products with AI features get reverse-engineered, then the back ends hammered for all they are worth. Poorly-coded support chatbots can be highjacked into serving fraudsters’ traffic instead of just their own. It doesn’t actually matter where the tokens come from, after all. As long as the fraudsters are obtaining them for free (or at least below their costs) then it’s profit.

That last point is one [Matt] zeroes in on with advice on how to mitigate this abuse. He goes into detail in his writeup but what it comes down to is recognizing that it’s a numbers game. Fraudsters depend entirely on obtaining tokens for free, or nearly free. So just like using an AI to keep phone scammers tied up, anything that raises friction increases the fraudster’s costs, in turn encouraging them to find an easier target.

25 thoughts on “Cheap AI Token Resellers: The Secret Ingredient Is Fraud

  1. wow, AI companies having problems with people who abuse and ignore agreements on large scales

    I feel like there’s irony here

        1. Exactly. They all trained their models on stolen materials and scraping websites without any consent or reimbursement for all the bandwidth, but cry when someone distills their models. Ugh.

          1. If it’s theft to train models on publicly available content then it’s also theft to run a search engine or even learn personally from said content.

          2. Also when you visit a website, there is often some sort of system that reimburses the website owner for the costs your visit has caused. Taking that out of the equation breaks lots of social contracts legal or otherwise. I like AI quite a bit. I just wish it were done above board from the beginning. I’m not super fond of how it was all open source until profit looked possible and then a bunch of opportunists came out of the woodwork.

          3. However let me also point out my own hypocrisy here. I hate advertisements. I always use ad blocking software. The internet is unviewable and unsafe without it, but while I block ads, I never take the content to then turn a profit while blocking those ads. Its a lame difference to some I’m sure, but I think it’s different. I didn’t use ad blocking software until those shaking banners came along.

          4. “If it’s theft to train models on publicly available content then it’s also theft to run a search engine or even learn personally from said content.”

            This is complete nonsense.

            The search engine comparison is just ridiculous: they don’t copy the content outside of what has been fair use forever (short summary) and they absolutely aren’t presenting other people’s work as their own.

            The “learn personally” comment is just anthropomorphizing what an AI model does. Humans learning is not model training: when you learn something it is in context of your unique experiences, because you’re a unique individual. You can’t extract out the exact information because you don’t experience it that way.

            Model training (on non public domain data) is just copyright infringement. The only reason courts don’t see it that way is because AI companies are rich.

          5. Pat, I’m not sure you appreciate how LLM training works, based upon your criticisms, they simply don’t match the reality. In the simplest form, if what you said were true, the final model would be exabytes.

          6. “In the simplest form, if what you said were true, the final model would be exabytes.”

            No, you don’t have to store complete copies of something for it to be copyright infringement. Training set extraction attacks on LLMs are well established, and honestly even if a company said “we’ve prevented that”, the simple fact that they had to prevent other people from extracting copyrighted data from their models tells you that the models are storing copyrighted data.

        2. @Pat
          Sorry, I have to disagree here. The models are trained over a lifetime in a short period. Context of individuality has NOTHING to do with factual information. Now, if I give you my opinion on what I just read, sure, but if you ask me a question and I tell you what I know on it, that’s based on what I remember about it. I’m not going to say “oh yea, so I was picking my nose and I was thinking about Chinese for lunch while I was reading that the sun is about 5 minutes further in the sky than we see it.” Also, to answer the search engine query, yes when I search for something I am gathering information on it. Modern LLM’s scrape articles and use search even post training for answering prompts. The way a model is built and the unique data they store is the part people have a problem with. Just like generative art. We imitate everything we see. A blind man, born blind, can not draw a dog without reference. How could he know what it looks like without ever seeing one?

          1. “but if you ask me a question and I tell you what I know on it, that’s based on what I remember about it. I’m not going to say”

            You’re not going to say it. But that’s actually how your brain works. Which is why eyewitness accounts and memory is so fallible. Your brain is not a computer. It does not store and fetch data, it links it to everything it experiences and because you’re an individual, the storage in your brain is unique.

            “Just like generative art. We imitate everything we see.”

            We don’t store copies of what we see. Even if we really, really, really want to, our brains simply don’t store information like that, and it’s demonstrable that LLMs can – otherwise training set extraction couldn’t exist, and it does.

  2. From what I understand, especially in mainland China (where a lot of frontier models are unavailable via normal means), token providers also are fronts for competitors. They sell tokens on the cheap and meanwhile capture the conversations that their users have in order to train their own models.

  3. I just finished the linked blog post/article. Wow. I’m almost speechless. It looks like customers are mostly big companies distilling Claude and Codex. Some people are making millions off of this. I’m not surprised at all that they’re also capturing the data. It’s a classic double dip. American corporations have done well teaching the rest of the world how to double dip.

    I also can’t help but think that this is all a result of attempting to keep China behind on the AI race. Why do humans keep inventing enemies and trying to keep other people down? If knowledge were free, this would be less of a problem. It probably would not disappear because people also tend to be greedy and hate work, but any step forward is better than pushing others backwards. What a mess.

    1. “Why do humans keep inventing enemies and trying to keep other people down?”

      This is way older than humanity. We really shouldn’t be so shocked by basic adversarial behavior. If the next industrial revolution is even a hundredth of the magnitude the hype suggests (I very much doubt the hype myself, but some of it is real..) then of course there is going to be a big fight about who gets the controlling share. Humans, aliens, hyperintelligent birds, silicon brains… I don’t think it’s realistic to expect any of them to ever overcome avarice (as a group, it’s possible in individuals). Maybe I’m a cynic.

      1. Overcoming avarice is difficult, maybe impossible, but is necessary for our survival. This is what makes open models so attractive. By taking the profit from just owning something that’s essentially not capable of being owned you remove a lot of froth — pointless human activity — from that thing which opens the door to all sorts of progress. (Where would we be without open source code? There’s a few of us left who know what the world was like ‘before’ — expensive, closeted, restricted, primarily tools of existing power structures.)

  4. AI is a hell of a word to call “computer metering”. Luckily the epoch of Luther brings us the notion that stored data can be copied easily, AI dystopia is not as monolithic as it is destabilizing

  5. Any normal company would have a security office that understands where the value is. So, perhaps the value of an AI company is not the tokens? Perhaps it does not matter that tokens get stolen. Investors don’t seem to look for AI companies to be profitable, yet, perhaps tokens used is the more important KPI.

  6. Selling cheap is nothing new. Pawning in general is not new either, if there is demand there will be supply.

    IMHO, you make sound as if other industries are not full of resellers of stolen (or bought on the forced cheap) goods. Let’s also not forget the classics, home builders using cheaper inferior substitutes sold for the inflated prices still. Cheaters can be found everywhere where there is promise of good profit made : ]

  7. AI access tokens is nice, but not the complete picture. The environment for the AI, how is accesses your data and can get context over your code is also important. I doubt these cheap api call are properly vs code integrated is my guess.

  8. My wife’s credit card number was recently stolen and the bank did not invalidate it quickly enough (took them a few days since they don’t cancel your existing card until your new one comes in the mail). I was surprised to see it was almost all spent on ai compute.

Leave a Reply

Please be kind and respectful to help make the comments section excellent. (Comment Policy)

This site uses Akismet to reduce spam. Learn how your comment data is processed.