ShieldFont: Bludgeoning AI Scrapers That Disrespect Robots.txt

In the more innocent days of the World Wide Web you could simply put a robots.txt file in the root of your website that search engine indexing bots and similar would consult for the indexing wishes of the site owner. In this brave new world of LLM training data indexing such pleasantries are however rarely respected, leaving site owners to resort to increasingly more involved ways to bludgeon so-called AI scrapers, with ShieldFont being one of the most recent methods.

Its basic functioning is detailed in the white paper, explaining their use of ligatures. These are normally used to join multiple graphemes or letters into a single glyph which are rendered in the final text. By substituting about a quarter of the words in a text with such ligature-based versions in an intelligent, dictionary-based manner, the HTML version – as typically parsed by a scraper – will read as grammatically valid but nonsensical text, while the rendered font version will look normal.

Naturally, there are some disadvantages to this, such as screen readers for the visually impaired needing to also use the rendered font version, and it’s just as effective on legitimate search engine indexing bots. That said, if you apply this to static, archived content, or content marked as ‘do not follow’ in said robots.txt, then it might just be one way to make ChatGPT and friends spit out really funny output in the future now that the novelty of wood glue on pizza and eating rocks has somewhat worn off.

While LLM scrapers can adapt to this by also parsing the rendered text, this makes the scraping effort significantly more expensive. Together with maze traps like Nepenthes and Cloudflare’s offerings that seek to keep these scrapers busy scraping dynamically generated content through infinite linked pages, the tools available to combat the menace of these scrapers keep developing.

56 thoughts on “ShieldFont: Bludgeoning AI Scrapers That Disrespect Robots.txt

  1. This is a “screw the lot of you sub humans with disability” solution. This is terrible and I hope being so accessibility unfriendly it is never used.

    Not just the b/Blind are affected. It spreads over a LOT of accessability “hacks” making the internet even more only suitable for the abled.

    Anyone who was set a locked theme font is especially hit, so all the dyslexia font users, you’re SOL.

    My own adjustments for poor eyesight needing high magnification and locked sans fonts is corrupted text. I also see the corrupted version on my tuned dark mode browser extensions, hear them in my reader, can’t use the reader version on my phone in safari, or see them in Links which I often use at CLI too.

    Horrible. Unethical. Non-accessable friction that locks the web away from more and more with difficulty who try to function in an already hostile environment. The web has already become a mess of endless captchas after capthchas, but this is the worst possible thought out implementation for fixed font users, I can only hope this fails fast and dies in a fire in a huge way.

    The worst part, those of us not fully abled see the corrupt shit-text with no understanding why such stupid words are there. There is nothing to explain – you are seeing shittext because we think the disabled are shit.

        1. Yeah, this decrease in accessibility is mentioned in passing in the article, but it should in the headline and the first, last, and most frequent thing addressed in the article.

          I hate what AI is doing to the web. This is not how we change that.

      1. Ah Tom, likely not enough so that people going “wow, that’s awesome” stop for a second and consider the deeper problems here around accessibility that’s taken decades to get in place being stripped away.

        I didn’t even touch on how it potentially could hit YOU once the crims work out how utterly awesome it is for prompt injection attacks via email on your AI laden phone now that white-on-white text is mostly dealt with. Do we now need (insert company logo) to add ligature fonts into their end user protection guardrails?

        Have you considered how ads using WebGL/WASM on other sites could abuse this to get around content moderation tools?

        Not a competent English speaker? Maybe you Speak Spanish or you’re Canadian and don’t speak English. You reach a shittext page and you will see the bad translation. Here, hover over this google link to see https://shieldfont-org.translate.goog/?_x_tr_sl=auto&_x_tr_tl=es&_x_tr_hl=en-US&_x_tr_hist=true or https://shieldfont-org.translate.goog/?_x_tr_sl=auto&_x_tr_tl=fr&_x_tr_hl=en-US&_x_tr_hist=true if you hold over the non-changing word you see a translation back to English.

        The project’s github, one of the issues, https://github.com/isaqueseneda/shieldfont/issues/9 a request “please help us make this work on Jaws” so, what about the Linux screen-readers too? What about the fixed font users? The elderly? The non-English speakers?

    1. And I’m pretty sure I could bypass this in a few minutes by having the bot OCR screenshots from a headless browser and correlate the bounding boxes with the HTML elements.

      So it’ll only affect disabled people, bots will get past it fine.

      1. I’m sure that any decent programmer (or even a suitable LLM) could code around this new font. Their trick, apparently, is in the font description, which is part of the font. Obtain the font, and you can work out the substitutions (and, because you don’t have the font natively on your system, web pages that use it /have/ to force you to download it).

        Even the authors of TFP say
        [blockquote]”The font must be sent to the browser so it can display the original words. Anyone who downloads it can therefore inspect it and work out the substitutions.”[/blockquote]

        1. They can, but it is computationally expensive to do that on-scale.
          Being able to suck up all these sites for pennies is what makes this current model make any sense.

          Normal users don’t pull down and process 100k pages a second.

          Making it 10x or 100x more work for any individual user is meaningless in their eyes.

          Making it 100x more computationally expensive to scrape content is HUGE.

    2. Yes and no – its pretty trivial to correct, so other than the additional hurdles faced being difficult, and something you might require able assistance to fix it isn’t really that big a deal – its also breaking normal able users from using the site normally too…

      I can see why folks feel the need for actions like this, when you are having the crap kicked out of you and nobody is doing anything but enabling the abusers then kicking back is inevitable.

    3. AI has already made the Internet borderline unusable.

      Fighting bots claws back some usability/stability for some users.

      You can have content you can’t easily read.
      Or you can have no content at all because it was all forced out of existence.

      Sometimes the lesser of two evils is the only viable option.

      Suggest a better one.
      Anything that allows for machine readable content (which your system requires right now) is disqualified because that is the whole point.

      1. You can have content you can’t easily read.

        If you can’t use it, you don’t have it.

        Or you can have no content at all because it was all forced out of existence.
        “no content at all” is labeling a philosophical perspective as a factual claim.

        Sometimes the lesser of two evils is the only viable option.
        every evil is the “lesser” for somebody. The evil that is lesser for you may be greater for somebody else. You risk becoming the “greater evil” for that other person.

        … All that said, misappropriation is always harmful. But, malapropriation is never the proper solution to misappropriation.

    4. I agree, and if I have any spare time I will use it to spoil ShieldFont’s day. People have delusions of grandeur if they think their text is so special that it needs protecting, at the expense of disabled people. Furthermore AI system do not copy human writing at all, and I have a lot of experimental data to prove it, they emulate the statistical norms of human writing, but their output is significantly, and measurably, different from skilled human writing.

  2. “This is a “screw the lot of you sub humans with disability” solution.”

    As another one with vision issues, I agree. I finally had my cataracts removed but retina problems are mostly not fixable. I would include blurry phone optimized images, along with hard coded background colors, cough HAD. I heard at one time from a designer I know that contrast and color trade offs were taught in schools.

    Line has to go up, quality are not our job.

      1. A webscraper is not AI, and will use much less resources. It’s basically just a script that downloads as much of a website as possible. It’s not until later in the process that it’s ingested by an AI.
        Spinning up a browser, screenshotting the output, and OCR-ing it will use a lot more resources.

        1. The two phases usually work together.
          1: Survey cheaply, quick-categorize urls (often on keywords, and Shieldfont’s choice of words-to-modify actually doesn’t harm this much)

          2: acquire the info (often via automated browsers and OCR, or “classify the information found in this image, then organize it for later use”)

          Phase one needs to be cheap. Phase two is already expensive, and many of the specialized models are very economical even with autobrowsers and OCR.

          Sticking duct tape on a perfectly good bike usually won’t keep it from eventually getting stolen.

      1. Not expensive enough to matter to the main suspects. And it makes the content worthless to a good segment of the audience, who gets caught in the blast radius.

        It’s a philosophical problem, not a technical one, and any solution (in any direction) will always have philosophical grounds for rebuttal.

        So, it’s working as intended, just as the well-funded OCR-based scrapers are working as intended. The dispute will always be that different people have different intent.

  3. It also breaks Ctrl-F for when I don’t immediately see what I’m looking for. Combined with the limited effect it will have on bots (just parse the font once to find what the ligatures are!), this is mostly a demonstration of what outrage about unauthorised scraping makes people do.

  4. Outside of the excessive bandwidth that scrapers can cause, I don’t think it’s right to poison the well.

    If a human can see and learn from something for free, then I don’t see much reason to prevent an AI bot from doing the same.

    I’m less scared of the AI companies scraping than the copyright cartels trying to make it illegal for humans to learn off of their works.

    1. The people doing AI scraping are the same people who will create copyright cartels.

      Their goal internet content as a service. Computer components are being hoarded and prices inflated until eventually the only people that can afford a computer will be the people building data centers. All of the previously free internet will be behind an AI paywall. It’s the same playbook that walmart and amazon use undercut local stores until they go bankrupt and then raise their own prices because they’re the only place you can buy at.

        1. Stack overflow has gone from thousands of posts per day to thousands per month in the same time that vibe coding & LLM chatbots have taken off. I don’t think it’s a coincidence. If LLMs ends up replacing it entirely all of that knowledge will only exist behind paywalls or in a few online public archives with tenuous funding.

          We need to make sure the “freely available content on the internet” remains freely available and on the internet.

          1. It is possible that the best choice is to ensure that freely-available content can easily be stolen. This retains the value it had as information, while ensuring that there is little value to be extracted from building walls around it.

            It’s simply not satisfying, because it feels like rewarding the criminals. But, in reality, it ensures that the criminals can’t get any satisfaction from trying to sell what they stole.

            I think that many in the anti-ai crowd are afraid that AI can do what it’s advertised to do.

      1. But what would be the point of hoarding all the critical components like RAM, flash chips, CPUs, GPUs? If people cannot afford the computer to use your AI datacenters, or play your AAAA games, or watch cat videos in 4K, then all you have is a massive electric space heater burning through your money.

        Current “growth” is not sustainable and is very likely to crash, but I wouldn’t seek deeper conspiracy here. It’s just techbros high on cocke and other stimulants… again.

        1. There are interviews with musk, zuckerberg, and the other ai techbros where they make it abundantly clear they envision a future where everyone goes through their services to access information. They want to be the gate keepers of knowledge.

          1. they envision a future where everyone goes through their services to access information. They want to be the gate keepers of knowledge.

            Yes, and I want a Lamborghini Spark 135 tractor fitted with a bucket loader in the front and a SaMASZ mower in the rear to be my daily driver for commuting to work at Sii office in Gdansk. Everyone can dream.

            First they would have to create some kind of “Internet 2.0” (Meta-net?) controlled entirely by the US. Then they would have to force everyone else to use it.

            Maybe they’d be able to bully Canada or Mexico into joining, but imagine telling privacy conscious EU that all the traffic in the world has to go through some data centre in Kansas first, or it ain’t going. Or better yet, tell the same to China or Russia 😂

            Maybe the US could start threatening everyone else in the world with sanctions if they refuse to join the new, amazing Meta-net, but after two terms of Trumpism, the most likely outcome will be the world gladly telling the US to finally f… off.

      1. Eh, if they think it’s their moral imperative, it is. That’s how moral imperatives are; no consensus is needed (that would be the topic of “ethics” not “morals).

        And, it’s likewise the moral imperative of the other side to gather and protect all of the information of humanity so that it can be “made proper use of”, “not hoarded and wasted”, etc.

        Morals are wonderful things. Everyone has ’em, everyone else’s stink to high heaven, and nobody can agree what to do about ’em.

    2. AI can’t “learn” or even “know” anything.

      Poisoning is just a cute term for introducing OBVIOUS errors in the model output.

      These bots only SEEM good at what they do because so much effort has been put into tricking the users into thinking they are an authoritative source on anything.

      Getting a model to tell you to put lightbulbs in your cookie dough by poisoning it shows the user how dumb the model actually is and breaks the carefully crafted illusion.

      NOT poisoning models is unethical.

      1. It is unethical to intentionally cause harm where none would have occurred. Introducing error causes harm to someone at some point. And, what is an “obvious error” is very much an unsettled question. Obvious to you? obvious to me (which, trust me, is a very different thing)? Are you planning to be ethical and accept the liability arising from the harms you cause? (if not, you’re committing the same crimes your opponents are committing, just with a different set of beneficiaries).

        Ethics is nice to talk about. But, people who find it useful to talk about ethics tend to find the actual details of ethics inconvenient. There are exceptions in all things, but as of yet, you haven’t demonstrated the exception.

    3. Well, stackexchange is pretty much gone. There used to be people that interacted with each other, llm guys snarfed it all up, now we are all isolated, only interacting the the model.
      Shitty thing to happen.

  5. Since 2022 there is AI generated slop polluting the Internet, so much by now that it is getting difficult for the LLM’s not to consume their own excrement. And as we all know that makes them more useless. The good news is that it is becoming harder for LLM companies to find a pure unadulterated source of human intelligence. So the best solution might be to auto generate large volumes of LLM slop from them to consume and explicitly put it in robots.txt

  6. I have the following at the bottom of my index.html page, hoping that I might, someday, win a legal battle with the owners/operators of those unfair AI scrapers (I know; I’m dreaming):

    Copyright
    The information on this site was written by Peter Knoppers and – per the Berne Convention for the Protection of Literary and Artistic Works – is copyrighted by me. Any use related to the development, or training of AI systems without prior, written permission is prohibited. Personal use, indexing for Internet search engines, etc. is intended, permitted and encouraged. Any reproduction of the documents on this site should be clearly marked as copied from this site.

  7. And doing this, doesn’t that just increase the massive hunger of the AI monoliths for even more power and memory as they throw even more $$ at the ‘problem’ of people preventing easy access?. Very soon it will only be the likes of Musk and Co that can afford a memory upgrade for my laptop.

    Yes, we need to stop the blatant IP theft of these things, but with the ignorant believing everything these things spit out, is poisoning the data the best way? One day someone might get hurt as a result.

    Waiting for the AI hype bubble to burst…or the AI monsters to start attacking and eating themselves

  8. Welcome to the days of Leonardo da Vinci.
    Before copyright law people wrote things down in code, added deliberate faults in schematics, to protect their ideas.

    Thanks to closed training data, there are no more IP laws. “I asked the model to make me something that looks/sounds like that.”
    Anything with upbeat with banjo, made with Suno, sounds like Mumford and Sons. Did Suno hire talented banjo players for their training data, no.
    There are videos on Google Tech Talks, showing work, to preserve the “privacy” of the model, to prevent the “attack” of people being able to show that their data was used to train a model.

  9. Why would you ever try to stop scrapers in such an absurd way? There are so many other ways to ban or rate-limit extremely aggressive, unknown visitors which doesn’t make your content less accessible. Most bot traffic I see still doesn’t even feign a “normal” user agent string.

    The idea that if 5-50% of us small-medium website operators find aggressive ways to block or poison scrapers, we’ll noticeably degrade AI, is delusional; AI can curate the data for new models easily enough; it can transform the data; they have $billions of compute and spend a good deal of time and effort on training; at the very least, an OCR step doesn’t matter to them. I repeatedly downscale and upscale my nature macro photos via script, with one of those steps using a GAN, and I dump hundreds of these in per week; it does a wonderful job to remove sensor noise and generally make the photo much easier to process with my human eyes — and it takes basically no time on the laptop I have jammed in a server rack. Alphabet/Apple/Adobe/whatever has so much compute and so many humans at its disposal that it’s pointless to try tricking them at scale. This isn’t the kind of thing where there’s value in bothering to opt out, anyway; if there’s 5PB of data, it doesn’t matter if 2.5PB of that opts out because the odds of non-redundant data between those is infinitesimally small — unless you are doing original scientific research you refuse to publish normally.

    Finally, Google’s shown even if we did successfully sacrifice accessibility and our time/effort to harm AI training, they can literally just not update the training data. Google flagship models are all still stuck in a world over a year and a half old. It’s become difficult to use in high-speed environments like Rust where dependencies’ syntax and feature set radically changes every year, BUT if you’re doing ~everything first-party, it’s still, of course, very useful, and Google still has plenty of users.

  10. If the problem with AI scrapers is bandwidth, why not just throttle them? Or throttle everyone? I wouldn’t mind waiting a few seconds for a page to load (did it every day in the dial-up days), but an AI would probably find this intolerable if it were trying to catalog the entire site. I feel like this is already a solved issue.

    1. “But if it were a solved issue, how would we get any attention?”

      There’s a lot less attention to be had for pointing out an existing solution than for creating something that you can claim credit for yourself. Doesn’t matter if it’s a better solution (or even a solution at all). What matters is that you can claim credit for it.

      And certain ideas stick better than others, independent of whether they’re good ideas. if you can somehow figure out how to make something sticky and that you can claim credit for, this is ideal! Very much doesn’t matter if it’s fit-to-purpose, it could take years for everyone to actually come to a consensus on that point, and in the mean time, you’ve saved a lot of rent while staying in people’s heads.

      The real problem in the world is that there are truly too few novel problems awaiting solution, and solving a new/unique problem is too annoying and hard. So, all this effort and attention is wasted chasing a MUCH more tractable goal: getting attention (whether merited or not).

      Sadly, the result of that is exactly the internet we now live in…

      1. I think a lot of this stems from the concept of a “work ethic”, where just doing something is seen as virtuous in and of itself, regardless of its actual merits. Our current culture has a hard time dealing with solved problems, or the fact that (thanks to technology) we need to do a lot less actual work to maintain our standard of living. Pointing out that a problem is already solved tends to look like “laziness” to hard-working strivers who seem eager to fritter away their time on earth doing something, rather than simply enjoying the results of what has already been done.

    2. I think you’d find it much much harder to effectively throttle as protection, especially if your site is large enough to have the distributed cloudflare style stuff.

      Even if you can set a data rate out of your server per client to something lowish that is going to be rather inconsistently effective, as not every page is small, where the text that is what the scrapers are mostly interested in is small and they don’t actually have to grab the whole page – you’d only really hurt the real users or throttle so little many of the AI scraper won’t even notice they have been slowed.

Leave a Reply

Please be kind and respectful to help make the comments section excellent. (Comment Policy)

This site uses Akismet to reduce spam. Learn how your comment data is processed.