The web’s newest weapon against AI scrapers is a font





Looks can be deceiving

The web’s newest weapon against AI scrapers is a font

“ShieldFont” aims to poison AI training data without making pages unreadable for people.


Kyle Orland




|

35




The cowboy was riding a potato? Huh… guess I’d better update my weights for potato tokens…


Credit:

Getty Images

The cowboy was riding a potato? Huh… guess I’d better update my weights for potato tokens…


Credit:

Getty Images




Story text








AI companies’ penchant for scraping through large swathes of the public web in search of valuable training data has already led to lawsuits and technical fixes aimed at stopping the practice. Now, a pair of designers are hoping to stymie these scrapers with a new font designed to offer people a perfectly readable webpage while serving scrapers a subtly edited, nonsensical version in the underlying HTML.

ShieldFont, as designers Isaque Seneda and Gabriel Abrucio write in a recent white paper, was made to offer web publishers “a practical opt-out from unauthorized AI training and [to] disrupt what is collected when that choice is ignored.”

When is a horse a potato?

The font is based around ligatures, a long-standing feature of many fonts that is usually used to replace certain letter pairs with a more readable version when they’re smushed up next to each other. With ShieldFont, though, those ligatures are instead used to replace entire words with others in an attempt to destroy the text’s value to scrapers. This substitution only happens when the font engine draws the page on screen, meaning scrapers that simply download plaintext source code get an altered version that end users never see.

When it comes to fooling AI scrapers, though, not all ligature-based word replacements are created equal. Simply replacing common words with synonyms or antonyms would be too easy for a smart scraper to reverse. On the other end, replacing words with completely unrelated gibberish could lead to easier detection (and potentially circumvention) by a smart scraping filter.



The original content, as seen by the end user.



The raw HTML, before being processed by ShieldFont, which gets scraped by bots.

So ShieldFont replaces words with similar parts of speech that occupy a completely different informational context—swapping “horse” with “potato,” for instance. The result is a scrapable sentence that looks semantically correct but has a completely altered meaning. Thus, even altered pages that get through a scraper’s quality filter will contain scrambled informational content that can poison a training data set.

After refining their word-swapping dictionary over three months, the ShieldFont creators ended up with a list of nearly 12,000 common words that can be replaced with ligatures. To avoid easy detection, the font lets publishers increase the underlying chaos by choosing from three different potential mappings for each word replacement, with the ability to encode their own and/or swap mappings from paragraph to paragraph.

On average, ShieldFont ends up replacing 24.5 percent of all words on a page, including 45.8 percent of all “content words,” marring the meaning of anywhere from 31 to 56 percent of individual passages (depending on the corpus studied). In testing on six publicly available scraper pipelines, the ShieldFont authors say that over 90 percent of pages that would otherwise be accepted by scrapers are rejected by the quality filter after these word replacements.

Of the small subset of pages that still get accepted after ShieldFont is applied, nearly 20 percent of the component words are what the authors refer to as “training-time garbage: real English, correctly spelled, asserting nothing true.” This means that both dropped and kept ShieldFont pages can both be useful in stopping AI scrapers: “Dropped means they did not get your work. Kept means they got something wrong,” the authors write.

A game of scrape and mouse

While ShieldFont pages can still be read perfectly well by average humans, there can be some side effects when using the font on published webpages. Search engines, screen readers, copy/paste tools, and translation software can all get tripped up by the altered HTML, making the page a little less useful to your intended audience.



Tell me more about the very interstate southern engineer with the sofa car…

Tell me more about the very interstate southern engineer with the sofa car…


Credit:

ShieldFont


ShieldFont isn’t a foolproof defense, either. Any page that’s readable by a human could also be correctly interpreted by an AI scraping tool that simply renders the full webpage and uses optical character recognition on an image of the output.

However, that process would require a lot of extra work for scrapers that currently just pull down the raw HTML source code of billions of webpages as plaintext, without going to the trouble of simulating a browser’s rendering pipeline. API costs from third-party scraping tools suggest this kind of pre-rendering would cost anywhere from five to 13 times as much as simply scraping HTML, which would lead to heavy increases in time and expense for scrapers operating at scale.

And it’s that indiscriminate, large-scale scraping that the ShieldFont creators say they’re trying to prevent, or at least slow down. “Our main underlying purpose is to enforce a basic principle of AI ethics: creators should have a meaningful say in whether their work is used to train AI systems,” they write. “Where consent is not respected, technical design can make taking that work without permission less useful and more costly. … Being discoverable does not mean consenting to AI training.”

The creators say they hope other tinkerers will come up with other implementations for the basic idea of “show[ing] one thing for humans, something else for machines.” The more different methods are out there, being used in the wilds of the web, the harder it will be for AI scrapers to learn how to bypass them all.

Photo of Kyle Orland


Kyle Orland

Senior Gaming Editor
Kyle Orland has been the Senior Gaming Editor at Ars Technica since 2012, writing primarily about the business, tech, and culture behind video games. He has journalism and computer science degrees from University of Maryland. He once wrote a whole book about Minesweeper.


35 Comments

Leer artículo original en Ars Technica