Product
Protection Packs
A Protection Pack is a curated, versioned rule set that PubTrust maintains: keywords, domains, URL patterns and a small number of regexes, each carrying a weight and a written justification. 16 packs, 14 languages, 12,034 rules. You enable the ones you want, per site, and you can edit any of them into a policy of your own.
How does a Protection Pack decide an ad is a violation?
Matched weights are summed, and the policy fires at 1.0. Every rule carries a float weight reflecting how conclusive it is in its own language. Terms with essentially no innocent use sit at 1.0–1.4 and can fire alone — German freispiele (free spins), Swedish casino utan spelpaus. Strong terms that want corroboration sit at 0.5–0.9. Weak signals that are only meaningful in combination sit at 0.2–0.4. Ordinary words that exist purely as corroboration sit at or below 0.2 and can never matter on their own. So one term rarely blocks anything; a creative that stacks several does.
The calibration principle
Over-blocking is this product's primary failure mode, and the corpus is built around that. These rules run over creative text on pages that are, by definition, news pages. A publisher writes about elections, war, alcohol, drugs, addiction, crime and gambling regulation every single day. Several packs — Weapons, Political, Alcohol, Shock — are made almost entirely of words that are also front-page editorial vocabulary. In those packs the category noun is nearly worthless as a signal and the commercial construction carries everything:
gun 0.1 buy ammo online 1.2
wine 0.2 wine club subscription 0.9
bitcoin 0.3 guaranteed 300% return in 24 hours 1.3The best signals are the ones advertisers are legally required to print
Mandated regulatory warnings appear in the advertising and essentially nowhere else. German Glücksspiel kann süchtig machen. French jouer comporte des risques and the 09 74 75 13 13 helpline. Czech Ministerstvo financí varuje. English begambleaware.org. Polish graj odpowiedzialnie. Danish StopSpillet. These carry the highest weights in the corpus and are the first candidates to be inlined into the tag for pre-bundle blocking. It is a pleasing inversion: the compliance apparatus that legitimate operators must carry is what makes their advertising the easiest to identify.
The 16 packs
Core
Gambling
Casino, sportsbook, lottery, bingo, poker and the affiliate industry around them. The largest pack in the corpus and the most regulation-sensitive, because "gambling" means something different in every one of our 14 markets. Licensed national operators are weighted low on purpose so the pack cannot block them on its own — see Global Coverage.
Adult
Sexually explicit and adult-services creatives on mainstream inventory. Calibrated hard against a specific failure: news sites publish health, court and relationship coverage daily. Anatomy terms sit at or below 0.2 — blocking breast-cancer coverage is the mistake that would embarrass this product. Identity terms (gay, lesbian, trans, queer, bi) are absent by design; they are never an adult signal.
Crypto & Forex Scam
Fake celebrity-endorsed trading robots, boiler rooms, pig-butchering funnels, giveaway and airdrop drainers, seed-phrase phishing. The asset is not the signal, the promise is: bitcoin, crypto, wallet, broker are capped at 0.2–0.35 because they are both ordinary editorial vocabulary and the vocabulary of lawful licensed advertisers. Guaranteed returns, implausible figures and secret loopholes carry the weight.
Malware & Scam
Scareware, tech-support scams, phishing lures, prize funnels, PUP installers. Bare virus, malware, antivirus, phishing, hacker are held at 0.1–0.2 — eight of them can co-occur in one legitimate security article, and security vendors advertise on these pages constantly. The second-person alert construction is what earns 0.9 and above.
Regulated
Alcohol
Retail constructions rather than category nouns. port, scotch, malt, cider, ale, stout, bitter, proof, shot, round and pint are deliberately absent as bare terms. So is all quit-drinking, alcoholism and recovery vocabulary: blocking a treatment charity's public-health ad would be a product failure, not a success.
Tobacco & Vape
Because tobacco advertising is banned or near-banned across the EU and UK, live creatives are overwhelmingly vape, e-liquid, nicotine pouches and cross-border grey market — so the weight sits on vape retail vocabulary. Smoking-cessation language is excluded entirely.
Pharma
Rogue online pharmacies: no-prescription pill shops, counterfeit ED and opioid sales, steroid and SARM sellers, unlicensed peptide shops. The transactional phrase is the signal — nobody writes "buy viagra online" editorially. GLP-1 brand names are held at 0.2–0.25 because they have been one of the largest news subjects of the last two years.
CBD & Cannabis
CBD oil and gummies, hemp flower, delta-8/HHC/THCA, headshops, seed banks. Legalisation is front-page politics in several markets and a licensed dispensary industry advertises lawfully in much of the US, so bare category nouns stay at 0.2–0.35. pot is a cooking pot, weed is a garden weed, and joint is omitted entirely.
Weapons
Retail advertising of firearms, ammunition and restricted knives. The catastrophic failure mode here is obvious: a news site reports on war, terrorism, crime and defence procurement every day. gun, rifle, weapon, ammunition, missile sit at or below 0.2 and cannot fire alone. Hunting and sport shooting are lawful, widely advertised activities and are treated as a category marker you may choose to allow, not a violation.
Political
Viewpoint-neutral by design, and this is a hard constraint. The pack contains no party names, no politician names, no movement or ideology labels and no issue vocabulary in any language. It detects the mechanics of paid electoral advertising — statutory disclaimer formulas, donation solicitation, the vote imperative — so you can label it, route it for review, or block it during a pre-election blackout under the EU Political Advertising Regulation and the DSA. It is a compliance control, not a censorship tool, and it will stay that way. It also sits lower overall than any other pack, because election coverage is the news.
Deceptive and low-quality
Weight Loss
Miracle fat-burners, keto and ACV gummies, detox teas, and the fake-doctor and fake-news-article funnels that sell them. The implausible promise is the signal, not the topic — commercial slimming programmes, gyms and meal-kit services are lawful advertisers buying against the same pages.
Clickbait & Chumbox
Content-recommendation creatives: disbelief constructions, outrage bait, celebrity-death-hoax teasers. Almost every rule is a multi-word phrase, because this pack runs over text on a publisher's own site and must never fire on the publisher's own headlines. Listicle markers stay low deliberately — real newsrooms and real advertisers both use them all day.
Fake Endorsement
Creatives falsely claiming a celebrity, broadcaster, newspaper or regulator endorsed a product, usually fused to a crypto or diet offer. We deliberately do not enumerate celebrity names: they go stale in weeks, they risk defaming real people, and the endorsement construction catches the format regardless of who is named. Broadcaster and newspaper names are capped at 0.3 — they are legitimate advertisers and constant editorial subjects.
AI-Slop Creative
Mass-produced generative creative that premium inventory does not want. The pack detects artefacts of machine generation, not the subject "AI". artificial intelligence, machine learning, chatgpt, neural are pinned at 0.15–0.25 and can never reach threshold alone; the bare two-letter token ai is omitted entirely, because it is a substring hazard in English (said, maintain, claim, plaid) and an ordinary word or particle in several of our other languages.
Shock & Gore
The disgust-marketing genre — toenail fungus, skin parasites, "look what came out of her ear" — plus gore-bait. blood, death, victim, crash, cancer, surgery, outbreak and their relatives are omitted entirely rather than weighted low: a news site prints them hourly. Self-harm and suicide vocabulary is excluded on purpose; a false positive there would be actively harmful.
Fake UI
Creatives imitating operating-system, browser or page chrome to steal a click: fake download buttons, fake close buttons, fake system dialogs, cursor arrows painted into the creative, fake progress bars. Bare ok, next, close, play, submit and download are omitted entirely — they are the legitimate call-to-action on real ads and on your own page furniture. The signal lives in the full sentences only a fake UI writes. This is the keyword half of a two-part defence; Runtime Integrity catches the behavioural half.
Why is Runtime Integrity not one of the packs?
Because it has no keywords and no opt-out. Forced navigation, pop-unders, document.write hijacking, autoplay with sound and CPU-hogging creatives are attacks on the reader and on the page, not categories of advertising, and no publisher has ever wanted them. So Runtime Integrity is an always-on policy on every site rather than a pack you enable. The one genuinely optional item in that group — blocking specific competitor domains — ships instead as a Blocklist policy you author yourself.
Which taxonomy does PubTrust report in?
Every violation carries an IAB Ad Product Taxonomy 2.0 node (cattax = 8) alongside our own pack key, so you can reconcile PubTrust's data with your SSP blocklists and your own reporting. Ad Product Taxonomy is the right primary mapping because it classifies ads, not pages, and it has direct nodes for nearly everything we detect: 1361 Gambling with children for Casinos, Lottery and Sports Betting, 1448 Non-Fiat Currency → Cryptocurrency Exchanges, 1001 Adult Products and Services, 1291 Dieting and Weightloss, 1544 Tobacco, 1049 Cannabis, 1474 Politics, 1002 Alcohol. We also carry IAB Content Taxonomy 3.1 (cattax = 9) for page context, including its Sensitive Topics branch.
A note on GARM, because you will be asked
PubTrust does not claim GARM compliance and no honest vendor should. GARM was discontinued in August 2024, and following the WFA/X settlement in July 2026 the WFA stated it will not form or restart GARM or a similar initiative. Its eleven Brand Safety Floor categories survive as the Sensitive Topics branch of IAB Content Taxonomy 3.1, which is versioned, machine-readable and actively maintained. That is the citation we use. Conformance to a dissolved body is a claim that fails the first time somebody checks it.
The parsing trap we handle
cattax defaults to 1 when absent, and taxonomy 1 is deprecated for lacking Sensitive Category Designation flags. A bare cat: [...] with no cattax is legally Content Taxonomy 1.0, not the current version. Our parser handles 1, 2, 6, 7, 8, 9 and the vendor range rather than assuming. It is a small thing that silently corrupts category data in a lot of pipelines.
Where IAB does not reach
Forex and binary options, chumbox and tabloid recommendation, and fake-news clickbait have no IAB Ad Product node. Those map to Google's public AdX sensitive-category dictionary instead — 10 Get Rich Quick, 37 Sensationalism, 30 Black Magic/Astrology/Esoteric — which is small, flat and derived from exactly the adversarial creative population we are classifying.