Developer fieldnotes

AI scores for ad domains: evidence, uncertainty, and useful thresholds

An AI score becomes useful when its target, evidence, and uncertainty are visible. Learn how to evaluate domain classifications and choose thresholds that reflect the cost of mistakes.

AI Scores Explained: neon typography and a glowing scoring motif, framed and branded AdBlockList.com.

Ad Block List Lab / Fieldnotes

A number beside a domain can make a classification look more settled than its evidence warrants. Before using an AI score to block a request, ask what the number predicts, which observations produced it, and what happens when it is wrong. This article proposes a practical evaluation workflow for teams building or assessing ad domain classifiers. It is an educational design discussion: AdBlockList.com does not publish verified scores for live domains or claim that a particular model can identify every advertisement.

Define the prediction in a sentence

Start with a label a reviewer can apply consistently. For example, your proposed task might be to identify hostnames observed delivering advertising resources during a specified observation period. That definition is narrower than identifying every company involved in advertising, every unwanted request, or every privacy concern.

Write a companion exclusion rule. Decide how to handle mixed purpose hosts, measurement services, payment infrastructure, and inactive domains. A host that sometimes serves an advertisement may also deliver content a user requested. Your model's label and your blocking policy need to preserve that distinction.

Record the unit being classified. A hostname, a registrable domain, a complete URL, and an individual request context are different units. If training labels describe individual requests but enforcement blocks entire domains, a strong request classifier can still produce unsuitable rules. The domain block list guide explains the scope decisions that need to accompany a classification.

Build evidence records before scoring

For each reviewed example, capture the observation date, the context in which the resource appeared, and a concise account of the relevant behavior. Distinguish direct observations from another list's classification. Two imported lists may ultimately rely on the same original report, so agreement alone is weak evidence of independent confirmation.

Keep an explicit uncertain category during annotation. Reviewers should be able to say that a request's purpose could not be established, or that the host serves several purposes. Forcing every ambiguous example into a binary label hides disagreement inside the training set.

Choose a short annotation guide with examples at the boundaries. Have reviewers independently classify a shared sample, then discuss disagreements before expanding the dataset. Preserve the original labels as well as the resolved decision. That history can expose a vague definition or a recurring observation gap more usefully than repeatedly adjusting the model.

Give the number a documented meaning

A model output on a zero to one scale is not automatically a reliable probability. If you present scores as probabilities, evaluate whether examples assigned similar values have corresponding observed outcomes in suitable held out data. State the evaluation population and time period beside that claim. A decimal alone does not establish what confidence means.

The NIST AI Risk Management Framework 1.0 discusses validity, reliability, measurement limitations, and the importance of the context in which an AI system is used. Applying those ideas here, a useful scorecard should pair classification results with evidence age, known gaps, and the consequences of the proposed action. The workflow in this article contains our suggested implementation choices.

When calibration has not been established, label the value as a model score and explain its intended use, such as ranking items for review.

Test on examples the model has not effectively seen

Set aside evaluation data before choosing a threshold. Keep related examples together where they could otherwise reveal the answer across splits. For instance, similar subdomains associated with one service can make a random split look easier than discovering unfamiliar services. Document the grouping method so another reviewer can reproduce it.

Add a later observation period to test whether the classifier still works after the original collection window. Keep the earlier benchmark for comparison. Investigate changes by category, language, hosting pattern, and evidence source when those groups are relevant and sufficiently represented. An overall score can hide a weakness in the precise group your deployment encounters most often.

Compare the AI system with a modest baseline, such as the existing reviewed list or a simple ruleset. Evaluate the complete operational workflow, including manual review time. A more elaborate model needs to contribute something measurable to your intended decision, whether that is finding overlooked candidates or reducing review effort.

Test missing and untrusted inputs

Include examples with incomplete observations, unavailable pages, and conflicting metadata. Decide whether the system should decline to score them, flag a limitation, or send them to a separate queue. Missing evidence should remain visible to whoever acts on the output.

If a language model reads page text during classification, treat that text as untrusted input. Build evaluation examples in which page content tries to instruct the model to change its classification or reveal unrelated information. Keep the classifier's task and allowed outputs constrained, and validate the result before it becomes a rule. The resulting test record should show both the adversarial input and the expected safe behavior, making future model changes easier to review.

Measure errors in units people understand

For the chosen threshold, count correctly identified ad domains, incorrect ad classifications, and missed ad domains in your labeled sample. Precision asks what share of predicted positives are actually positive. Recall asks what share of labeled positives were found. Report the counts alongside these fractions so a small sample cannot masquerade as decisive evidence.

Also test the action produced by the classification. A false positive that breaks an essential login deserves different attention from one that removes an optional decorative resource. Maintain a small collection of representative user journeys and record whether each still works with proposed rules active.

Review a sample of negative predictions too. A queue containing only high scores tells you little about what the model misses. Keep uncertain judgments visible in the evaluation report and explain whether they were excluded, separately counted, or resolved through further observation. That choice affects what the reported measurements mean.

Choose thresholds for specific actions

Use separate action bands instead of making one cutoff do every job. A broad candidate threshold can feed a review queue. A stricter publication threshold can require both stronger model evidence and a completed review. A mixed purpose classification can route to a request level investigation even when its score is high.

SituationSuggested next actionEvidence to retain
Strong evidence, narrow targetReview for a scoped ruleObservation and test result
Conflicting observationsInvestigate before publishingBoth sides of the disagreement
Little recent evidenceQueue a fresh observationLast reliable observation date
Essential functionality involvedTest a narrower interventionUser journey and recovery steps

Select actual numeric cutoffs from your validation results and operating costs. Estimate how many candidates the review team can inspect and how costly an incorrect automatic block would be. Record the tradeoff when choosing a threshold. A copied cutoff from an unrelated model has no dependable meaning for your labels, score scale, or traffic mix.

After a limited rollout, compare the predicted review volume with the actual queue and inspect reviewer overrides. If the queue grows faster than the team can resolve it, revise the operating plan explicitly. Quietly lowering review standards changes the meaning of the published list even when the model itself stays unchanged.

Make decisions explainable and revisitable

Show reviewers the observations that support a recommendation, their dates, and any conflicting evidence. Separate a generated explanation from the underlying record. Fluent text should never substitute for an observation the reviewer can inspect.

Version the model, the label policy, and the published rule decision. A score can change because the model changed, because new evidence arrived, or because the policy was revised. Those are different events and deserve different explanations. The blocklist API design article covers how to carry these identities through releases.

Provide a correction path for mistaken classifications. Use confirmed reports to improve examples and policy, while retaining an independent evaluation set. Repeatedly tuning against every reported test case can create a convincing history without proving performance on new cases.

Use AI to support a clear decision

A useful scoring system makes uncertainty actionable. It tells a reviewer where evidence is strong, where more observation would help, and when the proposed block reaches beyond the classification's scope. Begin with a documented label, evaluate independent examples, and connect each threshold to an explicit action. The AI score planning guide offers a place to organize those decisions before a ranking number becomes part of a live enforcement policy.

Questions or a factual correction? Contact Ad Block List Lab.

Keep reading

Follow the next question.

Back to the Lab