A number beside a domain can make a classification look more settled than its evidence warrants. Before using an AI score to block a request, ask what the number predicts, which observations produced it, and what happens when it is wrong. This article proposes a practical evaluation workflow for teams building or assessing ad domain classifiers. It is an educational design discussion: AdBlockList.com does not publish verified scores for live domains or claim that a particular model can identify every advertisement.
Define the prediction in a sentence
Start with a label a reviewer can apply consistently. For example, your proposed task might be to identify hostnames observed delivering advertising resources during a specified observation period. That definition is narrower than identifying every company involved in advertising, every unwanted request, or every privacy concern.
Write a companion exclusion rule. Decide how to handle mixed purpose hosts, measurement services, payment infrastructure, and inactive domains. A host that sometimes serves an advertisement may also deliver content a user requested. Your model's label and your blocking policy need to preserve that distinction.
Record the unit being classified. A hostname, a registrable domain, a complete URL, and an individual request context are different units. If training labels describe individual requests but enforcement blocks entire domains, a strong request classifier can still produce unsuitable rules. The domain block list guide explains the scope decisions that need to accompany a classification.
Build evidence records before scoring
For each reviewed example, capture the observation date, the context in which the resource appeared, and a concise account of the relevant behavior. Distinguish direct observations from another list's classification. Two imported lists may ultimately rely on the same original report, so agreement alone is weak evidence of independent confirmation.
Keep an explicit uncertain category during annotation. Reviewers should be able to say that a request's purpose could not be established, or that the host serves several purposes. Forcing every ambiguous example into a binary label hides disagreement inside the training set.
Choose a short annotation guide with examples at the boundaries. Have reviewers independently classify a shared sample, then discuss disagreements before expanding the dataset. Preserve the original labels as well as the resolved decision. That history can expose a vague definition or a recurring observation gap more usefully than repeatedly adjusting the model.
Give the number a documented meaning
A model output on a zero to one scale is not automatically a reliable probability. If you present scores as probabilities, evaluate whether examples assigned similar values have corresponding observed outcomes in suitable held out data. State the evaluation population and time period beside that claim. A decimal alone does not establish what confidence means.
The NIST AI Risk Management Framework 1.0 discusses validity, reliability, measurement limitations, and the importance of the context in which an AI system is used. Applying those ideas here, a useful scorecard should pair classification results with evidence age, known gaps, and the consequences of the proposed action. The workflow in this article contains our suggested implementation choices.
When calibration has not been established, label the value as a model score and explain its intended use, such as ranking items for review.
Test on examples the model has not effectively seen
Set aside evaluation data before choosing a threshold. Keep related examples together where they could otherwise reveal the answer across splits. For instance, similar subdomains associated with one service can make a random split look easier than discovering unfamiliar services. Document the grouping method so another reviewer can reproduce it.
Add a later observation period to test whether the classifier still works after the original collection window. Keep the earlier benchmark for comparison. Investigate changes by category, language, hosting pattern, and evidence source when those groups are relevant and sufficiently represented. An overall score can hide a weakness in the precise group your deployment encounters most often.
Compare the AI system with a modest baseline, such as the existing reviewed list or a simple ruleset. Evaluate the complete operational workflow, including manual review time. A more elaborate model needs to contribute something measurable to your intended decision, whether that is finding overlooked candidates or reducing review effort.
Test missing and untrusted inputs
Include examples with incomplete observations, unavailable pages, and conflicting metadata. Decide whether the system should decline to score them, flag a limitation, or send them to a separate queue. Missing evidence should remain visible to whoever acts on the output.
If a language model reads page text during classification, treat that text as untrusted input. Build evaluation examples in which page content tries to instruct the model to change its classification or reveal unrelated information. Keep the classifier's task and allowed outputs constrained, and validate the result before it becomes a rule. The resulting test record should show both the adversarial input and the expected safe behavior, making future model changes easier to review.
Measure errors in units people understand
For the chosen threshold, count correctly identified ad domains, incorrect ad classifications, and missed ad domains in your labeled sample. Precision asks what share of predicted positives are actually positive. Recall asks what share of labeled positives were found. Report the counts alongside these fractions so a small sample cannot masquerade as decisive evidence.
Also test the action produced by the classification. A false positive that breaks an essential login deserves different attention from one that removes an optional decorative resource. Maintain a small collection of representative user journeys and record whether each still works with proposed rules active.
Review a sample of negative predictions too. A queue containing only high scores tells you little about what the model misses. Keep uncertain judgments visible in the evaluation report and explain whether they were excluded, separately counted, or resolved through further observation. That choice affects what the reported measurements mean.
Choose thresholds for specific actions
Use separate action bands instead of making one cutoff do every job. A broad candidate threshold can feed a review queue. A stricter publication threshold can require both stronger model evidence and a completed review. A mixed purpose classification can route to a request level investigation even when its score is high.
| Situation | Suggested next action | Evidence to retain |
|---|---|---|
| Strong evidence, narrow target | Review for a scoped rule | Observation and test result |
| Conflicting observations | Investigate before publishing | Both sides of the disagreement |
| Little recent evidence | Queue a fresh observation | Last reliable observation date |
| Essential functionality involved | Test a narrower intervention | User journey and recovery steps |
Select actual numeric cutoffs from your validation results and operating costs. Estimate how many candidates the review team can inspect and how costly an incorrect automatic block would be. Record the tradeoff when choosing a threshold. A copied cutoff from an unrelated model has no dependable meaning for your labels, score scale, or traffic mix.
After a limited rollout, compare the predicted review volume with the actual queue and inspect reviewer overrides. If the queue grows faster than the team can resolve it, revise the operating plan explicitly. Quietly lowering review standards changes the meaning of the published list even when the model itself stays unchanged.
Make decisions explainable and revisitable
Show reviewers the observations that support a recommendation, their dates, and any conflicting evidence. Separate a generated explanation from the underlying record. Fluent text should never substitute for an observation the reviewer can inspect.
Version the model, the label policy, and the published rule decision. A score can change because the model changed, because new evidence arrived, or because the policy was revised. Those are different events and deserve different explanations. The blocklist API design article covers how to carry these identities through releases.
Provide a correction path for mistaken classifications. Use confirmed reports to improve examples and policy, while retaining an independent evaluation set. Repeatedly tuning against every reported test case can create a convincing history without proving performance on new cases.
Use AI to support a clear decision
A useful scoring system makes uncertainty actionable. It tells a reviewer where evidence is strong, where more observation would help, and when the proposed block reaches beyond the classification's scope. Begin with a documented label, evaluate independent examples, and connect each threshold to an explicit action. The AI score planning guide offers a place to organize those decisions before a ranking number becomes part of a live enforcement policy.



