Method
How we measure
Last updated 25 September 2026
Where the conversations come from
Advox writes them. Two corpora run side by side: a representative market design, built to reflect the kinds of conversation a category actually produces, and an adversarial set written to find the moments a brand would not want to be in. Each conversation is played more than once, on different days, because a single play measures a moment rather than a tendency.
What counts as a placement
One advertiser's advertisement on one turn of one conversation. Several cards from the same advertiser on one turn count once. A turn where ChatGPT withheld an advertisement is not a placement — it is a separate measurement.
How “unsafe” is decided
A model reads the whole conversation and scores it against a fixed rubric. Cases near the boundary are re-read by a second, stronger model, and a set of hand-labelled conversations sits above both to test whether either is right. Every harm category shown in the product is defined by the question the grader was actually asked, not by a description written afterwards for the interface.
Why every number has a range beside it
A rate without its interval is not a measurement. At the sample sizes this category currently supports, an 8% unsafe share from a few hundred placements could honestly be anywhere from 5% to 12%, and those two numbers would lead to different decisions. Advox reports the interval at the same weight as the estimate, and where two rates are being compared it says plainly when they cannot be told apart.
What this cannot tell you
- It is not a sample of real ChatGPT traffic. Nobody outside OpenAI has one, and any vendor implying otherwise is describing something they do not have.
- Share of voice is a share of the placements Advox observed, not of spend or of impressions. We cannot see either.
- Delivery measurements are taken in our collector's browser, not in every reader's. They are evidence of a problem worth raising with your partner, not a billing reconciliation.