I have been reading about ChatGPT Ads and I keep coming back to one problem. We have years of benchmarks for Google and Meta, but very little for a surface where the user is in the middle of a conversation, not scrolling a feed or typing a search query.
My working hypothesis is that the intent signal may be strong but the path to conversion is harder to see. A click might be the end of a long research chat, so a last-click view could undervalue the channel or credit it for something it only nudged.
So the real debate is about setup, not just performance. Do you start with a small fixed test budget and a written stop rule, such as a set cost per qualified lead or a set number of conversions by a date? Or do you wait until you have a proven baseline on another channel and then compare? Waiting feels safer, but you may never get clean data if you only look at the channel after everyone else has already decided.
I would like to hear how others define a kill threshold for a channel where they do not yet trust the attribution. What did you measure first, and what did you ignore?