The same questions, on the same engines, every week, so the numbers stay comparable and the plan stays honest.
Every question runs 13 ways, every week. Assistants are measured on both halves: what they answer from memory and what they answer after searching. Three apps are measured in a live session, exactly as a customer sees them.
Every answer is read: who is named, in what position, with what sentiment, from which sources, and who actually got picked. What changed since last week shows up first.
Every action carries severity, effort and horizon, and links back to the question it came from.