Two AI engines answering the same B2B buying question surfaced the same tracked brands 81 percent of the time on one platform and 43 percent on the other, according to measurement data Treyci released on September 3. The gap is roughly 2×, and it lands squarely inside the shortlist stage of a purchase.

Treyci’s analysis covered more than 1,200 scored answers, drawing on roughly 100 buying-intent prompts per category across ChatGPT, Perplexity, Gemini and Grok, with every prompt repeated three times per engine each month. The company didn’t disclose which two engines produced the 81–43 split, which category, or the prompt set. Founder Keith Schilling, previously an AEO/GEO practitioner at PayPal, framed the discipline problem bluntly: “Marketing teams are making decisions about AI search based on one screenshot of one answer from one engine.”

The absence problem is what makes this different from a normal channel-measurement story. A brand missing from an AI answer produces no impression, no session, and no analytics trace of the buying conversation. It’s a shortlist decision no one at the vendor ever sees.

Two adjacent data points sharpen the picture. Treyci’s scan of 100 B2B SaaS companies found only 41 had published an llms.txt file for AI crawlers. And Profound data reported by MarketingProfs, drawn from 1,724 prompts, showed Claude used web search in 93 percent of responses versus 13 percent for Claude Code, with brand overlap of roughly 20 percent between the two.

The Brainlabs study of AI-referred traffic quality measures buyers who arrived; Treyci measures buyers who never did. Google’s new per-site AI visibility reports illuminate one engine. The rest of the shortlist remains dark to its own vendors.

Sources