The Missing Science in Generative Engine Optimization

Most GEO platforms can tell you whether AI mentions your brand. Four gaps still stand between measuring visibility and actually improving it — and the real bottleneck is the speed of experimentation.

Oliver Bennett from RevAI Team

Kartik Hosanagar

Given the rapid rise in the use of AI search, Generative Engine Optimization (GEO) is an emerging focus for digital marketers. It refers to the science and art of informing and influencing AI engines about your product. Most GEO measurement platforms follow the same playbook. They do three things:

  • estimate the prompts customers are likely to ask AI systems

  • measure whether your brand appears in AI answers

  • identify the sources and citations those answers rely on

These are useful capabilities. But they leave out several things that matter if your goal is to improve visibility rather than simply measure it. Here are four gaps that still exist.

1. Nobody has the full picture of what users ask AI

Every GEO platform begins with prompts.

The problem is that only the AI providers themselves know the complete set of prompts people actually type into ChatGPT, Claude, Gemini, Perplexity, and other systems. Google Search has long published keyword data through products like Keyword Planner. Advertisers know roughly how often people search for "running shoes" or "CRM for startups." Nothing comparable exists for AI.

Most GEO platforms therefore estimate prompts using indirect signals. Some rely on opt-in consumer panels. Others extrapolate from Google search volume. Others use their own heuristics. Each approach is reasonable. None is complete and free from biases and errors. As a result, there is simple no established methodology for prompt generation today. 

2. AI answers are distributions, not facts

Most GEO dashboards ask a question once, record the answer, and treat that answer as truth.

Large language models do not work that way. First, language models are probabilistic systems. Ask ChatGPT "What's the best CRM for a startup?" in three parallel tabs and you will receive three different answers because of this probabilistic nature.

Second, AI systems personalize responses using prior conversation history.

Third, users rarely ask identical questions. One founder asks for the “best CRM for a seed-stage company.” Another asks for a 100-person startup. Tiny changes in wording often produce different recommendations.

This creates an important statistical problem. If your dashboard says your visibility increased from 36 percent to 42 percent, did your optimization actually work? Or are you simply seeing statistical variation?

3. Attribution largely disappears inside AI conversations

Marketing teams are built around attribution. AI makes attribution much harder.

Imagine someone asks ChatGPT: "What's the best executive program for CTOs?"

The model recommends Wharton's CTO program. The user spends fifteen minutes asking follow-up questions about tuition, curriculum, faculty, and admissions. Only after finishing that conversation do they visit Wharton's website and enroll.


A ChatGPT conversation recommending the Wharton, Cambridge Judge and UC Berkeley CTO programs in response to a question about the best leadership program for CTOs

Inside Wharton's analytics, that visit may appear as direct traffic or perhaps a branded Google search. The AI conversation that shaped the decision is invisible.

As AI increasingly becomes the research layer before website visits, traditional attribution systems lose visibility into a growing portion of the customer journey.

4. Measurement without experimentation does not create learning

This is the biggest gap.

Suppose your dashboard shows weak visibility. Now what? Your team studies the data. Someone proposes writing a new blog post. Someone else suggests restructuring product pages. Someone suggests editing YouTube video descriptions. Alt-Text for images, and so on. 

Marketing drafts new content. Brand reviews it. Legal reviews it. Marketing publishes it through the CMS. You wait several weeks for search and AI systems to absorb the changes. Finally you measure again. If the change fails, you repeat the process. Each learning cycle can easily take six to eight weeks. After several iterations, you may finally discover something that works. But by then, the underlying AI models may already have changed and you are back to the drawing board.

The bottleneck is the speed of experimentation.

From dashboards to flight simulators

Instead of testing every idea in production, marketers should first test ideas inside a simulator.

The workflow looks very different. First, measure what AI systems are doing. Next, have AI agents analyze the gaps and generate hundreds of possible interventions. Instead of deploying every one of those ideas, evaluate them inside an LLM simulator that has been trained on your domain. Most ideas will fail in simulation. A handful will consistently outperform the rest. Those become the strategies you ship into production.

The simulator does not replace real-world testing. It dramatically reduces the number of expensive (and failed) experiments you have to run. Instead of learning every six weeks, you can eliminate weak ideas in hours and reserve production for the most promising candidates.

The result is faster learning, lower cost, and a growing advantage over competitors.

Where GEO is heading

The next generation of GEO will focus on experimentation. The winners will be the companies that learn faster than everyone else.

At Bodhium Labs, that is the problem we are building for. We are developing LLM simulators that let marketers test optimization strategies before they publish them. By filtering out weak ideas before they reach production, teams can spend less time waiting, learn more from every experiment, and compound their advantage over time.

About the author

Oliver Bennett from RevAI Team

Kartik Hosanagar

Co-founder

Wharton Professor, Prev. Founder at Yodle (acq. by Web.com), Jumpcut (acq)