GEO and AEO promise that if you rewrite your content for language models, AI systems will mention you the way SEO once put you on page one. A study by Fenil Suchak, cofounder and CEO of OpenFunnel (YC F24), suggests that playbook misses the mark for at least one critical buyer type: the AI agent running vendor research on someone’s behalf.
What the Study Found
Suchak’s team ran 200 controlled buyer sessions through Openbenchmarks For Agents, a public independent benchmark hub, to observe how frontier reasoning models pick a vendor for a specific technical purchase, in this case a lookalike-audience API.
The numbers are hard to ignore:
- When agents retrieved the independent benchmark, they incorporated it into their final decision in 44 of 45 sessions, a 98% usage rate.
- The vendor leading the independent benchmark was chosen up to 86% of the time on specific, in-market queries, despite having almost no SEO authority behind it.
- Vendor pages got fetched too, but agents frequently set them aside as claims they could not verify.
- Backlink authority was a poor predictor of which vendor got recommended.
- Stronger reasoning models pulled the benchmark into their decisions far more than weaker ones.
Why Agents Behave This Way
Suchak notes he cannot see inside the models, but the session transcripts point in a consistent direction. A reasoning model given a purchasing task appears to look for comparable, structured, third-party numbers it can cite back to whoever assigned the task. A vendor page claiming “industry-leading match rates” gives the agent nothing to work with. A table showing match rates across eight vendors under identical test conditions is directly usable, and in these sessions, agents used it.
The precedent exists in human buying too. Database selection shifted toward public performance comparisons like TPC. Enterprise software purchases lean on analyst evaluations. The difference with agents is speed: the verification step a human might skip on a small purchase, the agent ran on every session observed.
The Operator Takeaway
Suchak is careful to frame this as an early signal from one category and one purchase type, not a settled rule. But if the pattern holds, here is what he recommends:
- Get represented in independent benchmarks in your category and make sure the data on you is accurate and recent. When agents cannot find structured comparisons, they piece together whatever comparable data exists, and you have no say in what that is.
- Publish numbers someone could check. Latency figures, coverage rates, accuracy on a named methodology. Verifiable specifics got used in these sessions. Superlatives got skipped.
- Start measuring agent traffic on your own properties. Most analytics setups still treat AI crawlers and agent sessions as noise. Knowing what agents fetch from your domain is the only way to test any of this against your own funnel.
- Treat measured performance as a growth input, not just an engineering metric. Whether improving a benchmarked number beats another quarter of content production will vary by company, but it is now a comparison worth running.
The narrower claim from the study: in 200 sessions across one B2B category, agents consistently preferred independent measurements over vendor-written pages, and the preference got stronger as buyer intent got more specific and the model got more capable. If agents keep taking on more of the research step in B2B buying, the highest-leverage page about your company may be one you did not write.
