5 criteria to find an AI marketing agency worth the retainer

man writing on whiteboard

Every agency pitch deck says “AI-powered” now. That phrase tells you almost nothing useful. It appears on the site of a team running ChatGPT as a glorified Google Doc replacement, and on the site of a team that has rebuilt its entire creative and reporting pipeline around it. From the outside, the decks look identical.

The question that actually separates those two categories is sharper: where in the decision-making process does AI operate, and what evidence shows it produces better outcomes than the agency got before? That is the version worth asking on every call.

Why “Do you use AI?” is the wrong opening question

Ask any agency in 2026 whether they use AI and the answer is yes. That question stopped being a real filter about the same time “do you have a website” did. A meaningful share of marketing teams now have AI running somewhere in their workflow, which means the yes/no version filters out almost nobody.

The productive version is narrower: which specific parts of the workflow does AI touch, and where does a human step in before anything reaches you? An agency that can answer that tool by tool is telling you something real. An agency that answers with “we’re AI-first” and nothing more specific is letting the phrase do work that substance should be doing.

“The productive question is never ‘do you use AI?’ but ‘where does AI actually operate, and what evidence shows it produces better outcomes?'”
3D rendered ai text on dark digital background

Criterion 1: Tool and workflow transparency

A real answer names the actual tools and models in use. It separates what is automated from what a person still reviews, explains how client data is handled inside that pipeline, and breaks out the strategy fee from pass-through platform and API costs. Those are two different line items. An agency that bundles them together to obscure the actual markup is a pricing problem, not a technology one.

What good looks like: the agency names specific tools (not just “AI”), describes the human review step by name and by person, and can tell you which parts of a deliverable are AI-drafted versus AI-assisted versus fully human. If they cannot answer that breakdown for their own process, they have not actually thought through where AI creates value versus where it just creates output.

Criterion 2: Verifiable outcomes tied to revenue

This is the criterion that separates agencies selling activity from agencies selling results. Ask for a case study with a named metric, a measured change, and a defined timeframe. Something checkable.

“Massive AI-driven gains” is marketing copy. “Contribution margin improved 14% over a 90-day window on this specific account” is a claim you can verify. The deeper version of this question, especially for eCommerce: does the agency report on business-level metrics like Marketing Efficiency Ratio and contribution margin, or only platform-reported ROAS? Platform ROAS is the easiest number to inflate and the least connected to what you actually keep. An agency that leads with it is optimizing your evaluation of them, not your business.

graphical user interface

Criterion 3: Named AI-search visibility capability

This is the newest criterion, and the one most agencies still do not have a real answer for. A growing share of buyer research now happens inside ChatGPT, Perplexity, and Google’s AI Overviews before a shopper ever clicks a paid ad or a traditional search result. Zero-click search behavior has been climbing, which means ranking on page one of classic search results matters less on its own than it used to.

The discipline built around this is called Answer Engine Optimization or Generative Engine Optimization: structuring content so AI engines cite your brand in the answers they generate, rather than just crawling and ranking it the old way. An agency without a real answer here is not behind on a minor tactic. It is missing an entire discovery channel that is actively growing.

What good looks like: a named service or workflow for AI-search visibility, not a vague mention of “staying current,” but an actual method involving structured data, direct-answer content blocks, and consistent entity presence across authoritative sources.

Pro tip: Before your first call, open ChatGPT or Perplexity and search “best [your product category] brands.” If your brand does not show up, that is a real gap costing you buyers who never reach your paid ads at all. Ask every agency on your shortlist directly how they would close it.

Criterion 4: Diagnosis before prescription

A real agency asks about your current channel mix, margins, average order value, and lifetime value before recommending anything. If the recommendation arrives before the diagnosis, every prospect gets the same “paid social plus email” pitch regardless of their specific numbers, that is a sign you are being sold a product, not given a strategy.

This matters more with AI in the mix, not less. AI tools make it cheap to generate a plausible-sounding strategy deck fast. That makes the diagnosis step easier to skip and easier to fake, which is exactly why it is worth checking for directly instead of assuming it happened. What good looks like: the agency asks for your actual numbers, not vanity metrics, before proposing a channel mix, and can explain why your specific margin structure or catalog size points toward one strategy over another.

Criterion 5: Category and vertical fit

An agency that has scaled businesses in your category will ramp faster than a generalist learning your vertical from scratch. This has always mattered, but it matters more now because AI tooling has made execution more commoditized. The differentiator has shifted from “can they run ads” to “do they understand your category’s specific economics”: return-rate-adjusted ROAS for apparel, or replenishment timing for consumables, for example.

What good looks like: specific, named experience in businesses comparable to yours in catalog size, margin structure, and channel mix, not just adjacent.

The six questions to bring to every call

Use the same list across every agency on your shortlist so the answers are actually comparable. Score the specificity of the answers, not just whether they gave one.

  1. Walk me through your workflow tool by tool. Where does AI touch this, and where does a human review it before I see it?
  2. Show me a case study with exact numbers and dates, tied to revenue or margin, not just traffic or platform ROAS.
  3. What is your AI-search visibility capability, and how do you measure it?
  4. What do you need from me (channel mix, margins, AOV) before you will recommend a strategy?
  5. How do you separate your strategy fee from platform and API costs in your pricing?
  6. What is your experience with businesses at our specific catalog size and margin structure?

The bottom line

Choosing an AI marketing agency in 2026 is not about finding the one with the most confident AI language on its homepage. It is about finding the one that can answer five specific questions with evidence instead of adjectives: where AI sits in their process, whether outcomes are verifiable and tied to revenue, whether they have a real AI-search capability, whether they diagnose before they prescribe, and whether they actually know your category.

Any agency that answers all five clearly is operating at a genuinely different level than one that just added “AI-powered” to its tagline.

Stay on top of AI & Automation with BizStack Newsletter
BizStack  —  Entrepreneur’s Business Stack
Logo