Do ChatGPT and Claude Recommend Your St. Louis Business? I Ran 30 Tests. They Agreed Once. (2026)

An original test. Ten questions, three AI assistants, one metro. Run 2026-08-28 and 2026-08-29.

Your Google rating is 4.9 and you still cannot tell whether an AI assistant would send anyone your way.

So I ran the test instead of guessing. Ten questions, written down and locked before a single one was asked, phrased the way a person asks rather than the way a keyword tool phrases things. Ten verticals I sell to, from custom jewelers to roofers to family law. Then I put the same ten to ChatGPT, to Claude, and to Perplexity, logged every business each one named, and read what it cited.

Across ten questions, ChatGPT and Claude named the same business first exactly once. On four of the ten they shared a single name out of five or six each. And on one question, one assistant ranked businesses using a source the other assistant told the reader was paid placement.

How do AI assistants decide which local businesses to recommend?

Each one looks for whoever it considers the credible judge in that specific trade, and they do not agree on who that is. For roofers, ChatGPT used manufacturer certification. For the same roofers, Claude used review depth and whether the company had a local physical address. For attorneys both landed on peer-nominated ratings. For HVAC, ChatGPT used a curated roundup site that Claude, the same day, called a lead-generation site where placement is paid for.

There is no single answer to optimize toward, because there is no single answer. There are two or three different machines reading different sources and reaching different conclusions about the same street.

Which raises the fair question: what did each of them read?

What decided each of the ten answers?

Nine different source classes across ten questions, and that is just one assistant. Here is ChatGPT's full run.

Question asked What it ranked on What it cited
Best custom jeweler in St. Louis one community thread Reddit
Best med spa near Clayton review volume med spa directories
Best custom home builder in St. Louis judged awards Home Builders Association of St. Louis
Best HVAC company in St. Charles County review volume, accreditation BBB, a curated roundup site
Best family law attorney in St. Louis peer recognition Super Lawyers, Best Lawyers, Martindale AV
Good roofer in O'Fallon manufacturer certification Owens Corning contractor profiles
Best kids orthodontist in St. Peters board certification, the practice's own site copy American Association of Orthodontists
Wedding photographer in St. Louis style category, press mentions own sites, Martha Stewart Weddings, Brides
Best winery near Augusta operational detail on the winery's own site the winery's own domain
Best landscaping in Wentzville review volume, published service list BBB, Houzz, Angi

Method: ChatGPT, all ten questions, 2026-08-29, in Temporary Chat, which the interface labels "Unpersonalized" and which ignores memory, custom instructions and plugins. Source class is the citation attached to its top picks, not every link in the answer.

Read the roofing row against the HVAC row. Both are home services, and most marketing advice treats them as one bucket. The assistant did not. For HVAC it counted reviews and checked BBB. For roofing it went to owenscorning.com, read the shingle manufacturer's certified-contractor profiles, and told the reader that Correll Roofing is "an Owens Corning Preferred Contractor and BBB Accredited."

Claude, run the same day, added source classes ChatGPT never touched: a DOE Housing Innovation Award, factory-authorized dealer status for Daikin and Trane, a judicial appointment as a county Family Court Special Master, and, on the winery question, four separate local news outlets reporting on who owns the Augusta wineries now.

Different judges in every trade, and different judges per assistant within the same trade. Which is where the numbers get uncomfortable.

Do different AI assistants recommend the same businesses?

Barely. On ten questions asked the same day, ChatGPT and Claude named the same business first once.

Question ChatGPT named first Claude named first Names shared
Custom jeweler Adam Foster Fine Jewelry Adam Foster Fine Jewelry 2
Med spa near Clayton The Shine Spa The Saint Louis Med Spa 4
Custom home builder Hibbs Luxury Homes named nobody 3
HVAC, St. Charles County Design Aire Renaud Brothers 2
Family law attorney Michelle J. Spirn Jennifer R. Piper 1
Roofer, O'Fallon Patriarch Roofing Brody Allen Exteriors 1
Kids orthodontist Otto Orthodontics Humphrey Orthodontics 3 of 3
Wedding photographer Sarah Corbett Photography Shane Michael Studios 1
Winery near Augusta Noboleis Vineyards Montelle Winery 4
Landscaping, Wentzville NxGen Outdoors Precise Outdoors 1

Method: both runs 2026-08-29, same wording, same connection, one in ChatGPT Temporary Chat and one in Claude Incognito. "Names shared" counts businesses appearing in both answers. The builder row's 3 comes from a labeled follow-up probe, described in Method, because Claude named nobody on the first pass.

Look at the pattern in the last column. Three questions produced strong agreement: med spas, kids orthodontists, wineries. Four produced almost none: family law, roofing, wedding photography, landscaping, where two assistants each named five or six businesses and overlapped on exactly one.

The difference is not review density, which was my working theory after the first two runs. It is how many plausible businesses exist. Clayton has a handful of med spas and St. Peters has four orthodontists, so any honest answer set is small and both assistants land on it. Wentzville has hundreds of people who will landscape your yard. When the honest answer set is two hundred, each assistant picks a different five, and there is no "the answer" to be optimized toward.

Some of the divergence is stark. ChatGPT's top HVAC pick, Design Aire, is Claude's fifth. Jerry Kelly, which Claude reported at 7,158 reviews, by far the largest review base in St. Charles County, does not appear in ChatGPT's answer at all. On roofing, ChatGPT's number one never appears in Claude's six, and Claude's number one, with 503 reviews, never appears in ChatGPT's four.

And on one question, the two assistants did not merely disagree. They contradicted each other about what counts as evidence.

Can you trust the "Top 10" list an AI cites?

One assistant ranked businesses on sources the other assistant told the reader are paid placement.

Asked about HVAC in St. Charles County, ChatGPT cited BBB and a curated roundup site: "a very recent 2026 independent roundup from Expertise also places local providers including Premier among its top St. Charles HVAC selections." Its landscaping answer cited Angi and Houzz.

Claude, asked the same HVAC question the same day, opened with this:

"There's no single 'best', and it's worth knowing that most of the 'Top 10 HVAC in St. Charles' lists you'll find (Angi, HomeAdvisor, Expertise, Today's Homeowner) are lead-generation sites where placement is tied to paying for leads, not to quality. Same for the blog posts written by HVAC companies themselves."

Then it did it again on attorneys, and caught something specific:

"Best Lawyers, Super Lawyers, and a Martindale-Hubbell AV Preeminent rating are peer-nomination based and reasonably meaningful. Thumbtack, Yelp, and most 'Top 10 Family Lawyers in St. Louis' blog posts are advertising, the Vantage Group list that ranks high on Google is a firm ranking itself."

A law firm publishing a "best family lawyers" list that ranks the law firm. One assistant identified it and threw it out. And on custom home builders Claude went further, about search generally: "at this tier the search results are mostly SEO copy, and the builders that show up first aren't necessarily the ones doing the best work."

For a business owner this is the part with real money in it. Buying placement on an aggregator can put you in one assistant's answer and get your whole source class discredited in another's, in the same week, to the same customer. It is not a strategy. It is a coin flip you paid for.

Which brings up the thing everyone assumes is the safe answer instead.

Do Google reviews still matter for AI recommendations?

They decide who makes the list far more than they decide the order, and three separate times an assistant argued out loud against using them.

On custom home builders:

"One thing I wouldn't do is choose based on Google stars alone. Some excellent custom builders do relatively few houses per year, so their online review counts can be tiny. Recent HBA results are more informative here because the program specifically judges custom homes on design, functionality, materials, workmanship and related criteria."

On family law attorneys:

"Super Lawyers says its selection process combines independent research and peer evaluation, while Best Lawyers is peer-review based, so I would give those professional credentials more weight than Google ratings alone."

On orthodontists for children:

"For kids, I wouldn't choose solely by reviews, the bigger question is whether the orthodontist recommends observation vs. Phase I/early treatment, explains why treatment is needed now, and gives you a clear total cost."

The ranking followed the reasoning. In the family law answer, the two attorneys with peer credentials and no review count given were placed first and second, and the four firms carrying 4.9 stars with no credential were placed third through sixth. A photographer with 74 reviews outranked one with 117. A 4.3-star winery was picked over a 4.4 on the assistant's own map.

Claude produced the sharpest version. Asked about landscaping in Wentzville, it named Best Lawn Care, noted its 969 reviews, and set it aside anyway: "nearly 1,000 reviews but is really a treatment company, weed control and fertilization, not landscaping." The winner had 149 reviews and was in the right category.

None of this makes reviews worthless. In med spas, HVAC and landscaping, review volume was the stated deciding factor for who appeared at all. Reviews get you considered. Something else decides where you land, and increasingly the assistants say so in the answer.

There is also the question of whether the numbers are right. In the med spa answer, ChatGPT cited two aggregators for the same business in one sentence and they disagreed, 665 reviews and 741, and it reported both without reconciling them. Claude, a day later, said 669. Perplexity had given a different business 1,478 and Claude said 1,481. Across nine businesses I could cross-check between assistants, eight matched within normal one-day accrual. The one that did not, 741, was presented with exactly the same confidence as the eight that did.

What else can an AI assistant do besides recommend you?

Three things showed up in this test that most owners have never considered, and each one costs something.

One. It can refuse to name anyone until the customer is qualified.

Asked for the best custom home builder in St. Louis, Claude named zero businesses. Instead it asked three questions: budget, whether you own land, and which part of the metro. Only after all three were answered did it produce a shortlist. ChatGPT, same question, named five builders immediately.

Claude asked a follow-up question in eight of its ten answers. The consequence for a business is quiet and large: in some categories you are not competing to be named, you are competing to be named after a qualification step you never see and cannot influence. A builder who does $600k homes was filtered out of that answer before any consideration of quality.

Two. It can name you as a warning.

In the landscaping answer, Claude flagged a business by name. Here is the wording, with the name removed: "[a named Wentzville landscaping company] shows up in a lot of searches at 4.5 stars, but there are billing and follow-through complaints mixed into the recent reviews. Worth reading those yourself before you call."

I am withholding that name. The finding is that an assistant will say this about a business, not which business it said it about, and I am not going to repeat an unverified claim about a local company's billing for a page view. Every other business in this article is named because an assistant recommended it, which is a different thing entirely.

That is a 4.5-star business, visible in search, appearing in the answer as a caution. It happened once in thirty runs, and it establishes the category. The recommendation surface is not only a place you can fail to appear. It is a place you can appear negatively, from your own review text.

Three. It can decide the customer meant a different city than the one they typed.

I asked for the best family law attorney in St. Louis. Claude answered:

"Venue probably isn't St. Louis. Lake Saint Louis is in St. Charles County, and Missouri family cases are generally filed where the parties live, so you'd likely be in the 11th Judicial Circuit (St. Charles County Family Court), not St. Louis County or the City."

The question named St. Louis. The assistant overrode it and answered for a different county, based on where it inferred the connection was. On the photography question the same thing was visible in the results card, which was titled "Wedding photographers near Lake Saint Louis" when the question said "the St. Louis area."

This one cuts both ways and it is the most important thing in the test for anyone outside the city core. A business in the county can be surfaced for a query that named the city. A business in the city can be dropped from a query that named it. Neither had anything to do with marketing.

So if the judges differ, the geography moves, and the customer may be qualified before you are considered, what is left to do?

Two things, and both are writing plainly on your own site about how you work.

Nearly every source class in this test is either earned slowly or not yours at all. You cannot decide to be in a Reddit thread or in a St. Louis Magazine piece about winery consolidation. Manufacturer certifications, board certifications, peer ratings and judged awards take years and qualifying, which is the point of them. But two things turned answers repeatedly, and a business owner can publish both this week.

One. Explain your own method, in plain language, on your own site.

Humphrey Orthodontics was recommended for younger children, and the reason ChatGPT gave was not a rating. It was the practice's own explanation of its clinical philosophy: they evaluate children around age seven and may recommend simply monitoring growth rather than starting braces immediately when treatment is not yet necessary. The assistant read that, understood it as a position, and recommended them for that specific customer because of it.

Claude's builder answer worked the same way, citing six builders' own domains and reading capability off them: RESNET partner, specializes in $1M+ custom homes, deliberately low volume since 1986. Its jeweler answer ranked on CAD renderings before casting, in-house bench work, a GIA graduate gemologist, four-week turnarounds. Perplexity named The Diamond Family off their 3D portfolio and CAD design page.

Every one of those is a sentence a business chose to write, and none of them is a claim about being the best.

Two. Publish the operational specifics.

ChatGPT picked Noboleis Vineyards over a higher-rated neighbor and justified it with published facts: an 84-acre property, panoramic hilltop views, outdoor seating, wine flights, pizza and shareables, live music on weekends. Then it quoted Saturday hours and built an itinerary around them.

Hours turned up in three separate verticals across two assistants. Claude used them to separate five med spas ("open Saturdays, which most of the others aren't", "only open Wed to Fri, so booking takes planning"), flagged an orthodontist as "closed Fridays, and Monday starts at 10am," and warned that Augusta wineries run reduced off-season hours. Service lists did the same work in landscaping, where the rating decided who made the list and the published service list decided what each one got recommended for.

A star rating tells an assistant you are good. A service page and an hours page tell it what you are good at and when, and "best for X" is the sentence that sends someone to your phone.

Where should a business start, based on its own vertical?

Count your competitors first, because that decides whether any of this is winnable.

If your category is small and well-defined, meaning a customer's honest shortlist is five businesses, the assistants converge and the lever is real. That was true for med spas, kids orthodontists and wineries. Reviews and category fit get you onto a list that is mostly the same list everywhere.

If your category is credential-governed, reviews will not carry you past someone holding the credential. Roofers, attorneys, orthodontists and custom builders all ranked behind an external judge: a manufacturer, a peer body, a board, an association, a court. If you hold that credential and it is not visible anywhere a machine can read it, that is your gap and it is a one-afternoon fix. If you do not hold it, that is a business decision rather than a marketing one.

If your category is crowded, with hundreds of plausible providers and no dominant judge, there is no single answer to optimize toward. Two assistants named twelve landscapers between them and agreed on one. That is not a failure of your marketing, it is the shape of your category, and the honest play is the two levers above: be legible about what you specifically do, so that when an assistant is looking for someone who does drainage or hardscape or heirloom redesign, you are the one who said so plainly. Which is harder than it sounds around here, since nearly half the local businesses I checked this summer had no website at all.

And in every category, check your own review text. One assistant read recent reviews closely enough to warn a customer about billing complaints at a 4.5-star company. Your rating is not the only thing being read.

None of this tells you whether any of it is working for you. That takes a different instrument: checking whether AI assistants are sending you any traffic at all, which your analytics can answer today, and which I ran on my own site with a result I did not enjoy.

If I run the same question twice, do I get the same answer?

No, and what changes versus what holds is the most useful thing in this test.

I asked the custom home builder question a second time, same wording, about forty minutes after the first, in a fresh session. The source held: the St. Louis Home Builders Association. The top pick held: Hibbs Luxury Homes. The reasoning held, Custom Home of the Year awards quoted by price category, down to the same brackets. The refusal to name one winner held.

The roster did not. The first run named five builders, the second named seven, and two companies appeared that had not been there before. The first run's blunt warning about Google stars did not reappear at all.

The mechanism is stable. The roster is a snapshot. Which judge an assistant consults, and what it weights, reproduced perfectly. Exactly who got named did not.

Worth knowing before you go run these queries yourself, because you will see a name I did not. That does not mean the pattern is wrong. It means the list reshuffles and the door it comes through does not.

The assistants are not judging your marketing. Each one is hunting for whoever it thinks already judges your industry, and then repeating that verdict to a customer standing in your parking lot with a phone. They do not agree on the judge. One of them will tell that customer the list your competitor paid to be on is advertising.

Find out who reads what in your trade. Everything else is guessing at machines that already told you, out loud, what they trust.

The checklist our paid audit runs is published in full. All 160 checks, website to AI readability, free, no email, no gate. Run them on your own business, or have us run them and write the plan: The Blueprint.

Method

The query set was written and locked on 2026-08-28, before any assistant was asked anything. Ten questions, one per vertical, phrased the way a person asks an AI assistant rather than the way a keyword tool phrases a term. No question was added or dropped after results came back.

ChatGPT, 2026-08-29, all ten completed, in Temporary Chat.

Claude, 2026-08-29, all ten completed, in Incognito chat, model shown as Opus 5.

Perplexity, logged out, 2026-08-28. Three of ten completed. Questions four onward returned "Sign up and repeat your request," the anonymous rate limit. That is its own small finding: a stranger gets roughly three free answers before the wall.

Declared deviation on session state. The locked plan called for logged-out sessions, which is what a stranger sees. ChatGPT and Claude both require accounts, so the substitutes were their unpersonalized modes. ChatGPT's Temporary Chat ignores memory, custom instructions, files and plugins and does not write to history. Claude's Incognito chat is described in its own interface as not saved, not added to memory, and not used to train models. Both are the closest available approximation to a stranger's view. Neither is identical to logged out.

★ What this test cannot tell you, and it matters. Neither unpersonalized mode suppresses location. Claude stated the inferred location explicitly, writing "easiest drive from Lake St. Louis," "closest full-service shop to Lake Saint Louis," and, on the family law question, "venue probably isn't St. Louis. Lake Saint Louis is in St. Charles County." Both runs came from the same connection in the western St. Charles County area, and both produced results weighted toward it. So this is what these assistants tell someone asking from the St. Louis area, not what they tell someone anywhere. A reader running these same questions from a different place should expect a different roster, and the same is true for their own customers.

Declared deviation on the builder question. The protocol is one question, no follow-ups. Claude answered the custom home builder question by naming zero businesses and asking three qualifying questions instead. Because a zero-name answer tells a business owner nothing about how to be named, I ran one labeled follow-up on that question only, answering $1.5M+, already own a lot, West County / Wildwood, chosen to match the tier ChatGPT had framed. It is not part of the ten-question set and it is marked as a probe wherever it is used.

Reproducibility. The custom home builder question was deliberately asked twice on ChatGPT the same day in separate fresh sessions. The cited source, top pick, ranking basis and refusal to name a single winner were identical. The roster grew from five names to seven. Claims here about what an assistant ranks on are reproducible. Claims about exactly which businesses were named are a snapshot of that run on that date.

Recorded per run: which businesses were named and in what order, what was cited, whether ratings or review counts appeared, whether the assistant deferred to a directory, and its own stated criteria where it gave any.

Review-count cross-check. Nine businesses had review counts reported by more than one assistant. Eight matched within normal one-day accrual, including one, Precise Outdoors at 4.9 with 149 reviews, that matched exactly. The exception was The Shine Spa, where ChatGPT published 665 and 741 in the same sentence; Claude reported 669 a day later, which fits 665 and does not fit 741.

Not verified here: whether every named business is currently in good standing, and whether any assistant's characterization of a business is fair. Businesses are named because an assistant named them, not because this article endorses them. One business name is withheld: the only one an assistant characterized negatively. The claim was not independently verified, so it is not repeated with a name attached.

Counts: ChatGPT named 50 distinct businesses across its ten answers. Claude named 53. Nine distinct lead source classes appeared in ChatGPT's ten, and Claude added several more, including a federal energy award, factory dealer status, a judicial appointment and four local news outlets.

Frequently asked questions

How do AI assistants decide which local businesses to recommend?

Each one looks for what it treats as the credible judge in that specific trade, and different assistants pick different judges. In a ten-question test across two assistants, judges included a manufacturer's certified-contractor list, peer-nominated attorney ratings, a trade association's judged awards, a federal energy award, factory-authorized dealer status, a court appointment, local news reporting, and Google review volume. Reviews decided who appeared more often than they decided the order.

Do ChatGPT and Claude recommend the same businesses?

Rarely. Asked the same ten questions on the same day from the same connection, they named the same business first once. On four of the ten they shared exactly one name out of five or six each. Agreement was high only where the category was small, such as med spas and children's orthodontists, and near zero in crowded categories like landscaping and wedding photography.

Can an AI assistant recommend against my business?

Yes. In one of thirty runs, an assistant named a 4.5-star landscaping company and told the reader there were billing and follow-through complaints in its recent reviews, and to read them before calling. Recent review text is being read, not just the star average.

Do Google reviews affect AI recommendations?

They decide who makes the list more than they decide the order. Three separate times an assistant argued explicitly against choosing on stars, once saying it would give professional credentials more weight than Google ratings alone. One assistant set aside a business with nearly 1,000 reviews because it was in an adjacent category to the one asked about, and recommended a company with 149.

What can a small business do to get recommended by AI assistants?

Two things a business fully controls turned answers repeatedly: explaining how you work in plain language on your own site, and publishing operational specifics such as services, hours, and what is on the property. An orthodontic practice was recommended for younger children because its site explained when it recommends waiting instead of starting treatment. A winery was picked over a higher-rated neighbor on the strength of published details about its property, food, live music and Saturday hours.

If I ask an AI assistant the same question twice, do I get the same businesses?

Not necessarily. Asking one question twice in a single day, forty minutes apart in fresh sessions, produced the same cited source, the same top pick and the same stated reasoning, but a roster that grew from five businesses to seven. The mechanism is reproducible. The exact list is a snapshot.

Does my location change which businesses an AI recommends?

Substantially, and it can override the city named in the question. Asked for the best family law attorney in St. Louis, one assistant replied that the venue probably was not St. Louis, because it had inferred the person was in St. Charles County, and answered for a different circuit court. This happened in a session mode that suppresses memory and history but not location.

Work with Hit My Algo

Every client starts with The Blueprint: 160 checks across everything your brand does online, and the plan that comes out of them, $489 one time. The same checklist is free on that page, no email, no form, no gate, if you would rather run it yourself.

After that, we build and run the whole marketing system on accounts you own: content, socials, funnel, and analytics, through The Handshake Framework. See everything we run, the full ladder, book a call, or ask us anything, no call needed.