Which AI Engine Names the Most Tools? We Asked 8 the Same 16 Questions

The same 16 buying questions, put to 8 AI engines, produced wildly different numbers of named tools. GPT-OSS 120B named 115 distinct products. Llama 3.3 70B via Groq named 40. That is a 2.9x gap in how many companies each engine is willing to put in front of a buyer.

If an AI engine will not recommend your product, this is the number that decides where to start. A blended visibility score cannot see it.

Distinct tools named, 16 identical questions
GPT-OSS 120B115
Gemini 2.5 Flash-Lite104
Perplexity101
Cohere Command-A97
Qwen 3.689
ChatGPT76
Llama 4 Scout58
Llama 3.3 70B via Groq40

GPT-OSS 120B names 2.9x as many distinct tools as Llama 3.3 70B via Groq on the same questions. Of the 268 tools named across all 8 engines, 128 were named by exactly one of them.

How many different tools each engine will name

Every engine below answered the same 16 questions. "Distinct tools" counts the different products it named at least once. "Per answer" is how many it lists in a typical response. "Only here" counts tools that no other engine named anywhere in this set.

#EngineDistinct toolsPer answerOnly hereExamples it alone named
1 GPT-OSS 120B
model memory
115
7.7 28 HyperWrite, ContentBot, Peppertype.ai
2 Gemini 2.5 Flash-Lite
live web
104
7 28 Jasper AI, Sintra AI, AdCreative.ai
3 Perplexity
live web
101
6.8 23 Bitrix24, Semrush One, Sight AI
4 Cohere Command-A
model memory
97
6.6 15 Wordtune, Sudowrite, ClosersCopy
5 Qwen 3.6
model memory
89
6.3 12 Brandwatch, WordPress, Contentful
6 ChatGPT
live web
76
5.4 11 Monday CRM, Writer, Kit
7 Llama 4 Scout
model memory
58
4.1 10 HubSpot Sales, Plausible, Matomo
8 Llama 3.3 70B via Groq
model memory
40
3 1 WordLift

Breadth is the size of the door, not the quality of the answer

Narrow, Llama 3.3 70B via Groq

40 tools

Decisive. Answers with a short list.

  • 3 tools in a typical answer
  • 1 tool nobody else named
  • Almost entirely the consensus set
  • Very hard for a newcomer to enter

Wide, GPT-OSS 120B

115 tools

Expansive. Answers with a long list.

  • 7.7 tools in a typical answer
  • 28 tools nobody else named
  • Reaches well past the consensus set
  • Room for a name it has not used before
Neither column is "better". A buyer who wants one answer is served by the left; a vendor who is nobody's obvious answer needs the right. "Are we visible in AI" is the wrong question. Visible where, and how crowded is that door is the right one.

Does live web search make an engine more generous?

The obvious hypothesis is that engines with live retrieval name more tools, because they can see the current market rather than recalling it. The measurement does not support it.

Live retrieval · 3 engines

94

average distinct tools named

Model memory · 5 engines

80

average distinct tools named

Retrieval changes which tools appear, not how many. A retrieval engine surfaces what was written about recently; a memory engine surfaces what was written about a lot. Similar-length lists, different contents, and a considerably less convenient result for anyone selling one blended visibility number.

The tools only one engine has ever heard of

128 tools were named by exactly one engine: real products that exist in a single answer stream. For a vendor that position is easy to mistake for success. Check the engine that names you and visibility looks solved; check any other and you are absent.

The pattern also runs the other way. Llama 3.3 70B via Groq named just 1 tool that nobody else did, which means its shortlist is almost entirely drawn from the consensus set. An engine like that is nearly impossible to break into with a new name, because it is not reaching for new names at all.

No engine is a substitute for another

If two engines returned broadly the same tools, a vendor could treat them as one target and halve the work. Across all 28 possible pairs, none of them does.

The most similar pair in the entire set is Llama 3.3 70B via Groq and Llama 4 Scout, and they share just 48 percent of the tools they name between them. The median pair shares 28 percent. The least similar, Llama 3.3 70B via Groq and Perplexity, share 19 percent.

Engine pairTools in commonShare of their combined list
Llama 3.3 70B via Groq + Llama 4 Scout3248%
Cohere Command-A + Qwen 3.65643%
Perplexity + ChatGPT5242%
Cohere Command-A + GPT-OSS 120B6241%
Gemini 2.5 Flash-Lite + Perplexity5739%
the other 22 pairs fall between these and the lowest
Llama 3.3 70B via Groq + Perplexity2319%

The detail worth pausing on is which pair came first. Llama 3.3 70B via Groq and Llama 4 Scout are models from the same family, and they still disagree about more than half of what they name. If two releases of the same lineage diverge that far, the assumption that any two engines can be treated as one audience does not survive.

The practical reading: there is no shortcut where optimising for a popular engine quietly covers the rest. Every engine you care about is separate work, and the 28-pair table above is the reason a single blended AI-visibility score cannot describe a real position.

The 12 tools every engine names, and the 128 only one does

Counting tools per engine says how wide each door is. Counting engines per tool says something sharper: how much of the market is genuinely settled.

Across the shared question set, 8 engines named 268 distinct tools between them. The distribution is not close to even.

Named byToolsShare
8 of 8 engines124%
7 of 8 engines93%
6 of 8 engines135%
5 of 8 engines166%
4 of 8 engines135%
3 of 8 engines2911%
2 of 8 engines4818%
1 of 8 engines12848%

128 tools, 48 percent of everything named, were known to exactly one engine. At the other end, only 12 tools, 4 percent, were named by all 8. The market as AI describes it is a very small agreed core wrapped in a very large contested fringe.

The core itself is worth naming, because it is the hardest real estate in the category to take:

Those 12 are not the best tools in their categories by any measure taken here. They are the ones every engine reaches for regardless of what it was asked, which is a different and more durable kind of position. Displacing one is not a content problem.

Breadth moves by category, which breaks the simple advice

The obvious action from the table near the top of this page is "target the widest engine". The data does not quite support that, and the honest version is more useful.

GPT-OSS 120B is the widest engine overall, but it is the widest in only 2 of the 16 categories measured. Breadth is not a stable trait an engine carries into every subject. It is something an engine does in a particular category, and the ranking reshuffles as the subject changes.

The categories themselves vary as much as the engines do. Lead generation drew 28 distinct tools across all engines; Project management drew 13. A crowded, unsettled category leaves room in every engine's answer. A consolidated one closes the door on all of them at once.

CategoryDistinct tools namedNarrowest engineWidest engine
Lead generation
28
38
GEO / AI visibility
26
18
AI chatbots
24
38
Scheduling
24
38
Marketing automation
23
38
Landing page builders
22
38
AI copywriting
19
38
CRM
18
38

Read that as a room-availability table rather than a quality ranking. In Lead generation the engines collectively found 28 tools worth mentioning, and even the narrowest engine listed 3. In the tightest categories the narrowest engine gets down to 1, which is a shortlist with no vacancy on it.

So the instruction is not "target the widest engine". It is: find the widest engine in your category, and check whether your category has any room in it at all before spending anything on the attempt.

How this was measured, including what we had to correct for

Each engine was asked the same buying questions, and every tool named in the answer was recorded in order. Two properties of the raw data would have made a naive version of this article wrong, and both are worth stating plainly.

The samples were not equal. Some engines answered 48 questions and others answered 16. Distinct-tools-named rises with the number of questions asked, so comparing across unequal samples measures our sampling schedule rather than the engines. Every figure on this page uses only the 16 questions that all 8 engines answered. On the full uneven data the headline gap looks considerably larger; that version of the number is not reported here because it is not about the engines.

Some captured names are not names. Engines that answer in prose sometimes yield fragments rather than products, such as a label like "Tools:" attached to a name, or a name trailed by a clause of description. A filter drops entries that are too long, contain sentence punctuation, or begin with a stop word. It rejected 2 percent of raw mentions. This matters more than the size suggests, because prose-heavy engines produce more fragments, so the noise inflates precisely the engines that would otherwise look most expansive.

Uncorrected, not reported

4.0x

Unequal samples, unfiltered names. Bigger, and about our sampling schedule rather than the engines.

Corrected, reported here

2.9x

Shared 16-question set, name filter applied. Smaller, and about the engines.

Both corrections push the same way: they shrink the gap. The finding survives them, which is the only reason it is worth publishing. A result that grows when you control for the instrument is a result to distrust.

What to do with this if you are the tool nobody names

  1. First · is there room

    Measure your category before your engine

    Categories here run from 13 to 28 distinct tools named. At the low end the narrowest engine lists 1. That is a positioning problem, not a visibility one, and content does not move it.

  2. Then · which door

    Widest in YOUR category, not overall

    GPT-OSS 120B leads the whole set and leads only 2 of 16 individual categories. The global ranking is the wrong instrument for a specific decision.

  3. Budget · no shortcuts

    Every engine is separate work

    Closest pair shares 48%, median pair 28%. No engine's coverage meaningfully carries to another, so treat any promise to lift "AI visibility" as one quantity with suspicion.

  4. Report · never average

    Per engine, with a date

    128 of 268 tools appear on exactly one engine. A single-engine check reports success while a buyer on any other engine cannot find you.

What none of this supports is the reflex of writing more content and waiting. The 12 tools every engine names are not there because they published more. They are there because they are the answer the corpus already contains, and entering that core is a market-position change that content follows rather than causes.

Which engines name you, and which do not?

The free AI Visibility check runs this measurement for your category: per engine, with a snapshot date, and no blended score. Given what the 28-pair table above shows about engine overlap, per-engine is the only report that means anything.

Run my free AI Visibility check ›

What this does not show

Breadth is not trafficAn engine with a wide door and few users is worth less than a narrow one your buyers actually use. Nothing here measures audience size.
One snapshotCaptures from a single audit window. A model update can move breadth substantially in either direction.
Scoped to B2B and GTMA category with a dominant vendor compresses every engine's list. Another vertical would produce different absolute numbers.
Naming is not endorsementAppearing in a longer list is not the same as being recommended first, and this page does not conflate the two.

Frequently asked questions

Which AI engine recommends the most different tools?

GPT-OSS 120B named 115 distinct tools across 16 identical buying questions, the widest of the 8 engines measured. Llama 3.3 70B via Groq named 40, the narrowest. Both answered exactly the same questions, so the gap is a property of the engine rather than of the sample.

Does using live web search make an AI engine name more tools?

Not reliably. In this snapshot the 3 engines with live retrieval averaged 94 distinct tools and the 5 answering from model memory averaged 80. Retrieval changes which tools appear far more than how many.

If my tool is invisible to AI, which engine should I target first?

The one with the widest door. An engine that names 115 tools across 16 questions has room for a newcomer in a way an engine naming 40 does not. Breadth is not the same as traffic, so weigh it against where your buyers actually ask.

What does it mean when only one engine names a tool?

Across the 8 engines, 128 tools were named by exactly one of them on this shared question set. For a vendor that is a foothold rather than a win: you exist in one answer stream and are absent from the rest, so a single-engine visibility check would report success and a buyer on any other engine would never see you.

How was this measured?

Every engine was asked the same buying questions and every tool it named was recorded in order. Figures on this page use only the 16 questions that all 8 engines answered, because 32 further questions went to a subset and counting those would measure our sampling rather than the engines. Captured names are filtered to remove prose fragments; that filter rejected 2 percent of raw mentions.

Bottom line

Eight engines, one set of questions, and a 2.9x spread in how many companies each is prepared to name. The engines do not merely disagree about which tool is best, which we have measured before. They disagree about how many tools deserve to be in the conversation at all, and that decides whether a newcomer has any route in.

If you are working on AI visibility, the first question is not what your score is. It is which door you are standing at, and how many people it lets through.

Data: The AI Recommendation Audit (2026), CC-BY, DOI 10.5281/zenodo.20767878. Figures computed on the 16 queries answered by all 8 engines. Related: the AI Search Disagreement Index measures whether engines agree; this page measures how wide each engine casts. See also why ChatGPT will not recommend your SaaS.

Save
Dashboard

From our network

Best AI Tools for Amazon Sellers - bagengine.comBest AI Courses 2026 - edubracket.comBest Accounting Software for Online Sellers - ceocult.com