Which AI Engine Names the Most Tools? We Asked 8 the Same 16 Questions
The same 16 buying questions, put to 8 AI engines, produced wildly different numbers of named tools. GPT-OSS 120B named 115 distinct products. Llama 3.3 70B via Groq named 40. That is a 2.9x gap in how many companies each engine is willing to put in front of a buyer.
If an AI engine will not recommend your product, this is the number that decides where to start. A blended visibility score cannot see it.
GPT-OSS 120B names 2.9x as many distinct tools as Llama 3.3 70B via Groq on the same questions. Of the 268 tools named across all 8 engines, 128 were named by exactly one of them.
How many different tools each engine will name
Every engine below answered the same 16 questions. "Distinct tools" counts the different products it named at least once. "Per answer" is how many it lists in a typical response. "Only here" counts tools that no other engine named anywhere in this set.
| # | Engine | Distinct tools | Per answer | Only here | Examples it alone named |
|---|---|---|---|---|---|
| 1 | GPT-OSS 120B model memory |
7.7 | 28 | HyperWrite, ContentBot, Peppertype.ai | |
| 2 | Gemini 2.5 Flash-Lite live web |
7 | 28 | Jasper AI, Sintra AI, AdCreative.ai | |
| 3 | Perplexity live web |
6.8 | 23 | Bitrix24, Semrush One, Sight AI | |
| 4 | Cohere Command-A model memory |
6.6 | 15 | Wordtune, Sudowrite, ClosersCopy | |
| 5 | Qwen 3.6 model memory |
6.3 | 12 | Brandwatch, WordPress, Contentful | |
| 6 | ChatGPT live web |
5.4 | 11 | Monday CRM, Writer, Kit | |
| 7 | Llama 4 Scout model memory |
4.1 | 10 | HubSpot Sales, Plausible, Matomo | |
| 8 | Llama 3.3 70B via Groq model memory |
3 | 1 | WordLift |
Breadth is the size of the door, not the quality of the answer
Narrow, Llama 3.3 70B via Groq
40 tools
Decisive. Answers with a short list.
- 3 tools in a typical answer
- 1 tool nobody else named
- Almost entirely the consensus set
- Very hard for a newcomer to enter
Wide, GPT-OSS 120B
115 tools
Expansive. Answers with a long list.
- 7.7 tools in a typical answer
- 28 tools nobody else named
- Reaches well past the consensus set
- Room for a name it has not used before
Does live web search make an engine more generous?
The obvious hypothesis is that engines with live retrieval name more tools, because they can see the current market rather than recalling it. The measurement does not support it.
Live retrieval · 3 engines
94
average distinct tools named
Model memory · 5 engines
80
average distinct tools named
The tools only one engine has ever heard of
128 tools were named by exactly one engine: real products that exist in a single answer stream. For a vendor that position is easy to mistake for success. Check the engine that names you and visibility looks solved; check any other and you are absent.
The pattern also runs the other way. Llama 3.3 70B via Groq named just 1 tool that nobody else did, which means its shortlist is almost entirely drawn from the consensus set. An engine like that is nearly impossible to break into with a new name, because it is not reaching for new names at all.
No engine is a substitute for another
If two engines returned broadly the same tools, a vendor could treat them as one target and halve the work. Across all 28 possible pairs, none of them does.
The most similar pair in the entire set is Llama 3.3 70B via Groq and Llama 4 Scout, and they share just 48 percent of the tools they name between them. The median pair shares 28 percent. The least similar, Llama 3.3 70B via Groq and Perplexity, share 19 percent.
| Engine pair | Tools in common | Share of their combined list | |
|---|---|---|---|
| Llama 3.3 70B via Groq + Llama 4 Scout | 32 | 48% | |
| Cohere Command-A + Qwen 3.6 | 56 | 43% | |
| Perplexity + ChatGPT | 52 | 42% | |
| Cohere Command-A + GPT-OSS 120B | 62 | 41% | |
| Gemini 2.5 Flash-Lite + Perplexity | 57 | 39% | |
| the other 22 pairs fall between these and the lowest | |||
| Llama 3.3 70B via Groq + Perplexity | 23 | 19% | |
The detail worth pausing on is which pair came first. Llama 3.3 70B via Groq and Llama 4 Scout are models from the same family, and they still disagree about more than half of what they name. If two releases of the same lineage diverge that far, the assumption that any two engines can be treated as one audience does not survive.
The practical reading: there is no shortcut where optimising for a popular engine quietly covers the rest. Every engine you care about is separate work, and the 28-pair table above is the reason a single blended AI-visibility score cannot describe a real position.
The 12 tools every engine names, and the 128 only one does
Counting tools per engine says how wide each door is. Counting engines per tool says something sharper: how much of the market is genuinely settled.
Across the shared question set, 8 engines named 268 distinct tools between them. The distribution is not close to even.
| Named by | Tools | Share | |
|---|---|---|---|
| 8 of 8 engines | 12 | 4% | |
| 7 of 8 engines | 9 | 3% | |
| 6 of 8 engines | 13 | 5% | |
| 5 of 8 engines | 16 | 6% | |
| 4 of 8 engines | 13 | 5% | |
| 3 of 8 engines | 29 | 11% | |
| 2 of 8 engines | 48 | 18% | |
| 1 of 8 engines | 128 | 48% |
128 tools, 48 percent of everything named, were known to exactly one engine. At the other end, only 12 tools, 4 percent, were named by all 8. The market as AI describes it is a very small agreed core wrapped in a very large contested fringe.
The core itself is worth naming, because it is the hardest real estate in the category to take:
- Copy.ai
- Salesforce
- Mailchimp
- Klaviyo
- Ahrefs
- Outreach
- Amplitude
- Mixpanel
- Hootsuite
- Buffer
- Sprout Social
- Typeform
Those 12 are not the best tools in their categories by any measure taken here. They are the ones every engine reaches for regardless of what it was asked, which is a different and more durable kind of position. Displacing one is not a content problem.
Breadth moves by category, which breaks the simple advice
The obvious action from the table near the top of this page is "target the widest engine". The data does not quite support that, and the honest version is more useful.
GPT-OSS 120B is the widest engine overall, but it is the widest in only 2 of the 16 categories measured. Breadth is not a stable trait an engine carries into every subject. It is something an engine does in a particular category, and the ranking reshuffles as the subject changes.
The categories themselves vary as much as the engines do. Lead generation drew 28 distinct tools across all engines; Project management drew 13. A crowded, unsettled category leaves room in every engine's answer. A consolidated one closes the door on all of them at once.
| Category | Distinct tools named | Narrowest engine | Widest engine |
|---|---|---|---|
| Lead generation | 3 | 8 | |
| GEO / AI visibility | 1 | 8 | |
| AI chatbots | 3 | 8 | |
| Scheduling | 3 | 8 | |
| Marketing automation | 3 | 8 | |
| Landing page builders | 3 | 8 | |
| AI copywriting | 3 | 8 | |
| CRM | 3 | 8 |
Read that as a room-availability table rather than a quality ranking. In Lead generation the engines collectively found 28 tools worth mentioning, and even the narrowest engine listed 3. In the tightest categories the narrowest engine gets down to 1, which is a shortlist with no vacancy on it.
So the instruction is not "target the widest engine". It is: find the widest engine in your category, and check whether your category has any room in it at all before spending anything on the attempt.
How this was measured, including what we had to correct for
Each engine was asked the same buying questions, and every tool named in the answer was recorded in order. Two properties of the raw data would have made a naive version of this article wrong, and both are worth stating plainly.
The samples were not equal. Some engines answered 48 questions and others answered 16. Distinct-tools-named rises with the number of questions asked, so comparing across unequal samples measures our sampling schedule rather than the engines. Every figure on this page uses only the 16 questions that all 8 engines answered. On the full uneven data the headline gap looks considerably larger; that version of the number is not reported here because it is not about the engines.
Some captured names are not names. Engines that answer in prose sometimes yield fragments rather than products, such as a label like "Tools:" attached to a name, or a name trailed by a clause of description. A filter drops entries that are too long, contain sentence punctuation, or begin with a stop word. It rejected 2 percent of raw mentions. This matters more than the size suggests, because prose-heavy engines produce more fragments, so the noise inflates precisely the engines that would otherwise look most expansive.
Uncorrected, not reported
4.0x
Unequal samples, unfiltered names. Bigger, and about our sampling schedule rather than the engines.
Corrected, reported here
2.9x
Shared 16-question set, name filter applied. Smaller, and about the engines.
What to do with this if you are the tool nobody names
- First · is there room
Measure your category before your engine
Categories here run from 13 to 28 distinct tools named. At the low end the narrowest engine lists 1. That is a positioning problem, not a visibility one, and content does not move it.
- Then · which door
Widest in YOUR category, not overall
GPT-OSS 120B leads the whole set and leads only 2 of 16 individual categories. The global ranking is the wrong instrument for a specific decision.
- Budget · no shortcuts
Every engine is separate work
Closest pair shares 48%, median pair 28%. No engine's coverage meaningfully carries to another, so treat any promise to lift "AI visibility" as one quantity with suspicion.
- Report · never average
Per engine, with a date
128 of 268 tools appear on exactly one engine. A single-engine check reports success while a buyer on any other engine cannot find you.
What none of this supports is the reflex of writing more content and waiting. The 12 tools every engine names are not there because they published more. They are there because they are the answer the corpus already contains, and entering that core is a market-position change that content follows rather than causes.
Which engines name you, and which do not?
The free AI Visibility check runs this measurement for your category: per engine, with a snapshot date, and no blended score. Given what the 28-pair table above shows about engine overlap, per-engine is the only report that means anything.
Run my free AI Visibility check ›What this does not show
Frequently asked questions
Which AI engine recommends the most different tools?
GPT-OSS 120B named 115 distinct tools across 16 identical buying questions, the widest of the 8 engines measured. Llama 3.3 70B via Groq named 40, the narrowest. Both answered exactly the same questions, so the gap is a property of the engine rather than of the sample.
Does using live web search make an AI engine name more tools?
Not reliably. In this snapshot the 3 engines with live retrieval averaged 94 distinct tools and the 5 answering from model memory averaged 80. Retrieval changes which tools appear far more than how many.
If my tool is invisible to AI, which engine should I target first?
The one with the widest door. An engine that names 115 tools across 16 questions has room for a newcomer in a way an engine naming 40 does not. Breadth is not the same as traffic, so weigh it against where your buyers actually ask.
What does it mean when only one engine names a tool?
Across the 8 engines, 128 tools were named by exactly one of them on this shared question set. For a vendor that is a foothold rather than a win: you exist in one answer stream and are absent from the rest, so a single-engine visibility check would report success and a buyer on any other engine would never see you.
How was this measured?
Every engine was asked the same buying questions and every tool it named was recorded in order. Figures on this page use only the 16 questions that all 8 engines answered, because 32 further questions went to a subset and counting those would measure our sampling rather than the engines. Captured names are filtered to remove prose fragments; that filter rejected 2 percent of raw mentions.
Bottom line
Eight engines, one set of questions, and a 2.9x spread in how many companies each is prepared to name. The engines do not merely disagree about which tool is best, which we have measured before. They disagree about how many tools deserve to be in the conversation at all, and that decides whether a newcomer has any route in.
If you are working on AI visibility, the first question is not what your score is. It is which door you are standing at, and how many people it lets through.
Data: The AI Recommendation Audit (2026), CC-BY, DOI 10.5281/zenodo.20767878. Figures computed on the 16 queries answered by all 8 engines. Related: the AI Search Disagreement Index measures whether engines agree; this page measures how wide each engine casts. See also why ChatGPT will not recommend your SaaS.