> ## Content Index
> Fetch the complete content index at: https://seeindie.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Ubersuggest vs Semrush: I ran the same keywords through both, four times
- URL: https://seeindie.com/blog/ubersuggest-vs-semrush/
- Published: 2026-09-08T11:57:24.000Z
- Updated: 2026-09-08T11:57:24.000Z
- Description: Four head-to-head tests across three niches: the search volumes are identical in origin, one keyword in ten diverges by more than 5x, and an un-refreshed difficulty score is off by an average of 23 points.
- Author: Isabella-Viktoria Quilez
- Tags: tools, SEO, review

Every SEO tool comparison you can find is written from feature lists and pricing pages. This one is written from four head-to-head tests on the same keyword sets, on the same day, in three niches and two countries — plus one accident that turned out to explain most of the disagreement between them.

The short version 
- **The search volumes are identical in origin.** Both tools output Google Keyword Planner buckets — 100% of values, four rounds out of four. The cheap one does not invent its numbers.
- **Roughly one keyword in ten diverges by more than 5×**, clustering on misspellings, spacing and apostrophes, because the two tools normalise variants differently.
- **An un-refreshed difficulty score is off by an average of 23 points** in Ubersuggest and 7 in Semrush. Refresh changes difficulty and nothing else.
- **Refreshed against refreshed, the two tools largely agree** — rank correlation rises from 0.665 to 0.909\. Most of the "they disagree" story was stale data.
- **Single-digit difficulty means "never calculated"**, not "easy". Nothing on screen tells you which.

**Disclosures.** I pay for Ubersuggest myself on a lifetime Enterprise licence, and I use **Semrush One Pro+ ($299/month)** daily at work where my employer pays for it. Having both is how these tests were possible at all. **No affiliate links**; nothing here earns anything. There's no switching story being sold, and no reason to talk either tool down. 

On this page [What was tested](#setup) [The volumes](#volumes) [The 5× outliers](#outliers) [The refresh finding](#refresh) [What it can't tell you](#limits) [Which to buy](#verdict) 

## What was tested

|            | Niche                                    | Market             | Keywords compared           |
| ---------- | ---------------------------------------- | ------------------ | --------------------------- |
| **Test 1** | Gaming                                   | Google US, English | 47 volumes · 100 difficulty |
| **Test 2** | Lighting / LED / solar, B2B              | Google DE          | 54 volumes · 32 difficulty  |
| **Test 3** | Furniture / children's beds, consumer    | Google DE          | 84 volumes · 72 difficulty  |
| **Test 4** | Gaming, built to probe the failure modes | Google US, English | 43 volumes · 50 difficulty  |

Round four was designed rather than collected: fifty keywords in blocks, each block aimed at a failure mode the first three rounds had turned up — compounds against spaced variants, singulars against plurals, misspellings, names that are also ordinary words, long tail, and a control group of head terms both tools would certainly know.

**The result:** exactly two findings survive all four rounds. Almost everything else reverses direction between niches, which is itself the most useful thing here.

## What held everywhere: the volumes are Google's numbers

**Both tools output Google Keyword Planner buckets.** Keyword Planner doesn't return exact volumes; it returns rungs on a logarithmic ladder stepped by roughly ×1.22: … 880, 1,000, 1,300, 1,600, 1,900, 2,400, 2,900, 3,600, 4,400, 5,400, 6,600, 8,100, 9,900 …

|        | Semrush on the ladder | Ubersuggest on the ladder |
| ------ | --------------------- | ------------------------- |
| Test 1 | 6,082 / 6,082         | 36 / 36                   |
| Test 2 | 19 / 19               | 20 / 20                   |
| Test 3 | 40 / 40               | 38 / 38                   |
| Test 4 | 26 / 26               | 22 / 23                   |

**Near-perfect across four rounds, three niches, two countries and two languages.** Neither tool is modelling its own estimate — both pass the same Google source through. One exception in the whole series: `steam` came back at 2,740,000 in Ubersuggest, which isn't a rung on the ladder.

## What also held: about one keyword in ten is off by more than 5×

9–21%

of keywords diverge between the two tools by more than 5×

Tests 2, 3 and 4 — 5/54, 8/84 and 9/43

511×

largest single divergence, on one misspelt keyword

`stardew valey` — Semrush 720, Ubersuggest 368,000

Not a rung. Multiples.

| Keyword                          | Semrush | Ubersuggest | Factor   |
| -------------------------------- | ------- | ----------- | -------- |
| stardew valey *(misspelt)*       | 720     | 368,000     | **511×** |
| a compound product noun          | 140     | 27,100      | **194×** |
| teraria *(misspelt)*             | 1,900   | 165,000     | 87×      |
| a compound product noun          | 2,900   | 201,000     | 69×      |
| metroid vania *(spaced)*         | 720     | 27,100      | 38×      |
| rim world *(spaced)*             | 6,600   | 60,500      | 9×       |
| baldurs gate 3 *(no apostrophe)* | 49,500  | 246,000     | 5×       |

### The mechanism is visible

Look at what Ubersuggest reports for the *correctly* spelled versions. `stardew valley`: 368,000\. `terraria`: 165,000\. `metroidvania`: 27,100\. `rimworld`: 60,500.

**They're identical.** Ubersuggest quietly normalises the variant to the canonical term and hands you the big keyword's numbers. Semrush doesn't — it reports what the misspelling actually gets, which is 720\. 

Neither behaviour is wrong in itself, and there's a decent argument that normalising is the more useful default. But the consequence is sharp: **if you're deciding whether a misspelling or a spacing variant is worth targeting, Ubersuggest will tell you it has 368,000 searches when it has 720.**

One caveat that matters: **this isn't consistent across languages.** In the German set the same test ran the other way — a common misspelling came back at 70 in Ubersuggest against 5,400 in Semrush. So the honest rule isn't "Ubersuggest inflates variants". It's: **pull both forms of anything with a plausible variant, and if the two tools disagree by more than about double, one of them is answering a different question than you asked.**

## What didn't hold: "Ubersuggest rounds up"

One niche on its own would tell you Ubersuggest is systematically optimistic and you should take a rung off. Four niches say that's local, not general.

|        | Ubersuggest higher | Identical | Lower  | Median ratio |
| ------ | ------------------ | --------- | ------ | ------------ |
| Test 1 | 40 of 47           | 6         | 1      | \~1.2–1.5    |
| Test 2 | **14 of 54**       | 10        | **30** | **0.81**     |
| Test 3 | 73 of 84           | 7         | 4      | 1.23         |
| Test 4 | 37 of 43           | 0         | 6      | 1.50         |

Higher three times, **lower once**. There is no correction factor to apply, in either direction. What you can say is narrower and more honest: the two tools land on different rungs of the same ladder, and which one sits higher depends on the niche.

## Difficulty agreement looked erratic — until the cause turned up

|        | Spearman ρ | Direction                      |
| ------ | ---------- | ------------------------------ |
| Test 1 | +0.446     | Semrush milder (66 of 100)     |
| Test 2 | **−0.234** | Ubersuggest harsher (28 of 32) |
| Test 3 | +0.403     | roughly balanced               |
| Test 4 | **+0.909** | close agreement                |

A **negative** rank correlation means the tools ordered the same keywords roughly inversely. On that evidence the reasonable conclusion is that difficulty scores from two vendors are simply incomparable.

**They aren't. The fourth round found the cause, and it isn't the scoring.**

## The finding that explains most of it: an un-refreshed difficulty score is worthless

In the fourth round the same 50 keywords were exported **twice from each tool, about two minutes apart** — once with the numbers as they sat, once after hitting refresh.

| Field                  | Changed after refresh — Ubersuggest | Changed after refresh — Semrush |
| ---------------------- | ----------------------------------- | ------------------------------- |
| Search volume          | **0 of 50**                         | **0 of 44**                     |
| CPC                    | **0 of 50**                         | **0 of 44**                     |
| Paid Difficulty        | **0 of 50**                         | —                               |
| **Keyword difficulty** | **50 of 50**                        | **42 of 44**                    |

22.7

average point change in Ubersuggest difficulty after a refresh, on a 100-point scale

Maximum observed: 48 points

7.0

the same figure for Semrush

Maximum observed: 29 points

Volume, CPC and paid difficulty are byte-identical. **Only the difficulty moves — and it moves a lot.** Both moved predominantly *upward*, so a cached score systematically **understates** how hard a keyword is.

**Refresh before you decide anything — and don't bother refreshing for volume.** Search volume doesn't change. If you're about to use a difficulty score to choose what to write, an un-refreshed one can be wrong by a quarter of the scale. 

And what makes it a criticism rather than a quirk: **the tool shows a stale 20 in exactly the same way it shows a fresh 68.** There's a "last updated" date beside it, and that date is now confirmed to refer to the difficulty rather than the volume — which had been guesswork before. But nothing about the number itself tells you it's a cached guess.

### And it explains the erratic agreement above

Because both tools were exported in both states, the effect can be isolated. Same keywords, same tools, only the refresh varies:

| Spearman ρ                    | Semrush **not** refreshed | Semrush **refreshed** |
| ----------------------------- | ------------------------- | --------------------- |
| **Ubersuggest not refreshed** | **0.665**                 | 0.765                 |
| **Ubersuggest refreshed**     | 0.823                     | **0.909**             |

Monotonic in both directions. Refreshing either side raises the agreement; refreshing both takes it from **0.665 to 0.909**.

So "the two tools order keywords completely differently" is largely an artefact of comparing stale data against fresh. Refreshed against refreshed, they agree pretty well. The rounds that produced ρ of +0.45, −0.23 and +0.40 were almost certainly comparing one refreshed set against one cached one.

That matters for how you read *any* tool comparison, this one included: **unless the author says both sides were refreshed, a disagreement between two tools might just be a disagreement between two dates.**

### Single-digit difficulty means "never calculated"

| Keyword                          | Un-refreshed | Refreshed |
| -------------------------------- | ------------ | --------- |
| how to get wishlists on steam    | **5**        | 28        |
| steam next fest wishlist spike   | **4**        | 23        |
| steam capsule art size           | **5**        | 43        |
| how many wishlists before launch | **4**        | 34        |
| indie game marketing tools       | **4**        | 36        |

**A single-digit difficulty score doesn't mean "easy". It means the number was never calculated for that keyword.** It stays that way until you spend a refresh on it, and nothing on screen says so. Semrush has a milder version of the same behaviour — it returned 0.0 for one of these.

This is the trap most likely to cost a beginner a month of work, because "difficulty 4, volume 2,900" reads like the find of the week.

## Keyword difficulty is bucketed too, and nobody mentions it

In test 2, 102 keywords produced only **28 distinct** difficulty values — 23 of them on exactly 17, and 17 more on exactly 12, so the two commonest values covered **39% of the set**. In test 3, 125 keywords gave 44 distinct values with 16 sharing 62.

The scale runs 0–100 and looks finely calculated. In practice a large share of any set lands on a handful of numbers, so **within a cluster the score can't rank anything** — the same limitation as the volume buckets, just undocumented.

## Paid Difficulty and CPC measure the niche, not the tool

| Niche                | PD = 1       | PD = 100     | CPC = 0  | Median CPC |
| -------------------- | ------------ | ------------ | -------- | ---------- |
| Gaming               | 44 / 50      | –            | 32 / 50  | 0          |
| Lighting / solar B2B | 50 / 102     | 15           | 57 / 102 | €0.00      |
| Furniture, consumer  | **13 / 125** | **82 / 125** | 24 / 125 | **€0.52**  |

Paid Difficulty sits at 1 in gaming because **nobody is bidding on gaming keywords**. Point the same column at furniture and 82 of 125 come back at 100\. The field is reporting a real fact about the market, and the fact happens to be "there is no paid competition here".

So don't write these columns off after one look at one niche — but do check whether your niche makes them informative. In gaming they'll tell you almost nothing, and **that is itself worth knowing: if nobody is buying ads, organic is the only game.**

## Where Ubersuggest beats Semrush outright: the long tail

One clear advantage only shows up once you test outside English. **Semrush returned no difficulty value at all** for large parts of both German sets — **32 of 64** keywords in test 2 and **19 of 91** in test 3, almost all long-tail product terms and brand names. Ubersuggest returned a number for every one.

**But this didn't replicate in English.** Test 4, with a long-tail block built specifically to check it, produced **zero** empty Semrush values out of 50\. So this is a real advantage in German — and probably in any market where the expensive tool's database is thinner — rather than a general property.

## Two more things that show up once you look

**A lot of real search terms come back as zero.** Across 462 tracked keywords on two commercial sites, **21% and 9%** reported a volume of zero — and among the keywords one of those sites *actually ranks for*, **45% report zero**. Zero means *below the detection threshold*, not *nobody searches this*.

**The twelve-month window isn't the same window for every keyword.** `indie game` returned a history ending April 2026; `indie games` ended July 2026 — three months apart for a singular and a plural pulled in the same session.

## What these tests can't tell you

- **No ground truth.** All four rounds show that the tools disagree, never which one is right. Any sentence of the form "Ubersuggest is less accurate" would be unearned.
- **Three niches is not a sample.** The whole point is that results don't generalise — which applies to these three as well.
- **Rank tracking isn't judged.** In test 1 only five keywords ranked at all and three in both tools. At n = 3, nothing follows.
- **The sets aren't symmetrical.** Ubersuggest's exports were larger than Semrush's (102 vs 64, 125 vs 91), so distribution figures describe each tool's own set rather than a matched pair.
- **Different collection methods.** Test 1 came from position tracking, tests 2 to 4 from bulk exports.

### And one hypothesis that was disproved

Round four included a block of game names that are also ordinary words, on the theory that a matcher would grab the wrong entity. `dredge`, `cocoon`, `hades` and `inscryption` all came back within the normal spread of 1.2 to 1.8×. The 13,500× outlier in the German set — a small brand name matched to a much larger, similarly spelled company — was an individual case rather than a general rule about proper nouns.

The control group itself also wobbled: `game development` returned 823,000 in Semrush against 60,500 in Ubersuggest, a factor of 13.6 and this time with *Semrush* the higher one — in a block that was supposed to validate the method.

## Which one should you actually buy?

That's a different question from "which numbers are right", and it has a different answer depending on who you are.

**Semrush is worth its price if** you manage client sites, need bulk analysis at scale, want intent and SERP features in the same export you build briefs from, or need a backlink index you'd defend in a meeting. It drifts less when data is stale, and rank tracking runs daily across 500–5,000 keywords where Ubersuggest does 125–300 weekly.

**Ubersuggest is worth its price if** you have one or two sites, want the fastest possible route from a question to an answer, and can live with checking your own variants. It costs roughly a quarter as much, and for the decision most small operators actually make — which of four pages to write first — a rough sort order is enough. It also bundles an official MCP connector on a €29 plan where the cheapest Semrush tier with MCP is $139.95.

**And the honest third answer: they're not really substitutes.** For the same seed keyword, only **36 keywords overlapped** between them at 100+ volume. They surface largely different terms — which is why a cheap permanent second opinion has real value next to an expensive primary tool, especially since a lifetime licence never becomes a recurring line item. 

The full buying arguments live in the two reviews: [Ubersuggest](https://seeindie.com/blog/ubersuggest-review/) and [Semrush](https://seeindie.com/blog/semrush-review/).

## What to actually do with all this

1. **Refresh a keyword before you act on its difficulty score**, and never trust a single-digit one. Don't bother refreshing for volume — it doesn't change.
2. **Pull both spellings of anything with a plausible variant.** Roughly one in ten diverges wildly, and the tools normalise variants differently.
3. **Read the paid columns as a signal about your niche**, not about the tool.
4. **Run twenty of your own keywords through whatever you're considering**, next to one other source, before you trust it with a content plan. Other people's experience of a tool doesn't transfer to your niche — that's the one lesson four rounds taught most clearly.

## Sources and method

- Ubersuggest and Semrush bulk exports and position tracking reports, 12–13 August 2026, Google US and Google DE
- Round 4 keyword list built to probe specific failure modes; both tools exported the same day, once un-refreshed and once refreshed
- Analysis: Pearson and Spearman correlations, Google Keyword Planner ladder checks, fold-change distributions
- [Ubersuggest pricing](https://app.neilpatel.com/en/pricing?ref=seeindie.com) and [Semrush pricing](https://www.semrush.com/pricing/seo/?ref=seeindie.com), checked 13 August 2026