Ubersuggest vs Semrush: I ran the same keywords through both, four times
Four head-to-head tests across three niches: the search volumes are identical in origin, one keyword in ten diverges by more than 5x, and an un-refreshed difficulty score is off by an average of 23 points.
Every SEO tool comparison you can find is written from feature lists and pricing pages. This one is written from four head-to-head tests on the same keyword sets, on the same day, in three niches and two countries — plus one accident that turned out to explain most of the disagreement between them.
- The search volumes are identical in origin. Both tools output Google Keyword Planner buckets — 100% of values, four rounds out of four. The cheap one does not invent its numbers.
- Roughly one keyword in ten diverges by more than 5×, clustering on misspellings, spacing and apostrophes, because the two tools normalise variants differently.
- An un-refreshed difficulty score is off by an average of 23 points in Ubersuggest and 7 in Semrush. Refresh changes difficulty and nothing else.
- Refreshed against refreshed, the two tools largely agree — rank correlation rises from 0.665 to 0.909. Most of the "they disagree" story was stale data.
- Single-digit difficulty means "never calculated", not "easy". Nothing on screen tells you which.
What was tested
| Niche | Market | Keywords compared | |
|---|---|---|---|
| Test 1 | Gaming | Google US, English | 47 volumes · 100 difficulty |
| Test 2 | Lighting / LED / solar, B2B | Google DE | 54 volumes · 32 difficulty |
| Test 3 | Furniture / children's beds, consumer | Google DE | 84 volumes · 72 difficulty |
| Test 4 | Gaming, built to probe the failure modes | Google US, English | 43 volumes · 50 difficulty |
Round four was designed rather than collected: fifty keywords in blocks, each block aimed at a failure mode the first three rounds had turned up — compounds against spaced variants, singulars against plurals, misspellings, names that are also ordinary words, long tail, and a control group of head terms both tools would certainly know.
The result: exactly two findings survive all four rounds. Almost everything else reverses direction between niches, which is itself the most useful thing here.
What held everywhere: the volumes are Google's numbers
Both tools output Google Keyword Planner buckets. Keyword Planner doesn't return exact volumes; it returns rungs on a logarithmic ladder stepped by roughly ×1.22: … 880, 1,000, 1,300, 1,600, 1,900, 2,400, 2,900, 3,600, 4,400, 5,400, 6,600, 8,100, 9,900 …
| Semrush on the ladder | Ubersuggest on the ladder | |
|---|---|---|
| Test 1 | 6,082 / 6,082 | 36 / 36 |
| Test 2 | 19 / 19 | 20 / 20 |
| Test 3 | 40 / 40 | 38 / 38 |
| Test 4 | 26 / 26 | 22 / 23 |
Near-perfect across four rounds, three niches, two countries and two languages. Neither tool is modelling its own estimate — both pass the same Google source through. One exception in the whole series: steam came back at 2,740,000 in Ubersuggest, which isn't a rung on the ladder.
What also held: about one keyword in ten is off by more than 5×
stardew valey — Semrush 720, Ubersuggest 368,000Not a rung. Multiples.
| Keyword | Semrush | Ubersuggest | Factor |
|---|---|---|---|
stardew valey (misspelt) | 720 | 368,000 | 511× |
| a compound product noun | 140 | 27,100 | 194× |
teraria (misspelt) | 1,900 | 165,000 | 87× |
| a compound product noun | 2,900 | 201,000 | 69× |
metroid vania (spaced) | 720 | 27,100 | 38× |
rim world (spaced) | 6,600 | 60,500 | 9× |
baldurs gate 3 (no apostrophe) | 49,500 | 246,000 | 5× |
The mechanism is visible
Look at what Ubersuggest reports for the correctly spelled versions. stardew valley: 368,000. terraria: 165,000. metroidvania: 27,100. rimworld: 60,500.
Neither behaviour is wrong in itself, and there's a decent argument that normalising is the more useful default. But the consequence is sharp: if you're deciding whether a misspelling or a spacing variant is worth targeting, Ubersuggest will tell you it has 368,000 searches when it has 720.
One caveat that matters: this isn't consistent across languages. In the German set the same test ran the other way — a common misspelling came back at 70 in Ubersuggest against 5,400 in Semrush. So the honest rule isn't "Ubersuggest inflates variants". It's: pull both forms of anything with a plausible variant, and if the two tools disagree by more than about double, one of them is answering a different question than you asked.
What didn't hold: "Ubersuggest rounds up"
One niche on its own would tell you Ubersuggest is systematically optimistic and you should take a rung off. Four niches say that's local, not general.
| Ubersuggest higher | Identical | Lower | Median ratio | |
|---|---|---|---|---|
| Test 1 | 40 of 47 | 6 | 1 | ~1.2–1.5 |
| Test 2 | 14 of 54 | 10 | 30 | 0.81 |
| Test 3 | 73 of 84 | 7 | 4 | 1.23 |
| Test 4 | 37 of 43 | 0 | 6 | 1.50 |
Higher three times, lower once. There is no correction factor to apply, in either direction. What you can say is narrower and more honest: the two tools land on different rungs of the same ladder, and which one sits higher depends on the niche.
Difficulty agreement looked erratic — until the cause turned up
| Spearman ρ | Direction | |
|---|---|---|
| Test 1 | +0.446 | Semrush milder (66 of 100) |
| Test 2 | −0.234 | Ubersuggest harsher (28 of 32) |
| Test 3 | +0.403 | roughly balanced |
| Test 4 | +0.909 | close agreement |
A negative rank correlation means the tools ordered the same keywords roughly inversely. On that evidence the reasonable conclusion is that difficulty scores from two vendors are simply incomparable.
They aren't. The fourth round found the cause, and it isn't the scoring.
The finding that explains most of it: an un-refreshed difficulty score is worthless
In the fourth round the same 50 keywords were exported twice from each tool, about two minutes apart — once with the numbers as they sat, once after hitting refresh.
| Field | Changed after refresh — Ubersuggest | Changed after refresh — Semrush |
|---|---|---|
| Search volume | 0 of 50 | 0 of 44 |
| CPC | 0 of 50 | 0 of 44 |
| Paid Difficulty | 0 of 50 | — |
| Keyword difficulty | 50 of 50 | 42 of 44 |
Volume, CPC and paid difficulty are byte-identical. Only the difficulty moves — and it moves a lot. Both moved predominantly upward, so a cached score systematically understates how hard a keyword is.
And what makes it a criticism rather than a quirk: the tool shows a stale 20 in exactly the same way it shows a fresh 68. There's a "last updated" date beside it, and that date is now confirmed to refer to the difficulty rather than the volume — which had been guesswork before. But nothing about the number itself tells you it's a cached guess.
And it explains the erratic agreement above
Because both tools were exported in both states, the effect can be isolated. Same keywords, same tools, only the refresh varies:
| Spearman ρ | Semrush not refreshed | Semrush refreshed |
|---|---|---|
| Ubersuggest not refreshed | 0.665 | 0.765 |
| Ubersuggest refreshed | 0.823 | 0.909 |
Monotonic in both directions. Refreshing either side raises the agreement; refreshing both takes it from 0.665 to 0.909.
So "the two tools order keywords completely differently" is largely an artefact of comparing stale data against fresh. Refreshed against refreshed, they agree pretty well. The rounds that produced ρ of +0.45, −0.23 and +0.40 were almost certainly comparing one refreshed set against one cached one.
That matters for how you read any tool comparison, this one included: unless the author says both sides were refreshed, a disagreement between two tools might just be a disagreement between two dates.
Single-digit difficulty means "never calculated"
| Keyword | Un-refreshed | Refreshed |
|---|---|---|
| how to get wishlists on steam | 5 | 28 |
| steam next fest wishlist spike | 4 | 23 |
| steam capsule art size | 5 | 43 |
| how many wishlists before launch | 4 | 34 |
| indie game marketing tools | 4 | 36 |
A single-digit difficulty score doesn't mean "easy". It means the number was never calculated for that keyword. It stays that way until you spend a refresh on it, and nothing on screen says so. Semrush has a milder version of the same behaviour — it returned 0.0 for one of these.
This is the trap most likely to cost a beginner a month of work, because "difficulty 4, volume 2,900" reads like the find of the week.
Keyword difficulty is bucketed too, and nobody mentions it
In test 2, 102 keywords produced only 28 distinct difficulty values — 23 of them on exactly 17, and 17 more on exactly 12, so the two commonest values covered 39% of the set. In test 3, 125 keywords gave 44 distinct values with 16 sharing 62.
The scale runs 0–100 and looks finely calculated. In practice a large share of any set lands on a handful of numbers, so within a cluster the score can't rank anything — the same limitation as the volume buckets, just undocumented.
Paid Difficulty and CPC measure the niche, not the tool
| Niche | PD = 1 | PD = 100 | CPC = 0 | Median CPC |
|---|---|---|---|---|
| Gaming | 44 / 50 | – | 32 / 50 | 0 |
| Lighting / solar B2B | 50 / 102 | 15 | 57 / 102 | €0.00 |
| Furniture, consumer | 13 / 125 | 82 / 125 | 24 / 125 | €0.52 |
Paid Difficulty sits at 1 in gaming because nobody is bidding on gaming keywords. Point the same column at furniture and 82 of 125 come back at 100. The field is reporting a real fact about the market, and the fact happens to be "there is no paid competition here".
So don't write these columns off after one look at one niche — but do check whether your niche makes them informative. In gaming they'll tell you almost nothing, and that is itself worth knowing: if nobody is buying ads, organic is the only game.
Where Ubersuggest beats Semrush outright: the long tail
One clear advantage only shows up once you test outside English. Semrush returned no difficulty value at all for large parts of both German sets — 32 of 64 keywords in test 2 and 19 of 91 in test 3, almost all long-tail product terms and brand names. Ubersuggest returned a number for every one.
But this didn't replicate in English. Test 4, with a long-tail block built specifically to check it, produced zero empty Semrush values out of 50. So this is a real advantage in German — and probably in any market where the expensive tool's database is thinner — rather than a general property.
Two more things that show up once you look
A lot of real search terms come back as zero. Across 462 tracked keywords on two commercial sites, 21% and 9% reported a volume of zero — and among the keywords one of those sites actually ranks for, 45% report zero. Zero means below the detection threshold, not nobody searches this.
The twelve-month window isn't the same window for every keyword. indie game returned a history ending April 2026; indie games ended July 2026 — three months apart for a singular and a plural pulled in the same session.
What these tests can't tell you
- No ground truth. All four rounds show that the tools disagree, never which one is right. Any sentence of the form "Ubersuggest is less accurate" would be unearned.
- Three niches is not a sample. The whole point is that results don't generalise — which applies to these three as well.
- Rank tracking isn't judged. In test 1 only five keywords ranked at all and three in both tools. At n = 3, nothing follows.
- The sets aren't symmetrical. Ubersuggest's exports were larger than Semrush's (102 vs 64, 125 vs 91), so distribution figures describe each tool's own set rather than a matched pair.
- Different collection methods. Test 1 came from position tracking, tests 2 to 4 from bulk exports.
And one hypothesis that was disproved
Round four included a block of game names that are also ordinary words, on the theory that a matcher would grab the wrong entity. dredge, cocoon, hades and inscryption all came back within the normal spread of 1.2 to 1.8×. The 13,500× outlier in the German set — a small brand name matched to a much larger, similarly spelled company — was an individual case rather than a general rule about proper nouns.
The control group itself also wobbled: game development returned 823,000 in Semrush against 60,500 in Ubersuggest, a factor of 13.6 and this time with Semrush the higher one — in a block that was supposed to validate the method.
Which one should you actually buy?
That's a different question from "which numbers are right", and it has a different answer depending on who you are.
Semrush is worth its price if you manage client sites, need bulk analysis at scale, want intent and SERP features in the same export you build briefs from, or need a backlink index you'd defend in a meeting. It drifts less when data is stale, and rank tracking runs daily across 500–5,000 keywords where Ubersuggest does 125–300 weekly.
Ubersuggest is worth its price if you have one or two sites, want the fastest possible route from a question to an answer, and can live with checking your own variants. It costs roughly a quarter as much, and for the decision most small operators actually make — which of four pages to write first — a rough sort order is enough. It also bundles an official MCP connector on a €29 plan where the cheapest Semrush tier with MCP is $139.95.
The full buying arguments live in the two reviews: Ubersuggest and Semrush.
What to actually do with all this
- Refresh a keyword before you act on its difficulty score, and never trust a single-digit one. Don't bother refreshing for volume — it doesn't change.
- Pull both spellings of anything with a plausible variant. Roughly one in ten diverges wildly, and the tools normalise variants differently.
- Read the paid columns as a signal about your niche, not about the tool.
- Run twenty of your own keywords through whatever you're considering, next to one other source, before you trust it with a content plan. Other people's experience of a tool doesn't transfer to your niche — that's the one lesson four rounds taught most clearly.
Sources and method
- Ubersuggest and Semrush bulk exports and position tracking reports, 12–13 August 2026, Google US and Google DE
- Round 4 keyword list built to probe specific failure modes; both tools exported the same day, once un-refreshed and once refreshed
- Analysis: Pearson and Spearman correlations, Google Keyword Planner ladder checks, fold-change distributions
- Ubersuggest pricing and Semrush pricing, checked 13 August 2026