← All Writing
September 22, 20267 min read

How I Added Hybrid Search to My Site (Keyword + Semantic, Merged)

My search bar used to run keyword search and only fall back to semantic search when that found nothing. Now it runs both on every query and merges them into one ranked list.

YieldA search bar that ranks results by keyword match and meaning at the same time, merged with Reciprocal Rank Fusion in a single database function
DifficultyIntermediate (Postgres full-text search, one SQL function, and a before/after test on a copy of the real search data)
Total Cook TimeOne session to build, most of it spent testing — which is also how the one real bug got caught

Ingredients

Where Search Stood Before

In April I added semantic search to the site’s search bar. That post covers the difference between keyword search and semantic search, so here’s the short version: keyword search finds the words you typed, and semantic search finds what you meant. Keyword search is great with names (“Schwab,” “cron”) and blind to meaning. Semantic search understands that a schnauzer is a dog, but it’s shaky on names and acronyms it never learned.

I didn’t pick one. I stacked them: keyword matching ran first, instantly, in the browser, and only if it found nothing did the site ask for semantic results. That worked well enough that I stopped thinking about it, until I looked closely at what the keyword step was actually doing.

Why “One, Then the Other” Wasn’t Enough

How Hybrid Search Works: Rank Twice, Then Merge

Hybrid search runs both searches on every query and merges the two ranked lists. The tricky part is the merge. Keyword scores come out on one scale and semantic similarity on another. You can’t just add them; it would be like adding a GPA to an SAT score.

The standard fix is Reciprocal Rank Fusion (RRF), and it’s almost insultingly simple: ignore the scores and use only each result’s position. A result at position r on a list earns 1 / (60 + r). Add up what each result earned from both lists, sort, done. The 60 is the convention from the original RRF paper; it keeps the gap between #1 and #2 from being enormous.

query: “oauth”
# keyword rank / semantic rank → RRF score

Schwab OAuth post    #1 / #2  → 1/61 + 1/62 = 0.0325
Gmail MCP post      #2 / #3  → 1/62 + 1/63 = 0.0320
Push alerts post    —  / #1  →   0  + 1/61 = 0.0164

Semantic search alone put the push alerts post first for “oauth.” The Schwab post ranked near the top on both lists, so it wins the merge.

That’s the whole idea: results that both methods agree on rise to the top, and a result that only one method finds still gets a vote.

🔧 Developer section: the Postgres side

Testing It Before Touching the Real Database

I didn’t want to learn whether this worked by shipping it. So I copied the site’s real search data, every page with its real embedding, into PGlite, a full Postgres that runs inside a Node script. Same SQL, same pgvector, same embedding model for the queries. Then I ran 27 queries through three setups: the old cascade, semantic only, and hybrid.

The test caught a real bug on its first run. “firewall” scored zero on the keyword side even though a post’s description contained the word. The cause: the query was reduced to its root word (“firewall” → “firewal”), then handed to a function that reduced it again (“firewal” → “firew”), which matched nothing. A one-line fix I would never have spotted by eye.

QueryOld cascade — top resultHybrid — top result
“rag”Gmail client post (for the “rag” in “storage”)The Ask Goose RAG chatbot post
“oauth”Gmail MCP postSchwab OAuth post
“self-hosted”Gmail MCP post (first page containing the phrase)The push alerts post (second on both lists)
“parking”The 3D model follow-upThe original parking tracker post
“ai”Ask Goose, then posts matching the “ai” inside words like “email”AI content pipeline, Ask Goose, TL;DR AI summaries
“the”Six unrelated pagesOne weak semantic match (About)

Of 27 test queries, the top result changed on 7. On the other 20, hybrid agreed with the old search, which is what you want from a change like this.

Seven out of 27 isn’t dramatic, and I think that’s the honest result. On a small site, most queries have one obvious answer and any reasonable search finds it. Hybrid search earns its keep on the edges: short acronyms, names the embedding model doesn’t know, and queries where the old search found something and so never looked for the better thing.

What Hybrid Search Didn’t Fix

It isn’t always better. For “linux server security,” semantic search alone put my home-server security guide first. Hybrid put the headless Linux setup post first, because “Linux” and “server” are in its title, and the security guide second. Both are reasonable answers, but it’s a real case where the keyword vote pulled a slightly worse result up.

“pm” still returns nothing. Neither does “llm,” on the website of a product manager who writes mostly about LLMs. Keyword search can’t find “pm” because my pages say “product manager,” and semantic search can’t connect two letters to that phrase. Merging two empty lists still gives you an empty list. The fix there isn’t a better algorithm; it’s a small synonym list, or descriptions that use the words people actually type.

One shared word gets a full vote. Because RRF only looks at position, a post that matches a single query word can ride that into the results. “game theory” now shows my cron automation post third, because its description mentions game bots. It sits below both Numerator results, but it’s the trade-off: the same property that rescues “oauth” also lets in the occasional stray.

Test on a copy of your real data

Before shipping a search change, run a fixed list of queries against a copy of your real data, old versus new. It turns “feels better” into “changed 7 of 27, here’s which ones,” and in my case it caught a bug that would have quietly broken keyword search for any word that gets shortened twice.

If you’re choosing between the two

Don’t. If your content has names, product terms or acronyms (almost all content does), semantic search alone will miss some of them. If your readers describe things in their own words, keyword search alone will miss those. Postgres can do both in one query, and the merge is a single line of arithmetic.

← Back to all writing