Retrieval Engines, LLM Reasoning-Based Search, and the Spectrum Between Them
Post mostly written by Claude - but it's a good relevant topic for those building agentic systems.
One of the most common questions I get from teams building AI-powered knowledge systems is deceptively simple: Why do I need a retrieval step at all? Can't I just pass all my documents into an LLM and let it figure out what's relevant?
It's a fair question — and increasingly, it's the right instinct. But the full answer reveals a spectrum of approaches, each with distinct mechanics, tradeoffs, and sweet spots. Let me break down what's actually happening under the hood when a retrieval engine finds content versus when an LLM does, and why the best systems combine both.
What a Retrieval Engine Actually Does
When we talk about "retrieval" in the classical sense — systems built on BERT-style encoder models, BM25 indexes, or dense vector search — we're talking about a very specific operation. The system takes your query, converts it into a representation (a sparse term vector or a dense embedding), and then performs a similarity computation against a pre-indexed corpus of document chunks.
Here's what matters about this process at a granular level. The retrieval engine never "reads" your documents at query time. It read them at index time, when each chunk was encoded into a fixed-dimensional vector. At query time, it's performing math — dot products, cosine similarity, approximate nearest neighbor search — across those pre-computed representations. This is why retrieval is fast. A well-tuned vector search over millions of chunks can return results in tens of milliseconds. The computation scales sub-linearly with corpus size thanks to algorithms like HNSW or IVF indexing.
But there's a fundamental constraint baked into this architecture: the retrieval model can only match on patterns it learned during training. A BERT-based bi-encoder compresses an entire chunk into a single vector, typically 768 or 1024 dimensions. That vector is a lossy compression of the chunk's meaning. It captures semantic similarity well — "cardiac arrest" will match "heart attack" — but it struggles with reasoning. It cannot infer that a paragraph about "Q3 revenue declining 12% in the EMEA region" is relevant to a question about "which geographies are underperforming," because that inference requires understanding the relationship between revenue decline and underperformance, not just semantic overlap.
Retrieval also operates at a fixed granularity. You chose a chunk size at index time — maybe 512 tokens, maybe a page, maybe a section. That decision is frozen. If the answer to a question spans two chunks, or lives in a single sentence buried in a 2,000-token chunk, the retrieval engine doesn't adapt. It returns what it indexed.
What an LLM Actually Does
When an LLM processes a document to find relevant content, the mechanics are entirely different. The model reads the content token by token through its attention mechanism, building a rich contextual representation of every passage in relation to every other passage and the query. It's not performing similarity search — it's performing comprehension.
This means the LLM can do things a retrieval engine fundamentally cannot. It can follow chains of reasoning: "this paragraph defines the policy, that paragraph describes an exception, and together they answer the question." It can understand document structure — navigating a table of contents, recognizing that a section header signals a topic boundary, inferring that an appendix contains supporting detail for a claim made in the executive summary. It can resolve ambiguity: when you ask "what's our position on X," the LLM can distinguish between a paragraph that mentions X in passing and one that constitutes an actual position statement.
Perhaps most powerfully, the LLM can evaluate relevance at arbitrary granularity. It doesn't need pre-defined chunks. It can identify that the answer is a single sentence, or that it spans three pages, or that it requires synthesizing information from multiple sections. It adapts to the question.
But this comes at a cost, and the cost is straightforward: the LLM must process every token. There is no index, no pre-computation, no shortcut. If you have 100,000 tokens of content, the model must attend to all of them. This is an O(n) operation on input length (and with standard attention, O(n²) in the attention computation itself). For a 50-page document, this might take a few seconds and cost a few cents. For 10,000 documents, it's simply not viable — you'd be looking at minutes of latency and dollars per query.
The Real Differences, Side by Side
The distinction becomes clearest when you look at specific dimensions.
Latency is where retrieval engines dominate. Vector search returns results in single-digit milliseconds. An LLM processing even a modest 50,000-token context will take multiple seconds. At production scale with concurrent users, this difference is the difference between a snappy product and an unusable one.
Accuracy on straightforward lookups — "What is the effective date of policy X?" — tends to be comparable. Both approaches find the right chunk when the query terms and document terms overlap well. Retrieval may even edge ahead here because it's been optimized specifically for this pattern and is less likely to hallucinate.
Accuracy on complex questions — questions requiring inference, multi-hop reasoning, synthesis across sections, or understanding of document structure — is where LLMs pull ahead significantly. A retrieval engine will return the top-k chunks by similarity, but it has no mechanism to reason about whether those chunks actually answer the question. The LLM does.
Scalability tells the inverse story. Retrieval engines were designed to search millions of documents. An LLM's context window, even at 128K or 200K tokens, holds maybe a few hundred pages. You simply cannot brute-force your way through a large corpus with an LLM alone.
Cost follows a similar pattern. A retrieval query is computationally cheap — it's a vector lookup. An LLM query prices by input tokens. Processing large volumes of content through an LLM context window gets expensive quickly, especially at scale.
The Spectrum, and Why Hybrid Wins
In practice, the choice is not binary — it's a spectrum with three zones.
At one end, you have pure retrieval: fast, cheap, scalable, but limited to pattern matching and fixed-granularity chunks. This works well when your queries are predictable, your documents are well-structured, and the vocabulary between questions and answers overlaps naturally.
At the other end, you have pure LLM reasoning over full documents: highly accurate, flexible, capable of genuine comprehension, but slow and expensive at scale. This works well when you have a small corpus (say, a handful of documents per query), complex questions, and the budget and latency tolerance to support it.
In the middle — and this is where the most effective systems live — you have retrieval-augmented reasoning. The retrieval engine acts as a first-pass filter, narrowing millions of documents down to a manageable set. Then the LLM reads and reasons over that narrowed set. The retrieval step handles scale; the LLM handles comprehension.
But I'd argue there's an increasingly important variant that doesn't get enough attention: LLM-guided retrieval. In this pattern, the LLM doesn't just consume retrieval results — it actively directs the retrieval process. It reads corpus metadata, directory structures, and document titles. It formulates and reformulates queries. It decides what to retrieve next based on what it's already seen. Think of it as the LLM acting as the "brain" of the search process, with the retrieval engine as its "hands."
This pattern is powerful because it addresses one of classical retrieval's biggest blind spots: query formulation. Users rarely phrase their questions in the same language the documents use. An LLM can bridge that gap, generating multiple query reformulations, expanding abbreviations, reasoning about what kind of document is likely to contain the answer, and iterating based on initial results.
So What Should You Build?
The answer depends on your constraints.
If you need sub-second latency at scale across a large corpus, you need a retrieval engine as your foundation. There is no LLM-only architecture that will give you millisecond response times over millions of documents.
If you need high accuracy on complex, reasoning-heavy questions over a small document set, LLM-first approaches can be simpler and more effective than building a full retrieval pipeline. At small scale, the chunking, indexing, and embedding decisions you'd need to make for a retrieval system introduce more complexity than they solve.
For most production systems, you want the hybrid: retrieval for scale and speed, LLM reasoning for comprehension and accuracy, with the LLM increasingly involved in orchestrating the retrieval process itself. The frontier of this field isn't about choosing between retrieval and reasoning — it's about finding the right balance point on the spectrum for your specific use case, corpus size, latency requirements, and accuracy expectations.
The systems that win are the ones that know when to search and when to read.