Context windows now reach a million tokens. That's roughly a few thousand pages, enough to hold a company handbook, a year of contracts or a mid-sized codebase in a single prompt. So a fair question keeps coming up in 2026: do you still need RAG (retrieval-augmented generation), with its chunking, embeddings and vector databases? Or can you just put everything in the prompt?
The honest answer is "it depends", but on things you can measure: how big your knowledge base is, how often it changes, how many questions you ask per minute, and how much a wrong answer costs. This guide explains what the research says about long contexts, where retrieval still wins, the cost math with prompt caching, and a tested Python demo that shows one retrieval failure and how a simple fix solves it.
The two approaches in one paragraph each
Long context means sending the whole source material with every question. The model sees everything, so nothing is lost to bad retrieval. You pay for every token you send, and quality can drop as the input grows.
RAG means splitting your documents into chunks, indexing them, and at question time sending only the few chunks most likely to contain the answer. It's cheap per question and scales to any size, but if the retriever picks the wrong chunks, the model never sees the answer at all.
What the research says about very long inputs
A bigger window is not the same as reliable use of that window. Three findings are worth knowing:
- Lost in the middle. Liu and colleagues (Stanford and partners, published in TACL) found that models used information best when it sat at the beginning or end of the input, and performance dropped significantly when the relevant passage was in the middle of a long context.
- Advertised versus effective length. NVIDIA's RULER benchmark (2024) tested ten long-context models that all claimed 32,000 tokens or more. Only four kept satisfactory performance at 32,000 tokens, even though nearly all of them passed the simple "needle in a haystack" test.
- Context rot. Chroma's July 2025 study of 18 models, including GPT-4.1, Claude 4 and Gemini 2.5, found that performance became increasingly unreliable as input length grew, even on simple tasks. Irrelevant material hurt: models did noticeably better on a focused excerpt than on the full conversation history containing the same answer.
Newer models handle long inputs better than the ones in these studies, and providers keep improving. But the direction has held across every generation tested so far: more irrelevant text makes answers less reliable. Long context is a capability to use deliberately, not a reason to stop curating what goes into the prompt.
What the research says about RAG versus long context
A 2024 study by Google DeepMind and University of Michigan researchers (Li et al., "Retrieval Augmented Generation or Long-Context LLMs?") compared the two directly. When the models had enough resources, long context outperformed RAG on average answer quality. But RAG was far cheaper. Their hybrid, Self-Route, lets the model first try with retrieved chunks and decide whether it can answer; only the questions it can't answer go to the full long context. That kept quality close to long context at a much lower cost.
The practical lesson is that this isn't a contest with one winner. The best systems in 2026 combine both.
Retrieval failures are fixable: contextual chunks
The classic weakness of RAG is chunks that lose their meaning when cut out of the document. A chunk saying "The limit is 30 days from the purchase date" doesn't mention refunds, so a question about refunds may not find it. Anthropic's "Contextual Retrieval" write-up (September 2024) tackled exactly this by adding a short, document-specific description to each chunk before indexing it. In their tests, measured as the share of questions where the right chunk wasn't in the top 20:
- Contextual embeddings alone cut retrieval failures by 35% (5.7% to 3.7%).
- Adding contextual BM25 (keyword search) cut them by 49% (to 2.9%).
- Adding a reranking step cut them by 67% (to 1.9%).
The same write-up gives a useful rule of thumb: if your knowledge base is under about 200,000 tokens (roughly 500 pages), you can often skip RAG entirely, include the whole thing in the prompt, and use prompt caching to keep it cheap.
A tested demo: one failure, one fix, and the cost math
The script below needs no libraries or API key. It does three things: runs a small BM25 keyword search over two policy documents, with and without a contextual header on each chunk; compares the cost per question of long context (with caching) and RAG; and applies a simple routing rule. Prices are illustrative ($3 per million input tokens, with Anthropic's published cache multipliers: 1.25x to write a 5-minute cache, 0.1x to read it on most models). Swap in your model's real prices.
"""RAG vs long context: a small, dependency-free demo.
1) BM25 retrieval over chunks, with and without a contextual header.
2) Cost per query: long context (with prompt caching) vs RAG.
3) A simple router that picks a strategy.
"""
import math, re
from collections import Counter
# --- 1. Toy knowledge base: two policy documents split into chunks ---------
DOCS = {
"Refund policy": [
"Customers can request their money back for annual plans.",
"The limit is 30 days from the purchase date.",
"After that, credit is offered instead of cash.",
],
"Shipping policy": [
"Orders inside the UAE arrive in 2 to 4 working days.",
"The limit is 5 kg per parcel for standard delivery.",
"Express delivery is available in Dubai and Abu Dhabi.",
],
}
def tokenize(text):
return re.findall(r"[a-z0-9]+", text.lower())
class BM25:
def __init__(self, docs, k1=1.5, b=0.75):
self.docs = [tokenize(d) for d in docs]
self.n = len(self.docs)
self.avgdl = sum(map(len, self.docs)) / self.n
self.df = Counter(t for d in self.docs for t in set(d))
self.k1, self.b = k1, b
def score(self, query, i):
d, tf, s = self.docs[i], Counter(self.docs[i]), 0.0
for t in tokenize(query):
if t not in tf:
continue
idf = math.log(1 + (self.n - self.df[t] + 0.5) / (self.df[t] + 0.5))
s += idf * tf[t] * (self.k1 + 1) / (tf[t] + self.k1 * (1 - self.b + self.b * len(d) / self.avgdl))
return s
def top(self, query, k=1):
return sorted(range(self.n), key=lambda i: -self.score(query, i))[:k]
plain, contextual = [], []
for title, chunks in DOCS.items():
for c in chunks:
plain.append(c)
contextual.append(f"[{title}] {c}") # contextual header prepended
query = "What is the time limit for a refund?"
for name, chunks in (("plain chunks", plain), ("contextual chunks", contextual)):
idx = BM25(chunks).top(query, k=1)[0]
print(f"{name:18} -> {chunks[idx]}")
# --- 2. Cost per query (illustrative prices, USD per million input tokens) --
PRICE = 3.00 # base input price, illustrative
WRITE_5M = 1.25 # 5-minute cache write multiplier
READ = 0.10 # cache read multiplier
def lc_cost(corpus_tokens, queries_per_window):
"""Whole corpus in the prompt, cached: one write, then reads."""
write = corpus_tokens * PRICE * WRITE_5M
reads = corpus_tokens * PRICE * READ * (queries_per_window - 1)
return (write + reads) / queries_per_window / 1e6
def rag_cost(chunk_tokens=500, k=8):
"""Only the top-k retrieved chunks go into the prompt."""
return chunk_tokens * k * PRICE / 1e6
print()
print(f"{'corpus':>10} {'queries/5min':>13} {'long ctx $/q':>13} {'RAG $/q':>9}")
for corpus in (50_000, 150_000, 800_000):
for q in (1, 20):
print(f"{corpus:>10,} {q:>13} {lc_cost(corpus, q):>13.4f} {rag_cost():>9.4f}")
# --- 3. A simple router ------------------------------------------------------
def choose(corpus_tokens, fits_window=1_000_000, changes_daily=False):
if corpus_tokens > fits_window:
return "RAG (corpus does not fit)"
if corpus_tokens <= 200_000 and not changes_daily:
return "long context + prompt caching"
return "RAG, or hybrid: retrieve, then send whole matching documents"
print()
for c, ch in ((80_000, False), (150_000, True), (600_000, False), (5_000_000, False)):
print(f"{c:>9,} tokens, changes daily={ch!s:5} -> {choose(c, changes_daily=ch)}")
Output, run on 2026-10-07:
plain chunks -> The limit is 5 kg per parcel for standard delivery.
contextual chunks -> [Refund policy] The limit is 30 days from the purchase date.
corpus queries/5min long ctx $/q RAG $/q
50,000 1 0.1875 0.0120
50,000 20 0.0236 0.0120
150,000 1 0.5625 0.0120
150,000 20 0.0709 0.0120
800,000 1 3.0000 0.0120
800,000 20 0.3780 0.0120
80,000 tokens, changes daily=False -> long context + prompt caching
150,000 tokens, changes daily=True -> RAG, or hybrid: retrieve, then send whole matching documents
600,000 tokens, changes daily=False -> RAG, or hybrid: retrieve, then send whole matching documents
5,000,000 tokens, changes daily=False -> RAG (corpus does not fit)
Three things to notice:
- The retrieval failure is real and silent. Asked about the refund time limit, plain chunks returned the shipping weight limit, a confident wrong answer waiting to happen. Adding the document title to each chunk fixed it. Nothing about the model changed.
- Caching changes the long-context math, but only with traffic. A 150,000-token corpus costs about $0.56 per question if you ask once, and about $0.07 if 20 questions share the same cache window. The cache only lasts minutes, so steady traffic matters.
- RAG stays cheapest per question (about $0.012 for eight 500-token chunks here) and is the only option once the corpus outgrows the window.
Output tokens cost the same either way and aren't included. (For the full picture on agent bills, see how to cut AI agent costs.)
When to use which
Use long context (plus prompt caching) when:
- The material is small to medium, roughly under 200,000 tokens.
- It changes rarely, so the cached prefix stays valid.
- Questions need connections across the whole document: "what changed between these two contracts?", "summarize every risk mentioned anywhere".
- Getting it wrong is expensive and you'd rather pay more per question than miss a passage.
Use RAG when:
- The knowledge base is large, growing, or bigger than any context window.
- Content changes often (a product catalog, support tickets, prices).
- You handle many questions per minute and cost per question matters.
- Different users may only see certain documents. Retrieval lets you filter by permission before anything reaches the model, which is much safer than sending everything and hoping the model doesn't mention it.
- You need to show sources. Retrieved chunks make citations natural.
Use a hybrid when you want both: retrieve to find the right documents, then send those documents whole instead of tiny chunks. This is often the best default for business knowledge bases, because it keeps retrieval's cost and permission control while giving the model enough context to reason.
Six rules that hold either way
- Put the question after the material, and the most important documents near the start or end, not buried in the middle.
- Ask the model to quote the passage it used before answering. It's easy to check, and it reduces made-up answers.
- Combine keyword and semantic search. Product codes, names and numbers are often found by keywords, not embeddings.
- Add context to chunks: at minimum the document title and section heading, as the demo shows.
- Cache stable content and keep anything that changes (dates, user details) at the end of the prompt, after the cached part.
- Measure retrieval separately from answers. Build a small test set of real questions with the passage that answers each one, and check how often the right passage is retrieved. (Our guide to testing AI agents with evals shows how.)
A note on Arabic content
Keyword search in Arabic is harder than in English: prefixes like "ال" and "و", attached pronouns, and different spellings of the same word (أ/ا/إ, ة/ه, ى/ي) all reduce matches. If your documents are in Arabic, normalize letters before indexing, use an embedding model tested on Arabic, and include both Arabic and English terms for technical words. Then test with real questions from your users, in the dialect they actually write. (More in our guide to AI in Arabic.)
The takeaway
Million-token windows didn't kill RAG; they changed where the line is. For a few hundred pages that rarely change, put it all in the prompt and cache it. For anything larger, faster-changing, permission-sensitive or high-volume, retrieve first, then give the model whole relevant documents. And in both cases, measure: a retrieval failure looks exactly like a confident answer until you check.
Sources: Liu et al., "Lost in the Middle: How Language Models Use Long Contexts," Transactions of the ACL (2024); Hsieh et al., "RULER: What's the Real Context Size of Your Long-Context Language Models?" (NVIDIA, 2024); Chroma, "Context Rot: How Increasing Input Tokens Impacts LLM Performance" (July 2025); Li et al., "Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach," EMNLP Industry Track (2024); Anthropic, "Introducing Contextual Retrieval" (September 2024); Anthropic prompt caching documentation (checked October 2026). The demo uses illustrative prices.