LLM Integration for RAG: Context Windows, Token Budgets, and Tool Use
The engineering layer between your retrieval pipeline and the model itself — context windows, token budgeting, streaming, structured output, function calling, and what long-context models actually change.
Introduction
Prompt Engineering covered how to prompt an LLM well once retrieved context is ready to send. This post covers the layer just underneath that: the practical mechanics of actually integrating an LLM into a production RAG system — how much you can send it, how to manage that budget, how responses come back, and how a model can reach beyond generating text at all. None of this is retrieval-specific, exactly, but every RAG system runs into all of it the moment it leaves a notebook and becomes a real application.
Context Windows
A model’s context window is the maximum amount of text — measured in tokens, not characters or words — it can process in a single call, spanning the system prompt, injected retrieved context, conversation history, and the model’s own generated output combined, all counted against the same shared limit.
This is a hard, non-negotiable constraint, and it directly shapes decisions made in earlier post: it’s part of why chunk size and top-k selection matter — retrieve too much context and you can exceed the window entirely, or crowd out room for a useful answer even if you technically fit. It’s worth being precise about a common point of confusion: a context window isn’t measured in words — a “token” is roughly three-quarters of an English word on average, though this varies by language and content type (code and non-English text often tokenize less efficiently, consuming more tokens per unit of apparent content). Knowing your actual token count for a given piece of text, rather than estimating from word count, matters enough that production systems typically compute it precisely using the model provider’s own tokenizer rather than guessing.
Token Budgeting
Token budgeting is the discipline of deliberately allocating a model’s finite context window across its competing consumers, rather than injecting retrieved context until you run out of room by accident.
A typical RAG call has to fit several things into one shared budget: the system prompt and instructions, the retrieved context chunks, any conversation history if it’s a multi-turn interaction, the user’s current question, and headroom reserved for the model’s own response — a long, cut-off answer because the window filled up before the model finished is a real, avoidable failure mode. A deliberate token budget allocates a rough share to each of these upfront (for instance, capping retrieved context to some fixed portion of the total window, leaving explicit reserved space for the response) rather than treating context injection as open-ended.
This connects directly back to the top-k and threshold retrieval decisions: those aren’t purely relevance-quality decisions, they’re also token-budget decisions — the same retrieval quality question (“how many chunks are genuinely useful”) has to be balanced against a hard capacity constraint that has nothing to do with relevance at all. In multi-turn conversations specifically, token budgeting also has to account for accumulating history — a long conversation can eventually crowd out room for new retrieved context entirely unless older turns are deliberately summarized, truncated, or dropped as the conversation grows.
Streaming
Streaming is the practice of returning a model’s output incrementally, token by token (or in small chunks), as it’s generated, rather than waiting for the entire response to finish before showing anything to the user.
The motivation is almost entirely about perceived responsiveness rather than actual total speed: a RAG answer might take several seconds to fully generate, and streaming means a user sees the first words appear almost immediately, which substantially improves the experience of using the system even though the total time to the last token is roughly the same either way. There’s a genuine RAG-specific complication worth knowing, though: if you’re doing citation prompting or want to validate a response’s faithfulness before showing it, streaming raw tokens directly to the user makes that harder, since you’re displaying content before you’ve had a chance to check it — some production systems address this by streaming the answer while running faithfulness checks in parallel, then flagging or retracting the response after the fact if a check fails, rather than blocking the stream entirely and losing the responsiveness benefit.
Structured Output
Structured output means constraining a model’s response to a specific, machine-parseable format — commonly JSON matching a defined schema — rather than free-form prose, so the response can be reliably consumed by downstream code without fragile ad hoc text parsing.
This matters throughout a RAG pipeline beyond just the final user-facing answer: query rewriting and decomposition, the LLM-generated hypothetical answer in HyDE, or the contextual summary generated per chunk during contextual retrieval all benefit from constrained, structured output rather than needing to parse free-form text reliably. Modern LLM APIs increasingly support this natively — enforcing that output conforms to a provided JSON schema at the generation level itself, rather than merely instructing the model to “please respond in JSON” and hoping it complies — which is a meaningfully more reliable guarantee, since prompt-only instructions can still occasionally produce malformed or inconsistent output that breaks a parser.
Function Calling and Tool Use
Function calling is the mechanism that lets a model, instead of only generating text, request that a specific defined function be executed — describing which function to call and with what arguments — with the actual execution and its result handed back to the model to incorporate into its next response. Tool use is the broader pattern this enables: giving a model access to capabilities beyond pure text generation, of which a retrieval call is itself one clear example.
This is directly relevant to how a RAG system’s retrieval step actually gets triggered in more advanced architectures: rather than always retrieving before every generation (the fixed pipeline structure assumed throughout most of this series), a model can be given retrieval itself as a callable tool, and decide — based on the specific question — whether retrieval is even necessary for this particular query, what to search for, and whether to search again if the first attempt didn’t return enough to work with. This is the architectural seed of agentic RAG: a system where retrieval isn’t a fixed, single-shot preprocessing step but a tool the model actively and repeatedly reaches for as part of a broader reasoning process, deciding for itself when and how many times to invoke it — closely related to, and often implemented on top of, the multi-hop and recursive retrieval patterns.
Long Context Models
Context windows have grown substantially — from a few thousand tokens in early transformer-based models to context windows spanning hundreds of thousands, sometimes over a million tokens, in current-generation models. This shift is worth understanding precisely, because it’s tempting to read it as making much of this series’ earlier retrieval-precision work obsolete, and that reading is mostly wrong.
A genuinely large context window does reduce the pressure on some of what earlier posts covered — aggressive chunk-size minimization, for instance, matters somewhat less when a model can comfortably ingest far more raw context at once. But it doesn’t eliminate the underlying problems retrieval solves. A larger context window doesn’t mean better attention to everything within it — the lost-in-the-middle effect doesn’t disappear just because the window got bigger; if anything, a much longer context gives a model considerably more opportunity for something important to end up buried in an under-attended middle section. It’s also a straightforward cost and latency problem: processing hundreds of thousands of tokens on every single call, even when the model architecturally supports it, is measurably slower and more expensive than processing a focused, well-retrieved few thousand tokens — and doing so for a question a much smaller, precisely retrieved context could have answered just as well is wasted cost with no accuracy benefit to show for it.
The more accurate framing: long context models change what’s possible — entire documents, or many documents at once, can now genuinely fit where they previously couldn’t — without making precise retrieval unnecessary. Retrieval remains the mechanism for finding the right information efficiently and keeping a model’s attention focused on what actually matters for a given query; a bigger window is a larger canvas to work with, not a replacement for judgment about what belongs on it.
Closing Thoughts
This post’s topics rarely show up in a RAG system’s architecture diagram, but they determine whether that architecture actually survives contact with production traffic: a token budget that isn’t tracked will eventually overflow, a context window that isn’t respected will silently truncate answers, and a model given tool access without careful design can retrieve inefficiently or unpredictably. None of this replaces the retrieval and generation quality work covered in earlier posts — it’s the operational layer that has to hold correctly underneath it.