LLM.coPrivate, self-hosted LLM deployments
Legal AI infrastructure for firms
AI RFP discovery and response drafting
Automatic.coBusiness process automation
Secure AI virtual data rooms
What a Production RAG System Actually Costs in Year One
Most RAG proposals land in your inbox as a single number. A six-figure build, a vague monthly estimate, a line about "scales with usage." That number is almost never wrong on purpose. It is wrong because the person who wrote it has not decomposed the system into the four or five variables that actually move year-one cost: corpus size, query volume, model tier, and whether a human stays in the loop for evaluation.
This piece walks through the line items of a production retrieval-augmented generation system over its first twelve months, with current list prices and the trade-offs that matter when you compare proposals or build an internal budget. It assumes you have already decided to build. If you are still choosing between fine-tuning and retrieval, that is a different conversation.
So where does the money actually go?
The Five Line Items That Move the Total
A production RAG system has roughly five recurring cost centers and one upfront one. The upfront cost is the build: engineering time to connect sources, chunk and clean documents, stand up retrieval, write evals, and ship a UI or API. The recurring costs are embeddings, vector storage and query, LLM inference, evaluation and observability, and ops.
For a mid-sized internal system (say, 2–5 data sources, a few hundred daily users, hybrid search with reranking), a production chatbot with citations, monitoring and multiple sources typically costs $30,000 to $80,000 to build, usually over 8 to 12 weeks. That maps to a small senior pod: a backend engineer, an ML engineer, and part-time design and PM. Enterprise systems with SSO, audit logging, multi-tenancy, and content governance land higher, often $120k–$250k before the first month of production traffic.
Three variables dominate year-one run cost:
- Query volume and average output length. Inference is almost always the largest ongoing line.
- Model tier. A GPT-4o-class model costs roughly 10× a small model per token.
- Vector store choice at scale. Below ~10M vectors the numbers are rounding errors. Above 50M they diverge sharply.
Everything else — embeddings, evals, ops — tends to be a smaller, more predictable band.
Ingestion and Embeddings Are the Cheap Part
Teams spend disproportionate anxiety on embeddings because that is where the pipeline starts. In dollar terms it is usually the smallest line. OpenAI's text-embedding-3-small runs $0.02 per million tokens, and text-embedding-3-large runs $0.13, with a 50% discount on the Batch API for 24-hour turnaround. A 10-million-page corpus at roughly 1,000 words per page is about 13 billion tokens. On 3-small, embedding the entire thing costs about $260. On 3-large, about $1,700. Reindexing quarterly on 3-small is still under $1,100 per year.
Ingestion engineering is where the real work hides. Pulling from Confluence, SharePoint, Salesforce, S3, Postgres, and a handful of SaaS APIs; handling PDFs with tables; deduping near-identical versions; respecting ACLs so the retriever cannot leak a document the user is not allowed to see. Expect 25–40% of the build budget here. The API integration work is usually larger than the model work. If your corpus is well-governed and lives in two or three systems, cut that estimate in half. If it is a 15-year-old file share, double it.
Vector Store Choice Is a Function of Scale
Below ten million vectors, almost any option is fine and the price difference does not matter. Above that, the slope changes. At 10 million vectors, Pinecone Serverless costs roughly $70/month, Weaviate Cloud $135, Qdrant Cloud $65, and pgvector on RDS about $45; at 100 million vectors, Pinecone can reach $700+/month while self-hosted Milvus or pgvector stays under $100.
Two things move those managed numbers in practice. First, Pinecone introduced a $50/month minimum pricing floor in 2025, while Weaviate set a $25/month floor, which matters for small tenants and multi-environment setups (dev/stage/prod each pay the floor). Second, Pinecone's serverless meter has three dimensions: $0.30 per GB per month for storage, $4 per million write units, and $16 per million read units. A chatty app with heavy reindexing can run several multiples of the storage line.
The honest trade-off: managed services buy you backup, replication, and someone to call at 2 a.m. pgvector on an existing Postgres cluster is cheaper and keeps your data in one place, but your team now owns HNSW tuning, replica strategy, and the operational edges. For teams that already run Postgres at scale, pgvector often wins; PostgreSQL logical replication handles the multi-region story without a second system. For teams whose only database is a managed instance they rarely touch, a hosted vector DB is worth the premium.

Inference Is Where the Budget Actually Lives
Pricing a RAG system without a query volume estimate is guessing. The number you need is daily queries × average input tokens (prompt + retrieved context) × average output tokens. Retrieved context is usually the big one: a top-k of 8 passages at 500 tokens each is 4,000 tokens of input before the user even asks anything.
Current list for the common choices: GPT-4o is $2.50 per million input tokens, $1.25 per million cached input tokens, and $10.00 per million output tokens. Smaller tier models from OpenAI, Anthropic, and Google land roughly an order of magnitude lower. Prompt caching matters: in a RAG system, the system prompt and much of the retrieved context repeat across sessions, and a 2x discount on cached input is real money at scale.
A worked example. 5,000 queries per day, 5,000 input tokens, 400 output tokens, GPT-4o, no caching:
- Input: 5,000 × 5,000 = 25M tokens/day × $2.50/M = $62.50/day
- Output: 5,000 × 400 = 2M tokens/day × $10/M = $20/day
- Monthly: roughly $2,475. Annual: about $30,000.
Drop to a small-tier model for 70% of queries and route only hard ones to GPT-4o, and the same workload lands near $8,000–$12,000 per year. This is where routing, caching, and context compression earn their keep. It is also where prompt injection defenses belong in the SOW, because once you tier models, you need to guarantee that a malicious prompt cannot escalate itself into the expensive one.
Self-hosted inference on vLLM or TGI is viable above a certain volume. Rough rule of thumb: an A100 or H100 running Llama 3.1 70B or Qwen 2.5 72B pays for itself against hosted mid-tier pricing somewhere between 10M and 50M tokens per day, before you count the SRE time to keep it healthy.
Evals and Observability Are Not Optional
You cannot ship RAG without an eval harness, because retrieval quality is the whole product. On industry-standard CRAG benchmark testing, most advanced LLMs achieve under 34% accuracy, straightforward RAG improves accuracy to only 44%, and state-of-the-art industry RAG solutions answer just 63% of questions without hallucination. The 63% number is the one to anchor on when a vendor promises "near-human" accuracy out of the box.
Budget two things. A golden-set evaluation run against each release — a few hundred to a few thousand labeled Q&A pairs scored by an LLM-as-judge — typically costs $20–$200 per run depending on dataset size and judge model. If you ship weekly, that is maybe $500–$2,000 per month. The second line is production observability: tracing, retrieval quality sampling, hallucination flagging, and user feedback capture. LangSmith, Langfuse, Arize, or Phoenix sit in a $200–$2,000 per month band depending on volume and retention.
Add a quarterly human review: a subject-matter expert spending a day labeling a sample of production traffic. That is often the cheapest quality lever in the entire system.
Build vs Buy vs Partner Changes the Shape of Year One
The buy side of this question has shifted. Menlo Ventures found that 76% of enterprise AI use cases are now purchased rather than built internally, a sharp reversal from the prior year when 53% were built in-house. For a well-defined horizontal use case (customer support deflection, developer docs search), an off-the-shelf product may be the right answer. For anything that touches proprietary workflows, regulated data, or an existing product surface, custom still wins — and the gap between a demo and a production system is where most buyers underestimate.
Three delivery shapes to compare:
- In-house. Two engineers plus fractional ML and SRE time, 4–6 months to production. All-in cost usually $350k–$600k year one if you include loaded salary. You own the knowledge.
- Partner build, hand off. A senior pod delivers in 10–16 weeks for $120k–$300k, then transitions to your team. Our dedicated team cost breakdown covers the staffing math. Ongoing run cost sits with you.
- Managed delivery. Partner builds and operates, you pay a monthly retainer. Predictable, slower to internalize.
The partner-build model is usually the right default for a first production system, because the eval harness and ops runbook you inherit are worth more than the code. A good RFP for custom software should ask vendors to price each of the six line items separately, not a lump sum, and to show you an eval report from a previous build.
- 11. Source discoveryInventory corpora, ACLs, and refresh cadence · Data + domain SMEs
- 22. Ingestion & chunkingConnectors, parsers, dedupe, metadata · Backend engineering
- 33. Retrieval stackEmbeddings, vector store, hybrid search, rerank · ML engineering
- 44. Generation & routingModel tiering, prompt caching, guardrails · ML + application
- 55. Evals & observabilityGolden set, LLM-as-judge, tracing, feedback · ML + QA
- 66. Ops & handoffOn-call, runbooks, reindex schedule, cost review · SRE + product
A Realistic Year-One Number
Add it up for a representative mid-market system: 10M vectors, 5,000 daily queries, mixed-model routing, weekly releases, standard observability. Build at $150k, inference at $18k, vector store at $1,500, embeddings at $500, evals and observability at $12k, ops and on-call at $25k. Year one lands around $210k. The same system, in-house from scratch with no partner, lands closer to $450k when you honestly count salary.
The number any vendor hands you is only useful if you can see the five sliders behind it. Ask for the query volume assumption, the model tier assumption, the vector store at target scale, the eval cadence, and the on-call model. If a proposal cannot answer those five questions in writing, the number on the cover page is not a budget. It is a wish. For teams scoping this work with an outside AI development partner, the right first conversation is about those sliders, not the total.
