Legal AI infrastructure for firms
AI RFP discovery and response drafting
Automatic.coBusiness process automation
Secure AI virtual data rooms
Why More Custom AI Rollouts Will Be Done In-House
For two years, the default assumption was that enterprises would rent their way into generative AI. Pick a frontier model API, pay per token, let a SaaS vendor or consultancy handle the deployment plumbing, and move on. That assumption is quietly collapsing. The model-selection decision has stabilized, open weights have caught up on a lot of workloads, and the hard part has shifted: it is now the rollout itself — the deployment, the integration into existing systems, the compliance posture, the on-call rota — that decides whether the project returns anything.
What follows is a look at why that work is increasingly staying inside the enterprise in 2026, and what a credible internal rollout actually looks like.
The Economics Of Serving Have Flipped
Through 2024, GPU scarcity was the single best argument for handing inference to a vendor. That argument has thinned. H100 rental medians have fallen from more than $7 per hour in early 2024 to around $3.38 per hour by mid-2026, with marketplace floors well below that. On specialty clouds, on-demand H100 SXM5 pricing now runs roughly $2.20 to $3.00 per GPU-hour, with spot in the $1.10 to $1.80 range. AMD MI300X capacity, long a theoretical alternative, is actually bookable now at $1.71 per hour on TensorWave and $2.39 on Runpod.
Two things follow. First, the break-even math against per-token API pricing moved. For any workload with steady utilization, running your own endpoint on a reserved H100 or MI300X is no longer the exotic option. Second, the serving stack caught up to the hardware. vLLM has become the default open-source engine, with first-class integration into KServe, BentoML, and Ray Serve, and an OpenAI-compatible API that drops into existing clients. Hugging Face's TGI, the other common choice, entered maintenance mode in December 2025, which pushed most new deployments to vLLM or SGLang. That consolidation matters: there are now two or three serving engines that senior infrastructure engineers can actually standardize on, document, and hand off.
The Model Choice Stopped Being The Interesting Problem
Eighteen months ago, "which model" was the thing enterprises hired consultancies to answer. In 2026 it is a shorter conversation. The frontier closed-weight models are still ahead on the hardest reasoning tasks, but for the median enterprise workload — summarization over internal documents, structured extraction, classification, retrieval-augmented answers, workflow agents hitting internal APIs — the open-weight families now land close enough that the deciding factor is operational, not benchmark-driven.
That shifts the center of gravity. The expensive, high-judgment work in a rollout is no longer picking a model. It is the integration surface: how the model gets access to the right data, how it is observed, how its outputs are evaluated against regressions, how it fails safely when a tool call does. That work is tightly coupled to internal systems — identity, data lineage, feature flags, incident response — which is exactly the work that does not travel well across a vendor boundary. If you want a benchmark for what in-house delivery looks like, the dev.co piece on building an AI team that ships covers the staffing shape in more detail.

Regulation Rewards The Team That Owns The Logs
The compliance picture in 2026 is the second quiet forcing function. Under the EU AI Act, August 2, 2026 is the binding enforcement date for high-risk AI system obligations, covering Articles 9–17 for providers and Article 26 for deployers. Deployer obligations include human oversight mechanisms, Fundamental Rights Impact Assessments where applicable, and automated logs retained for a minimum of six months.
Those are not vendor-friendly requirements. Six months of per-request logs, with input references and the identity of the human reviewer, has to live somewhere auditable — and the enterprise is the entity on the hook, not the SaaS provider. The Cloud Security Alliance's reading of the deadline notes that over half of organizations lack systematic AI inventories, and the practical response for most engineering teams has been to pull the serving layer, the logging layer, and the evaluation layer back inside the perimeter where SOC 2, HIPAA, and the EU AI Act obligations already have owners. The NIST AI Risk Management Framework gives US buyers the parallel template, and most mature programs now map their controls to both.
Vendor Lock-In Is The New Shadow IT
Procurement teams have started treating AI vendor dependencies the way they treated single-cloud dependencies in 2015. The concern is not ideological. It is that model behavior changes under you, pricing changes under you, and the switching cost on a deeply embedded SaaS AI product is dominated by prompt-engineered glue code and non-portable eval harnesses that nobody wants to rewrite.
An internal AI platform does not remove this risk, but it localizes it. You still pick a base model, and you still pay for it. The difference is that the retrieval layer, the prompt templates, the tool schemas, the observability, and the evaluation suite live in your repository, under your CI, with your tests. Swapping a 70B open-weight model for the next one, or routing a subset of traffic to a frontier API for hard cases, becomes a config change rather than a vendor renegotiation. IBM's own argument that proprietary data is the competitive edge in generative AI reads the same way: the value lives in how your data is shaped and served to the model, which is work you cannot subcontract without giving up the edge.
The Agent Frameworks Finally Settled Down
The 2024 agent frameworks were a mess. The 2026 ones are not. Patterns around tool calling, structured output, memory, and graph-style orchestration have converged enough that a competent backend team can build a production agent without placing a long-term bet on any single framework. That is what makes in-house ownership realistic for a wider range of teams.
Agent rollouts also surface problems that no vendor demo will catch: prompt injection via user-controlled content, tool misuse under adversarial inputs, and privilege escalation when an agent is given write access to internal systems. The dev.co note on prompt injection defenses is the short version of what to require in any statement of work before an agent touches production. Those controls are easier to enforce when the agent lives in your stack, behind your service mesh, with your rate limits.
A Concrete Internal Rollout Checklist
If a team is pulling an AI rollout in-house, the following is the shape of the work that tends to get underestimated. Treat it as a baseline, not a maximum.
- Serving. Pick one engine (vLLM is the safe default) and pin a version. Decide tensor parallel size, max model length, and KV cache budget before benchmarking. Put the service behind an OpenAI-compatible gateway so client code is portable.
- Capacity. Model peak concurrent requests, not average. Reserve baseline capacity; use spot or burst for overflow. Document the fallback when GPU capacity is exhausted — degrade to a smaller model, queue, or return a typed error.
- Retrieval. Treat the index as a first-class service: schema, migrations, backfills, re-embedding plans. pgvector or a dedicated vector store, with the same deployment discipline as the rest of the database tier.
- Evaluation. A golden set in version control, run on every model and prompt change, with pass thresholds enforced in CI. No silent prompt edits in production.
- Observability. Per-request traces including prompt, retrieved context, tool calls, latency, token counts, and model version. Prometheus metrics for queue depth, cache hit rate, and GPU utilization.
- Audit and retention. Logs stored to meet the longer of SOC 2, your sector regulation, and the six-month EU AI Act floor. Reviewer identity captured where humans are in the loop.
- Safety. Input and output filters, tool allow-lists, and a documented process for incidents. Red-team the agent surface before launch, not after.
- Cost control. Per-tenant token accounting, caching for repeat prompts, and a kill switch at the gateway. The cost model for a production RAG system in year one is a reasonable template to start from.
What To Hand To A Partner, And What To Keep
In-house does not mean alone. Most enterprises moving an AI rollout internal still bring in outside engineering for the first six to twelve months — to stand up the serving platform, write the first evaluation harness, and transfer the on-call playbook. The useful partner here is one that writes production code and leaves, not one that charges per API call forever. That is the shape of the AI engineering work dev.co tends to be asked for, and it is a reasonable lens on what to procure. For buyers writing the SOW, the general RFP checklist applies with one addition: require that every artifact — model configs, prompts, eval sets, infra as code — is delivered into your repository, not a vendor portal.
The things worth keeping internal are the pieces that encode your business: the data pipelines, the evaluation suite, the prompt and tool schemas, the incident process. Those are where the actual advantage compounds, and they are the pieces that make switching models cheap when the next generation of open weights lands. The rollout phase is where that discipline is set, which is why more of it will be done in-house from here.
