LLM.coPrivate, self-hosted LLM deployments
Legal AI infrastructure for firms
AI RFP discovery and response drafting
Automatic.coBusiness process automation
Secure AI virtual data rooms
Building an AI Development Team That Actually Ships
Most AI team plans fail at the org chart, not the model. A founder hires two ML researchers, waits six months for a demo, and discovers that the demo cannot be deployed, cannot be evaluated, and cannot survive the first user who types in a prompt injection. The headcount was real. The structure was wrong.
Shipping an AI product is a distributed-systems problem with a probabilistic component attached. The ratio of modeling work to platform, product, and evaluation work is usually inverted from what new teams assume. That inversion is the subject of this piece.
So what does an AI development team that actually ships look like in 2026?
Start With the Product, Not the Platform
Before any role gets a title, decide which category of system you are building. The staffing follows from that.
- RAG over your own documents. The hard problems are retrieval quality, chunking, access control, and freshness. Modeling is light. A frontier model does the generation.
- Agentic workflows. The hard problems are tool contracts, state, failure recovery, and guardrails against prompt injection. Still light modeling, heavier backend and security engineering.
- Custom or fine-tuned models. The hard problems are data, training infrastructure, and offline evaluation. This is the only category where ML research hires are the first move.
Roughly 80% of the applications we see at DEV.co's AI practice fall into the first two categories. Teams keep staffing for the third. A March 2025 Bain survey found nearly 44% of executives citing in-house AI expertise as a key barrier to generative AI delivery, and part of that gap is a mismatch: companies are hunting PhDs when they need senior backend engineers who understand retrieval.
The Six Roles a Shipping Pod Actually Needs
A pod of six can carry a RAG or agentic product from zero to production in a quarter. The composition that works:
- Tech lead / principal LLM engineer (1). Owns the system design: gateway, retrieval pipeline, eval harness, guardrails service. Writes code. This is the role that keeps the pod from fragmenting into prompt tinkerers.
- Backend engineers (2). Build the ingestion pipeline, the orchestration layer, the tool handlers, the auth model. The majority of RAG code is boring plumbing in Python, Go, or TypeScript, and it needs people who have shipped boring plumbing before.
- AI / applied engineer (1). Owns prompts, retrieval tuning, model selection, and the day-to-day eval loop. Public job data confirms this role is distinct: RAG shows up in 39% of AI Engineer postings versus 16% of ML Engineer postings.
- Platform / MLOps engineer (0.5 to 1). Owns deployment, GPU or inference infrastructure, observability, cost telemetry. Fractional at first, full-time by the second production surface.
- Product manager (0.5 to 1). Owns the use case, the quality bar, and the kill criteria. AI PMs who can read eval scores are worth more than AI PMs who can write a vision deck.
- Designer (0.5). Fractional, but non-optional. Agent interfaces are new UX territory and uncertainty visualization is a design problem, not an engineering one.
Notice the ML-to-everything-else ratio. One applied engineer, four or five people doing platform and product work. If the model is the product, add researchers. If the product is the product, do not.

Who Owns Evals, Guardrails, and Retrieval
These three are where new AI teams quietly fail. Nobody owns them, so everybody writes a little bit of each, and none of it is production-grade. The split that works:
Retrieval belongs to a backend engineer with the AI engineer consulting. It is a data pipeline problem: ingestion, chunking strategy, embedding model choice, index management, permissions enforcement, freshness. Treat it like any other pipeline, with SLAs on data freshness, access-control-aware retrieval, and index refresh strategy. This is where things like API integration work and schema design matter more than clever prompting.
Evals belong to the AI engineer in a small pod, with the tech lead gating releases. The platform team supplies infrastructure. The platform team does not write rubrics. This boundary matters because the team closest to the user owns the definition of "good" for their surface; the platform team only owns the harness and the CI gates. Delay this split and the backlog of unowned rubrics becomes its own project.
Guardrails are shared. Input validation, output filtering, PII redaction, and injection detection are platform concerns. Policy content, refusal behavior, and tolerance thresholds are product concerns. A 2024 McKinsey Global AI Survey found that 63% of companies using generative AI do not have governance structures in place for the associated risks, and the pattern underneath is almost always diffuse ownership. Name an owner per layer, in writing, before the first production release. The same discipline applies to prompt injection defenses in an agentic SOW: if the contract does not name who owns the retrieval rail and the output filter, nobody will.
Ratios That Keep the Chart Honest
A chart with the right roles and the wrong ratios still breaks. The ratios below are the ones that get cited across staffing benchmarks in 2026, in decreasing order of how often teams violate them.
The platform ratio is the one most teams get wrong. Published benchmarks suggest roughly one MLOps or platform engineer for every four to six people building models, and thinner than that makes deployment itself the bottleneck. The product-management ratio is the second: one AI PM per six engineers keeps the roadmap from devolving into Slack arguments about which demo matters.
When a Pod of Six Beats a Department of Twenty
A six-person pod with full ownership of a single surface ships faster than a twenty-person department split across three. The reason is coordination overhead, not talent. One pod can hold the entire system in its head: retrieval, prompts, eval set, guardrails, gateway, UI. Twenty people cannot, so they build interfaces, then meetings, then reorgs.
The pod model works until one of three things happens: a second production surface emerges, cost or latency becomes a board-level number, or compliance requires a named governance owner. At that point the platform work splits off and a central team forms. Not before. Pre-building a platform for a product that does not exist is the fastest way to burn a year.
If staffing a full-time pod is not realistic yet, a dedicated external team can carry the first production surface while the first internal hires come online. Our breakdown of dedicated team cost structures covers that tradeoff honestly, and the year-one cost of a production RAG system frames the budget around the same six-role shape.
The Hiring Sequence That Avoids Dead Weight
Order of hire matters more than total headcount. The sequence below reflects what we see work for product-first AI teams, with each hire triggered by a bottleneck rather than a plan.
The counterintuitive move is hiring the platform engineer early, around hire four or five. Deployment debt compounds silently, and MLOps work rediscovered at hire twelve is twice as expensive. The second counterintuitive move is delaying the research hire until the problem genuinely requires a custom model. The 2025 PwC Global AI Jobs Barometer found AI-skilled workers commanding a 56% wage premium over peers in the same roles, more than double the prior year, and research-track hires sit at the top of that curve. Pay for them when the model is the product. Not before.
Governance belongs on the chart earlier than most teams want. McKinsey's 2024 State of AI survey reports 28% of AI-using organizations place CEO responsibility on governance and 17% place it with the board, which tells you that at the enterprise end, this is not a job to leave undefined. A fractional risk owner at ten people becomes a named role at thirty.
What a Shipping Chart Looks Like a Year In
A year after the first hire, a healthy AI org looks less like a research lab and more like a product engineering group with a specialty. One or two pods of five to seven. A thin platform team of three to five, owning the gateway, the eval harness infrastructure, and the guardrails service. A named governance owner, part-time or full. A product manager per pod. Model-building work concentrated in the people who actually need it, not spread thin across every engineer who wanted to touch an LLM.
The teams that ship treat AI as an engineering discipline with probabilistic components, not a separate species of work. The chart reflects that. If yours does not, the fix is usually fewer researchers and more backend engineers who understand retrieval, not the other way around.
