RAG 與知識系統

Self-hosted RAG vs. platform/cloud services: a build-vs-buy framework for Greater China enterprises

Build your own RAG or buy a platform? Most enterprises frame this as a technology choice, but it's fundamentally a problem of risk and cost allocation. Three factors matter: data sensitivity, team capability, and total cost of ownership (TCO). Add Greater China's compliance requirements, and you have a framework for deciding whether to build or buy.

By

Tenten AI 研究團隊

AI 基礎設施

Published

January 17, 2026

Read time

6 分鐘

RAG自建vs平台企業AI選型TCO資料合規知識系統

Last quarter we evaluated the technical architecture for a medical device manufacturer in South China. They planned to buy an off-the-shelf RAG platform, which made practical sense: it's fast and comes ready to use. But when we reviewed their data inventory, the picture changed. Clinical trial records, supplier contracts, unpublished regulatory filings all needed to go into the vector store. Once that data left their premises, compliance and legal wouldn't approve it.

That meeting had one conclusion: they'd been asking the wrong question from the start.

Build vs. buy is a risk problem, not a technology preference

The choice between build and buy fundamentally allocates risk and cost, not technical preference. Building means assembling your own embedding models, vector databases, retrieval and reranking pipelines, and LLM layers, deployed on your infrastructure or private cloud. Buying means using a managed RAG service or cloud provider's knowledge base product, outsourcing infrastructure while you focus on data ingestion and prompt tuning.

Most people start by comparing retrieval accuracy. That almost never determines success or failure. What matters is whether your data can leave your perimeter, whether your team can sustain operations, and what you'll actually spend over three years. These three factors guide the decision.

Data sensitivity: the threshold question

If your knowledge base holds regulated personal data, unpublished financial records, core process parameters, or content that crosses data residency boundaries, the buy option is often not viable. There's no need for further comparison.

This boundary is especially strict in Greater China. Mainland China's Data Security Law and Personal Information Protection Law require explicit review and approval for moving critical data across borders. The financial sector adds data localization and regulatory filing requirements. Taiwan's Financial Supervisory Commission has similar rules for financial institutions using cloud services, including outsourcing and data residency. Healthcare has separate requirements for personal information and medical records.

In practice, segment your data rather than make an all-or-nothing choice. The truly sensitive 20 percent stays on premises; public or low-sensitivity content can use a platform to move faster. Don't saddle your entire company with self-hosting complexity because of a small data slice. And don't move things outside your control just to save effort when they belong inside.

Team capability: sustaining the system

RAG isn't finished when it launches. Launch is when operations begin. Documents change, models update, retrieval quality drifts, and users ask questions you didn't anticipate. Building means you own all of it.

A team that can run self-hosted RAG long-term needs people who understand vector retrieval tuning, LLM application engineering, DevOps and GPU resource management, and someone tracking evaluation metrics. This isn't a side project for one engineer; it's a product line. If you can't answer right now whether you need to rebuild your index when your embedding model updates, the true costs of building will arrive by month four.

Platforms amortize this cost into the subscription fee. The tradeoff is straightforward: you trade control for it. Vendors raise prices, change APIs, sunset features, and you adapt accordingly.

Total cost of ownership: the three-year calculation

Most cost comparisons only look at platform monthly fees versus self-hosted GPU bills. This is the most common mistake. Real TCO includes labor, opportunity cost, migration costs, and compliance audits.

Cost ItemSelf-Hosted RAGPlatform/Cloud Service
Initial SetupHigh (architecture, selection, integration)Low (out of the box)
InfrastructureYou manage GPU/storage/vector DBIncluded in subscription
Labor & OperationsRequires dedicated team investmentVendor handles most of it
Cost per Unit at ScaleAmortizes with volume, predictableScales linearly with volume
Compliance & Data ResidencyFull control, easy to auditLimited by vendor architecture
Vendor Lock-In RiskLowHigh (significant migration costs)
Time to LaunchSlow (weeks to months)Fast (days to weeks)

The pattern is straightforward. Low volume, need for speed, and non-sensitive data: the platform's TCO almost always wins. High volume, sensitive data, and long-term operation: self-hosted crosses a threshold where per-unit cost reverses, and you stop subsidizing vendor margins. Where that threshold sits depends on your query volume and data scale. Calculate it precisely once, not estimated.

Hybrid models in practice

Most mid-to-large enterprises in Greater China end up in a hybrid arrangement, not pure build or pure buy. Keep the sensitive data core on premises for control and compliance. Use platforms for peripheral, public, and experimental scenarios to validate value quickly. Let business units get a win in low-risk cases first, then move revenue-critical systems to your on-premises production.

When we help with selection decisions at Tenten, we don't ask which vendor you want first. We classify your data, inventory your team's capabilities, calculate your three-year TCO on paper, then determine what to build and what to buy. A system nobody uses or that won't pass compliance review fails regardless of demo polish. What matters is whether it's live, people use it daily, and auditors sign off on it.

One stuck workflow
is enough to begin

Tell us what the team does today, where it breaks down, and what a better working day should look like.