Autopsy of a Grounded Enterprise AI PoC: From Demo Day to Stillborn
An enterprise AI PoC that signed a contract, earned standing ovation at Demo Day, then died one step away from production. When we took over, actual usage was 0%. This case reveals what killed it: not weak technology, but the absence of clear ownership.
By
Tenten AI FDE 團隊
前線部署工程
Published
June 19, 2026
Read time
5 分鐘

This PoC wasn't killed by weak technology. It starved from the simple fact that no one took responsibility for shipping it.
Late last year we got called in to save a project. A mid-sized manufacturing client wanted an internal knowledge QA system, field engineers could ask questions in natural language and get answers from equipment manuals and historical work orders. They'd hired a capable systems integrator who built the PoC over three months. The system answered fast and accurately. At Demo Day, the director signed off on a company-wide rollout. Six months later, when we took over, the system was still in the test environment. Actual users: zero.
Enterprise AI PoC failures follow predictable patterns. Nearly every failure mode that appeared here shows up again in other projects. This is a complete diagnostic of one patient's pathology.
Case Summary
| Item | Details |
|---|---|
| Chief Complaint | Knowledge QA system won't go live; 0% adoption rate |
| Timeline | 3-month PoC → demo approval → stuck in test environment for 6 months |
| Presenting Symptoms | "Waiting on IT to allocate resources," "Permissions haven't been set up," "Accuracy needs more tuning" |
| Root Causes | Fabricated demo dataset, no owner, no adoption design, no exit criteria |
| Status on Arrival | Model runs fine, but no department considers it their responsibility |
Death Cause #1: The Demo Ran on 'Curated' Data, 32 Handpicked Documents
The vector database held exactly 32 documents. Every one was cherry-picked by the integrator, sanitized, and reformatted into clean PDFs. The demo questions all had answers in those 32 docs.
The company's actual repair knowledge was scattered across 14,000 work orders, scanned images, Excel sheets, and the heads of senior technicians. When we tested the same questions against the real corpus, accuracy dropped from 95% to 41%. The demo had been running in a world that didn't exist.
Any accuracy number never tested against real, messy data isn't accuracy, it's a glamour shot. Benchmarks in week one should run against the actual documents in worst shape. A rough-looking demo beats discovering the gap after going live.
Death Cause #2: Nobody Owned Shipping It
The question was straightforward: who owns this system after launch? IT said the business unit requested it. The business unit said IT maintains the model. The integrator said their responsibility ended at PoC acceptance.
All three were right. Together, that meant nobody.
The PoC had a clear owner, the person collecting the project bonus. But getting 200 field engineers to actually use this every day wasn't in the contract. It wasn't in anyone's KPIs. No one woke up thinking it was their job. A system doesn't ship itself. Someone has to carry it to production and then stand guard afterward, watching metrics, fixing bad cases, pushing process changes.
Death Cause #3: Accuracy Became an Endless Excuse
The official story was 'accuracy needs more tuning.' It sounded responsible. It was a pit with no bottom because no one had defined what 'tuned enough' meant.
No acceptance criteria meant any bad case could freeze the launch indefinitely. They were forced to sit down and draw a line: if the knowledge QA system returns the right answer in the top-3 results and includes the source text, it passes. Remaining bad cases go to the backlog, they don't block launch. Three days later, it shipped. Two weeks of real production data beat six months of their assumptions.
Death Cause #4: Nothing Was Designed for Actual Usage
Even if it cleared the first three hurdles, this system would have gone dormant. It was built as a standalone web page. Engineers had to open another browser tab, log in again, and re-type their questions. Real people are on the shop floor wearing gloves, eyes on equipment. Who switches three windows to look something up?
Adoption grows inside existing workflows, not alongside them. It was embedded into the work order system they already used every day. A query pulls related manual sections automatically. Adoption didn't climb from pushing it. It climbed because users didn't have to change habits. A month after launch, weekly active users reached 63%.
The One Thing This Diagnosis Should Tell You
Looking back, none of these four causes were 'AI wasn't strong enough.' The model was good enough on day one. It died from a fabricated demo dataset, orphaned responsibility, perfectionism without a finish line, and confusing 'it runs' with 'people use it.'
We put engineers onsite to own the launch and adoption metrics, not hand over a completed demo. A beautiful demo doesn't count. Running on a screen doesn't count. Only shipping, and having real people using it daily, counts.

One stuck workflow
is enough to begin
Tell us what the team does today, where it breaks down, and what a better working day should look like.