Turning Court Discovery into a Viral Product: How Extend Rebuilt Elizabeth Holmes' 2016 Desk to Showcase Document Extraction

An analysis of how an interactive retro-OS simulation turned messy federal trial PDFs into an undeniable product showcase for modern document intelligence.

Published: 2026.10.09

From Criminal Court Exhibits to Interactive OS: Reconstructing the Theranos Desk

When the United States federal government prosecuted former Theranos chief executive Elizabeth Holmes, the discovery process unsealed thousands of internal corporate artifacts. The public record swelled with private text messages, slide decks, lab equipment logs, and email threads documenting the collapse of a multi-billion-dollar blood-testing startup. For years, these records existed as cumbersome, scanned PDFs buried inside government repositories and legal databases.

Bo Lau, an engineer at document intelligence company Extend, took that mountain of unstructured legal files and turned it into an interactive time capsule. Instead of publishing a conventional whitepaper or an academic case study, Lau rebuilt Holmes’ 2016 workspace as an interactive web app. Visitors land directly in front of an Apple MacBook Air running OS X El Capitan, complete with a period-accurate iPhone resting beside it.

From Raw Legal Discovery to Structured Interactive Simulation

The engineering pipeline behind Extend's retro-OS demonstration

1

1. Federal Docket Extraction

Ingesting thousands of heterogeneous trial exhibits, including scans, emails, and images.

2

2. Document Splitting and OCR

Parsing unformatted PDF batches into clean text, conversation threads, and metadata.

3

3. Interactive State Mapping

Binding parsed messages and slides to virtual macOS and iOS native application state.

Users can slide to unlock the virtual iPhone, read actual text exchanges entered into evidence during United States v. Holmes, open presentation files inside the simulated laptop, and inspect operational readouts from the Theranos Edison machine. The project was timed alongside a promotion for an upcoming watch party for Nathan Fielder’s documentary You Can See Everything, but its underlying function is an exercise in product engineering as marketing.

Traditional business software companies struggle to demonstrate what their data pipelines actually do. Telling prospective enterprise customers that software can parse, extract, and split difficult documents often falls flat because every legacy optical character recognition (OCR) vendor makes the exact same claim. By feeding thousands of chaotic, real-world trial exhibits through Extend’s data parsing pipeline and presenting the output as a fully navigable operating system, the company replaced abstract marketing claims with tangible proof.

Handling raw court discovery is a classic stress test for data extraction. Legal exhibits arrive in terrible condition: skewed smartphone screenshots, low-resolution black-and-white scans, overlapping handwriting, inconsistent table layouts, and misaligned page numbers.

To understand why an engineering project like the Theranos desk simulation stands out, one must look at how corporate teams, legal departments, and operations desks traditionally handle bulk unstructured documents versus how modern multimodal document models process them.

Metric and CapabilityManual Paralegal / Data EntryLegacy Template-Based OCRMultimodal Document Extraction (Extend / Modern AI)
Throughput Speed (per 1,000 pages)35–45 operating hours8–15 minutes45–90 seconds
Estimated Processing Cost$2,800–$4,500 (labor wages)$15–$35 (server and licensing)$8–$20 (token and inference pricing)
Handling of Unformatted ScansHigh accuracy, slow speedHigh error rate (breaks on layout drift)High accuracy (understands context and layout)
Contextual Text ReassemblyExcellent human judgmentPoor (chops text across columns)High (reconstructs threads and tables)
Setup and Maintenance OverheadHigh continuous labor costHigh setup time (needs custom regex templates)Low setup time (zero-shot prompt and schema mapping)
Data Export ReadinessManual spreadsheet entryRequires heavy programmatic cleaningDirect export to structured JSON / SQLite

Industry simulation based on enterprise paralegal labor rates ($75–$100/hr) versus baseline cloud OCR engines and modern document parsing pipelines.

When engineering teams attempt to ingest public records using legacy optical character recognition, the system frequently fails on basic structural logic. A standard OCR tool reads text left-to-right across an entire scanned page. If a court exhibit contains a side-by-side comparison of blood draw test logs next to an email snippet, the legacy tool will interleave the text from both columns into an unreadable mess.

Fixing this problem historically required engineers to write brittle bounding-box rules or regular expressions for every distinct document format. If the government scanned an exhibit at a five-degree tilt, the rules broke.

Modern document parsing engines bypass this fragility by reading documents visually and contextually at the same time. The engine recognizes that an iPhone screenshot contains chat bubbles, detects the sender, stamps the timestamps, and extracts the conversation into a clean database. That structured database is what allows an engineer to bind real evidentiary records to a clickable web app with zero manual transcription.

Why Unstructured Data Ingestion Still Paralyzes Enterprise Workflows

The Theranos desk simulation is entertaining on the surface, but the underlying operational bottleneck it addresses is one of the most expensive problems inside global enterprises: unstructured data debt.

When operational workflows rely on human beings to copy-paste information out of complex documents, businesses suffer predictable friction across three distinct vectors.

Operating Expenses: The Hidden Tax of Manual Data Extraction and Correction

Enterprise organizations run on documents that computers cannot naturally read: invoices, customs declarations, contracts, medical intake forms, and insurance claims. When an enterprise processes tens of thousands of these documents every month without automated parsing, operating budgets balloon.

Companies often hire offshore business process outsourcing (BPO) teams or deploy internal operations staff to manually inspect documents, type line items into enterprise resource planning (ERP) systems, and cross-reference dates. When errors inevitably occur—transposing a serial number, misreading an invoice subtotal, or missing a critical legal disclaimer—the cost to audit and remediate those mistakes is often triple the cost of the initial data entry.

Lead Time Bottlenecks: How Multi-Format Document Queues Stall Enterprise Decisions

In paper-heavy sectors like commercial insurance, freight logistics, and corporate compliance, document intake is the single largest driver of operational turnaround lag.

When a freight shipment arrives at a port, or when an insurer receives a complex commercial liability claim, the review process cannot proceed until all attached documentation is verified. If workers take forty-eight to seventy-two hours to manually parse supporting evidence, every dependent operation grinds to a halt. In competitive environments, a two-day delay in processing an application or clearing an import record directly leads to customer churn and expensive storage demurrage penalties.

Data Reliability and Accuracy: Why Hallucinations and Missed Edge Cases Kill Automation Pipelines

Early attempts to automate document intake using generic large language models (LLMs) created a different operational crisis: hallucination. If an enterprise feeds an eighty-page loan agreement into a standard chatbot prompt, the model might invent terms, miss footnotes, or drop specific numeric constraints.

In high-stakes industries, an extraction pipeline must provide exact citation provenance. Operators need to see precisely where an extracted value came from on the source page. Without verifiable source anchoring, automated ingestion pipelines fail internal risk and compliance audits. A functional document engine must guarantee that every parsed record matches the underlying source text with zero distortion.

Engineering As Marketing: How Technical Demos Outperform Generic Whitepapers

Extend’s project underscores an important trend in business-to-business growth strategy: building genuine software tools and interactive experiences delivers far higher brand recall than traditional marketing campaigns.

For decades, enterprise software marketing followed a predictable formula. Marketing departments published static whitepapers behind email-capture forms, posted generic blog posts, and booked trade show booths. Prospective buyers were forced to sit through thirty-minute sales presentations before seeing how the software actually worked.

B2B Go-To-Market Execution: Traditional Marketing vs. Interactive Engineering

Comparing engagement patterns between static marketing content and interactive technical showcases

Traditional B2B Content Funnel

Low Intent / High Churn
  • • Gated 20-page whitepapers that rarely get read
  • • Generic product videos showing curated mockups
  • • High friction requiring mandatory sales qualification calls
  • • Zero public proof of edge-case handling

Interactive Engineering Showcase

Viral Distribution
  • • Instant public utility and immediate browser interaction
  • • Real-world dirty data stress-tested in public view
  • • Organic developer sharing across social and industry feeds
  • • Self-evident proof of core product extraction capability
Editorial Verdict: Interactive software demonstrations convert passive readers into active evaluators by letting the technology prove its own capabilities in real time.

Today, technical buyers—such as lead engineers, product managers, and operations directors—dislike standard sales pitches. They want to see the system handle real edge cases immediately.

When an engineer demonstrates that an extraction tool can digest thousands of messy criminal trial exhibits, untangle nested threads, and output those records into a functional web application, the technical capability is proven in public. Prospective enterprise buyers do not need to take the company’s marketing copy at face value; they can test the responsive environment directly.

Modern development teams often orchestrate these automated document parsing pipelines by connecting ingestion APIs to automation platforms like Make, routing structured records directly into back-office databases without writing custom glue code from scratch. When the source data is cleanly parsed at the entry point, automating downstream actions becomes straightforward.

The Next Phase of Enterprise Document Processing and B2B Growth

The technical execution behind the Theranos desk simulator signals a broader reorganization across the enterprise software and data extraction market over the next twelve to twenty-four months.

Legacy OCR Vendors Face Radical Margin Compression

Legacy software providers that built their business models around charging premium per-page licensing fees for rigid, template-based OCR will face severe margin pressure.

Enterprise clients are no longer willing to pay tens of thousands of dollars in professional services fees just to configure template extractors that break every time a vendor changes an invoice layout. As modern multimodal document models and specialized parsing platforms become cheaper and easier to deploy, legacy optical character recognition will be treated as an obsolete technology. Providers that cannot deliver zero-template, context-aware document splitting will lose market share to agile platforms that process raw documents out of the box.

The Shift in Document Intelligence Economics

Projected efficiency gains from modern multimodal parsing over legacy OCR systems

90%

Setup Time Reduction

Eliminates weeks of manual regex and template mapping.

4x

Ingestion Throughput

Parallel processing of heterogeneous files in seconds.

-60%

Exception Handling Cost

Fewer human reviews required for skewed or damaged scans.

Three Non-Negotiable Capabilities for Tomorrow’s Document Platforms

Enterprise buyers evaluating document intelligence solutions must look beyond basic marketing claims and verify three core architectural requirements:

  • True Layout-Agnostic Parsing Without Templates: The platform must ingest unstandardized, multi-page PDFs containing mixed content—such as tables, handwritten notes, and low-resolution imagery—without requiring pre-configured bounding boxes or layout templates.
  • Auditable Lineage and Source Grounding: Every extracted data field must map directly to bounding coordinates on the original source document. If an auditor clicks an extracted figure, the software must instantly highlight the exact pixels where that value originated.
  • Product-Led Developer Tooling: The platform must provide robust programmatic access, clear documentation, and transparent testing environments. Solutions that require multi-week professional services engagements simply to test sample documents will be bypassed in favor of developer-friendly platforms that prove their accuracy on day one.

Interactive projects like Bo Lau’s Theranos simulation demonstrate that the era of abstract software marketing is fading. Whether parsing federal court exhibits or enterprise supply chain manifests, the companies that succeed will be those that let their engineering speak for itself.

Weekly Briefing

Weekly Tech & Business Data Briefing

Verified software analysis, practical gotchas, and essential supply-chain updates delivered weekly.

Unsubscribe with 1 click anytime. Zero spam.

* We may earn an affiliate commission from links in this report, at no extra cost to you and with zero impact on our benchmark data.