Startup Program:OcuFlow

Thestoryaspitched:aneye-surgeoncrisis,arealphasemodel,amovementskillmetric,andamuseum-stagefinale
StartLabs × OneAim Program — Sole Builder of the Product
13 surgical phases0.918 validation accuracy6 pitches, finale at 200+2026 (2 months)
Crisis
01The Problem

27 million eyes a year — and fewer hands to train

Cataract surgery is the most common operation on Earth — 27 million patients a year need it to keep their sight. And the math behind it is moving in the wrong direction: demand is climbing while the surgical workforce shrinks. Every future surgeon has to get good faster, with less senior time available to teach them.

+35%

demand in eye surgeries by 2030

from 700k/year today in Germany alone

−12%

ophthalmology workforce (US)

with comparable drivers in Germany

Sources: Berkowitz et al., Ophthalmology (2024), projection to 2035 · DOG (2025)

02The Status Quo

Surgical training still runs on a senior's spare time

How does a resident get better today? An attending watches and gives notes. That system has three structural problems — and none of them get better by having fewer surgeons:

Senior-dependent

Review is entirely dependent on a senior physician being available — the scarcest resource in the building.

Manual, on top of surgery

Reviews happen manually, squeezed on top of full surgical schedules. Feedback arrives late or not at all.

Inconsistent metrics

Every reviewer measures differently — outcomes across trainees, hospitals, and time are simply not comparable.

03Enter OcuFlow

Upload a surgery — get structure back, automatically

With OcuFlow, a cataract surgery video is structured into its surgical phases automatically — incision, viscoelastic, capsulorhexis, phaco, and on — so feedback can be anchored to the exact moments that matter. And this part is not a mock: a frozen ImageNet ResNet50 feeding a custom MS-TCN head (the published TeCNO recipe), trained on Cataract-1K — no turnkey pretrained model exists for this task. Below: the unedited output for one held-out case. Drag through it.

Model output — case 510418 segments · 11 phases · 15:55 surgery
Capsulorhexis

model confidence 96.1%

2:00 / 15:55

Drag the strip. Gaps between segments are idle windows (instrument exchanges) the model correctly leaves unlabelled. This is unedited output of the 4-fold-CV model — the same segments the demo plays.

0.918 ± 0.025

validation accuracy, 4-fold CV over all 56 cases

0.871 ± 0.030

macro F1 — above the paper's 0.78–0.85 baselines

Deliberately the published recipe — no class weighting, no augmentation tricks — so the numbers are defensible against the literature, not inflated. Trained locally on a MacBook M2 Pro: ~85 minutes, $0 cloud compute.

Per-phase F1 highlights

PhacoemulsificationF1 0.99

long, visually distinctive

CapsulorhexisF1 0.96

long, visually distinctive

HydrodissectionF1 0.95

long, visually distinctive

Lens PositioningF1 0.71

short, visually similar — the same hard case the original paper flags

04Economy of Movement

How much did the instrument move — and how smoothly?

Phases give feedback its where; instrument motion gives it its what. The pipeline computes, per phase, the motion metrics clinical studies already use for resident assessment — path length, velocity, direction changes, dwell time — plus SPARC, a single-number smoothness score from the movement-science literature.

It also computes per-frequency-band power (0.5–3 Hz, 4–7 Hz, 8–12 Hz). The hypothesis behind those features — that controlled fine-motor work and hand instability live in different bands — comes from the motion-analysis literature, and it stays a hypothesis until paired skill-score data validates it. Below is what can be shown honestly today: the real labelled tip track from case 5104, split with a simple low-pass into its smooth path and its high-frequency residual.

Instrument track — hand-labelled, 60 fps0 frames in window

Pipeline report — Incision I

Duration17.7 s
Path length0.74
Dwell time5.5 s
Efficiency31%
SPARC-18.34

Every point is a real labelled position of the instrument tip (in-house web labeller + Lucas-Kanade propagation). "Smoothed" is a 0.4 s low-pass of the same track; "Micro-motion" colors that path by how much the low-pass removed at each moment — part hand motion, part tracking noise, not a validated tremor measurement.

05Ground Truth

Four ways to get per-frame instrument positions

Cataract-1K only ships sparse masks (~every 5th frame). Motion analysis needs every frame, so four tracking/labelling approaches were built and honestly compared:

AGT-seeded Lucas-Kanade pipeline

the workhorse

Perceptual-hash GT frames back to their video position, seed tool tips, propagate LK re-anchoring drift at each GT frame. Batch-rendered tracked-motion video for 27 cases (3.3 GB).

BIn-house web keypoint labeller

built into the product

Manual sparse seeding + LK propagation + per-frame drag-to-correct, 5 API endpoints. Produced the 100%-coverage track powering the figure above.

CDUSTrack integration

research-grade fallback

Wrapped a semi-automatic research tracker in an isolated conda env with a converter into the project’s track schema.

DCVAT

evaluated & rejected

Dockerized, tested, documented why it isn’t enough: linear-only interpolation between keyframes can’t follow surgical motion.

06The Vision

Track your progress. Learn from the best.

Where this goes once paired skill-score data lands: every surgery a resident performs becomes a data point on their own curve, and every phase can be laid beside an expert's. Both views are built and interactive today — with illustrative data, shown exactly as they were pitched at the finale.

07The Business

A $7.5B market, entered at $9 a case

Surgical-training software sits inside a $7.5B market. The wedge is a $9 pay-per-case entry — low enough for a resident to expense — growing into the real revenue driver: $15–30k/year SaaS per clinic. The positioning claim from the deck: existing options are either affordable but passive (content channels) or active but expensive (Zeiss, simulator hardware). OcuFlow takes the empty quadrant.

The finale closed on the ask: a €1.2M pre-seed and clinical pilot partners — pitched to 200+ people at the Ägyptisches Museum.

Sources: Grand View Research (2024/25); SOM = 2% of Europe SAM · CAGR 5.7% to 2033

Market — as pitched
TAM $7.5BSAM $2.3BSOM$46M
Positioning
AFFORDABLEEXPENSIVEACTIVE LEARNINGPASSIVEOcuFlowContent channelsZeissHaag-Streit sims
08The Team

Three people, six pitches, one codebase

OcuFlow was a three-person team inside the StartLabs × OneAim program — two months paced by six pitches, ending on the museum stage. The pitching, positioning, and research were shared. The shipped product was not: every line of the platform — frontend, backend, model training, labelling tools, deployment — is mine.

David Vogenauer

AI & Computer Vision

Research direction on the vision pipeline

Chantelle Davis

Business & Go-to-Market

Market sizing, positioning, clinical outreach

Andrei Zitti

Product & Engineering

Sole builder of everything that runs — including this page and the demo it embeds

09Post-Op Notes

Honest scope & limitations

  • Phase recognition is real and validated on 56 Cataract-1K cases. Dataset licensing and public-asset constraints limit what can be bundled directly in the portfolio.
  • The economy-of-movement metric is implemented and runs, but its skill-prediction value awaits paired ground-truth skill scores (Cataract-LMM — access requested).
  • The progression and peer-comparison views are product concepts with illustrative data; demo reference medians are not validated cohort norms.
  • All ML trained on a single public clinic dataset; cross-clinic generalization is still unproven.
  • A weekend-MVP-turned-research-platform — not a regulated medical device. Market figures are as pitched, not audited.

Product built end-to-end by Andrei Zitti

OcuFlow · Smarter eye surgery training · StartLabs × OneAim, 2026 · with David Vogenauer & Chantelle Davis