Data moats and flywheels
LAST REVIEWED 2026-08 · SOURCED FROM 2 SESSIONS, APR–MAY 2026
Why data is the moat
Section titled “Why data is the moat”A deep tech GP’s framing: startups can’t out-algorithm the labs or out-compute the hyperscalers — data is the one ingredient a small company can genuinely own. And it doesn’t have to be exotic. The example told in the room: a high-schooler who scraped a retail-trading community’s ticker chatter for three years and sold it to hedge funds — unassailable, because history can’t be re-collected.
Two repeatable sources of differentiated data, from a seed investor whose thesis is exactly this:
- Collect what nobody publishes. Do the unglamorous work of gathering data that exists in the world but not in any dataset.
- Decades of domain fluency. Knowing how customers actually use data — what’s signal and what’s noise in a niche — is itself proprietary.
The corollary from the roundtable: “alpha is everything; ordinary data gives no signal.” Commodity enrichment sources are table stakes, not a moat.
Building the flywheel
Section titled “Building the flywheel”The mechanics that came out of the practitioner roundtable:
- Make feedback pay instantly. Users label your data when doing so immediately improves their own experience — that instant reward, not goodwill, is what powers explicit feedback loops.
- Interview your ICP about what’s missing. The people doing delivery and support know which context the dataset lacks; mine call recordings the way teams “study film.”
- Back-test creative sources. Permit and remodel records against subsequent outcomes, demographic age-out data, web-scraped alternative data, satellite imagery — the roundtable’s real-estate examples all shared one property: a testable predictive claim.
- Small data is enough to start. Learning is roughly log-linear: a small dataset surfaces the strongest signals and directional insight; granular personalization is what needs scale. Don’t wait for big data to ship the first loop.
Contracts feed the flywheel
Section titled “Contracts feed the flywheel”The first-customer playbook connects here: deep early discounts traded for utilization minimums and feedback obligations are how the flywheel gets its first turns — see First customer contracts.
For compliance-grade products, the roundtable’s reliability pattern: multiple models cross-checking with a tiebreaker, tasks decomposed into micro-steps with intermediate outputs as an audit trail, and human-in-the-loop encoded as a company value rather than a patch.
Pitching the moat
Section titled “Pitching the moat”Investors don’t want the nuance. The roundtable’s framing advice: “we’re building the largest dataset in X; it trains the best predictive models; customers stay because the product keeps getting better” — a network effect in one breath. If you keep having to explain why the data compounds, you’re pitching the wrong investors.
“It’s much better than… trap them, then please them. Right? Trap them in the correct way.”
— Managing partner, deep tech seed fund · session, Apr 2026, on products users can’t leave
Sources
Section titled “Sources”Two sessions: an AMA with the managing partner of a deep tech seed fund on defensibility (Apr 2026), and an in-person member roundtable on proprietary data and flywheels (May 2026). Roundtable material is presented without individual attribution.