Dec 5, 2013 · 50m · mad

Fireside chat with Dwight Merriman // Data Driven #8 // Sep 2012 (interviewed by Matt Turck)

Dwight Merriman · 35m spoken Matt Turck · 6m spoken AB Mendez · 1m spoken Michael Selick · 35s spoken Andrew Dinsmore · 11s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this fireside chat hosted by Matt Turck at the NYC Data Business Meetup, 10gen co-founder Dwight Merriman discusses MongoDB's NoSQL database architecture, commercial open-source monetization strategies, and the growth of New York City's enterprise tech ecosystem.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 13.7% of the talking time here. How this is scored →

Matt as informed peer 2.2 Guest teaching 4.8 Guest disagreement 0.6 Matt pushing back 0.6
05100:0015:0030:0045:001:48–8:33 · Matt as informed peer 2/10 Understanding NoSQL and Horizontal Scaling Matt prompts Dwight with foundational questions on NoSQL and horizontal scaling. Dwight delivers an extended technical breakdown of hardware evolution, commodity cluster mechanics, and document-oriented data models.8:33–15:47 · Matt as informed peer 3/10 Comparing NoSQL Categories and Relational Use Cases Matt probes whether NoSQL aims to completely replace relational databases and asks for categorization of NoSQL types. Dwight clarifies that one size no longer fits all and categorizes graph, key-value, and document stores.15:47–20:00 · Matt as informed peer 2/10 MongoDB 2.2 Release and Future Product Roadmap Matt asks about the MongoDB 2.2 release and product roadmap. Dwight explains new technical features including the aggregation framework, concurrency lock improvements, and operational tool requirements.20:00–23:45 · Matt as informed peer 3/10 Open-Source Strategy and Business Model Matt asks about open-source business strategies and community building. Dwight lightly reframes the timeline by noting MongoDB was open-source from inception, detailing market size and revenue compression dynamics.23:45–26:38 · Matt as informed peer 4/10 Transitioning from Startups to Enterprise Customers Matt asks how MongoDB navigated transitioning from early web startup adoption to selling to enterprise clients. Dwight explains early adopter psychology and connecting enterprise needs to his past DoubleClick experience.26:38–30:15 · Matt as informed peer 3/10 Startup Opportunities in Big Data and B2B Matt asks where untapped opportunities exist in B2B and big data. Dwight highlights Sequoia commentary on B2B underweighting and cites surrounding ecosystem opportunities like BSON and analytics tools.30:15–34:26 · Matt as informed peer 4/10 Enterprise Tech Ecosystem and Hiring in New York Matt questions whether enterprise tech startups face fundraising barriers at the seed level in New York. Dwight reframes the premise, arguing that New York VC interest has caught up and engineering hiring is favorable.34:26–39:57 · Matt as informed peer 1/10 Audience Q&A: Security Features and Government Adoption Audience members ask about security features for defense contracts and handling semi-structured vs unstructured data. Dwight re-defines unstructured data to semi-structured and shares client examples like Telefonica.39:57–42:12 · Matt as informed peer 0/10 Audience Q&A: Aggregation Improvements and Hadoop Integration An audience member asks about batch aggregation and Hadoop integration. Dwight explains native MapReduce vs the aggregation framework and distinguishes Hadoop engine compatibility from HBase competition.42:12–44:43 · Matt as informed peer 0/10 Audience Q&A: Monetizing Open Source and Subscription Conversion An audience question asks how to optimize free-to-paid conversion in open source. Dwight delineates lead generation challenges from buyer willingness to pay for mission-critical support and subscriber features.44:43–47:23 · Matt as informed peer 2/10 Audience Q&A: Google Spanner and Distributed Transactions When an audience member asks about Google Spanner, Matt steps in to ask for a public definition before Dwight explains the fundamental trade-offs between distributed ACID transactions and horizontal scale.1:48–8:33 · Guest teaching 6/10 Understanding NoSQL and Horizontal Scaling Matt prompts Dwight with foundational questions on NoSQL and horizontal scaling. Dwight delivers an extended technical breakdown of hardware evolution, commodity cluster mechanics, and document-oriented data models.8:33–15:47 · Guest teaching 6/10 Comparing NoSQL Categories and Relational Use Cases Matt probes whether NoSQL aims to completely replace relational databases and asks for categorization of NoSQL types. Dwight clarifies that one size no longer fits all and categorizes graph, key-value, and document stores.15:47–20:00 · Guest teaching 5/10 MongoDB 2.2 Release and Future Product Roadmap Matt asks about the MongoDB 2.2 release and product roadmap. Dwight explains new technical features including the aggregation framework, concurrency lock improvements, and operational tool requirements.20:00–23:45 · Guest teaching 4/10 Open-Source Strategy and Business Model Matt asks about open-source business strategies and community building. Dwight lightly reframes the timeline by noting MongoDB was open-source from inception, detailing market size and revenue compression dynamics.23:45–26:38 · Guest teaching 4/10 Transitioning from Startups to Enterprise Customers Matt asks how MongoDB navigated transitioning from early web startup adoption to selling to enterprise clients. Dwight explains early adopter psychology and connecting enterprise needs to his past DoubleClick experience.26:38–30:15 · Guest teaching 4/10 Startup Opportunities in Big Data and B2B Matt asks where untapped opportunities exist in B2B and big data. Dwight highlights Sequoia commentary on B2B underweighting and cites surrounding ecosystem opportunities like BSON and analytics tools.30:15–34:26 · Guest teaching 3/10 Enterprise Tech Ecosystem and Hiring in New York Matt questions whether enterprise tech startups face fundraising barriers at the seed level in New York. Dwight reframes the premise, arguing that New York VC interest has caught up and engineering hiring is favorable.34:26–39:57 · Guest teaching 5/10 Audience Q&A: Security Features and Government Adoption Audience members ask about security features for defense contracts and handling semi-structured vs unstructured data. Dwight re-defines unstructured data to semi-structured and shares client examples like Telefonica.39:57–42:12 · Guest teaching 5/10 Audience Q&A: Aggregation Improvements and Hadoop Integration An audience member asks about batch aggregation and Hadoop integration. Dwight explains native MapReduce vs the aggregation framework and distinguishes Hadoop engine compatibility from HBase competition.42:12–44:43 · Guest teaching 5/10 Audience Q&A: Monetizing Open Source and Subscription Conversion An audience question asks how to optimize free-to-paid conversion in open source. Dwight delineates lead generation challenges from buyer willingness to pay for mission-critical support and subscriber features.44:43–47:23 · Guest teaching 6/10 Audience Q&A: Google Spanner and Distributed Transactions When an audience member asks about Google Spanner, Matt steps in to ask for a public definition before Dwight explains the fundamental trade-offs between distributed ACID transactions and horizontal scale.1:48–8:33 · Guest disagreement 1/10 Understanding NoSQL and Horizontal Scaling Matt prompts Dwight with foundational questions on NoSQL and horizontal scaling. Dwight delivers an extended technical breakdown of hardware evolution, commodity cluster mechanics, and document-oriented data models.8:33–15:47 · Guest disagreement 1/10 Comparing NoSQL Categories and Relational Use Cases Matt probes whether NoSQL aims to completely replace relational databases and asks for categorization of NoSQL types. Dwight clarifies that one size no longer fits all and categorizes graph, key-value, and document stores.15:47–20:00 · Guest disagreement 0/10 MongoDB 2.2 Release and Future Product Roadmap Matt asks about the MongoDB 2.2 release and product roadmap. Dwight explains new technical features including the aggregation framework, concurrency lock improvements, and operational tool requirements.20:00–23:45 · Guest disagreement 1/10 Open-Source Strategy and Business Model Matt asks about open-source business strategies and community building. Dwight lightly reframes the timeline by noting MongoDB was open-source from inception, detailing market size and revenue compression dynamics.23:45–26:38 · Guest disagreement 0/10 Transitioning from Startups to Enterprise Customers Matt asks how MongoDB navigated transitioning from early web startup adoption to selling to enterprise clients. Dwight explains early adopter psychology and connecting enterprise needs to his past DoubleClick experience.26:38–30:15 · Guest disagreement 0/10 Startup Opportunities in Big Data and B2B Matt asks where untapped opportunities exist in B2B and big data. Dwight highlights Sequoia commentary on B2B underweighting and cites surrounding ecosystem opportunities like BSON and analytics tools.30:15–34:26 · Guest disagreement 1/10 Enterprise Tech Ecosystem and Hiring in New York Matt questions whether enterprise tech startups face fundraising barriers at the seed level in New York. Dwight reframes the premise, arguing that New York VC interest has caught up and engineering hiring is favorable.34:26–39:57 · Guest disagreement 1/10 Audience Q&A: Security Features and Government Adoption Audience members ask about security features for defense contracts and handling semi-structured vs unstructured data. Dwight re-defines unstructured data to semi-structured and shares client examples like Telefonica.39:57–42:12 · Guest disagreement 1/10 Audience Q&A: Aggregation Improvements and Hadoop Integration An audience member asks about batch aggregation and Hadoop integration. Dwight explains native MapReduce vs the aggregation framework and distinguishes Hadoop engine compatibility from HBase competition.42:12–44:43 · Guest disagreement 0/10 Audience Q&A: Monetizing Open Source and Subscription Conversion An audience question asks how to optimize free-to-paid conversion in open source. Dwight delineates lead generation challenges from buyer willingness to pay for mission-critical support and subscriber features.44:43–47:23 · Guest disagreement 1/10 Audience Q&A: Google Spanner and Distributed Transactions When an audience member asks about Google Spanner, Matt steps in to ask for a public definition before Dwight explains the fundamental trade-offs between distributed ACID transactions and horizontal scale.1:48–8:33 · Matt pushing back 0/10 Understanding NoSQL and Horizontal Scaling Matt prompts Dwight with foundational questions on NoSQL and horizontal scaling. Dwight delivers an extended technical breakdown of hardware evolution, commodity cluster mechanics, and document-oriented data models.8:33–15:47 · Matt pushing back 1/10 Comparing NoSQL Categories and Relational Use Cases Matt probes whether NoSQL aims to completely replace relational databases and asks for categorization of NoSQL types. Dwight clarifies that one size no longer fits all and categorizes graph, key-value, and document stores.15:47–20:00 · Matt pushing back 0/10 MongoDB 2.2 Release and Future Product Roadmap Matt asks about the MongoDB 2.2 release and product roadmap. Dwight explains new technical features including the aggregation framework, concurrency lock improvements, and operational tool requirements.20:00–23:45 · Matt pushing back 1/10 Open-Source Strategy and Business Model Matt asks about open-source business strategies and community building. Dwight lightly reframes the timeline by noting MongoDB was open-source from inception, detailing market size and revenue compression dynamics.23:45–26:38 · Matt pushing back 1/10 Transitioning from Startups to Enterprise Customers Matt asks how MongoDB navigated transitioning from early web startup adoption to selling to enterprise clients. Dwight explains early adopter psychology and connecting enterprise needs to his past DoubleClick experience.26:38–30:15 · Matt pushing back 0/10 Startup Opportunities in Big Data and B2B Matt asks where untapped opportunities exist in B2B and big data. Dwight highlights Sequoia commentary on B2B underweighting and cites surrounding ecosystem opportunities like BSON and analytics tools.30:15–34:26 · Matt pushing back 2/10 Enterprise Tech Ecosystem and Hiring in New York Matt questions whether enterprise tech startups face fundraising barriers at the seed level in New York. Dwight reframes the premise, arguing that New York VC interest has caught up and engineering hiring is favorable.34:26–39:57 · Matt pushing back 0/10 Audience Q&A: Security Features and Government Adoption Audience members ask about security features for defense contracts and handling semi-structured vs unstructured data. Dwight re-defines unstructured data to semi-structured and shares client examples like Telefonica.39:57–42:12 · Matt pushing back 0/10 Audience Q&A: Aggregation Improvements and Hadoop Integration An audience member asks about batch aggregation and Hadoop integration. Dwight explains native MapReduce vs the aggregation framework and distinguishes Hadoop engine compatibility from HBase competition.42:12–44:43 · Matt pushing back 0/10 Audience Q&A: Monetizing Open Source and Subscription Conversion An audience question asks how to optimize free-to-paid conversion in open source. Dwight delineates lead generation challenges from buyer willingness to pay for mission-critical support and subscriber features.44:43–47:23 · Matt pushing back 1/10 Audience Q&A: Google Spanner and Distributed Transactions When an audience member asks about Google Spanner, Matt steps in to ask for a public definition before Dwight explains the fundamental trade-offs between distributed ACID transactions and horizontal scale.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 73.3% · guest 26.7%0:00 · Matt 73.3% · guest 26.7%3:00 · Matt 5.1% · guest 94.9%3:00 · Matt 5.1% · guest 94.9%6:00 · Matt 9.7% · guest 90.3%6:00 · Matt 9.7% · guest 90.3%9:00 · Matt 15% · guest 85%9:00 · Matt 15% · guest 85%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 3.5% · guest 96.5%15:00 · Matt 3.5% · guest 96.5%18:00 · Matt 18.5% · guest 81.5%18:00 · Matt 18.5% · guest 81.5%21:00 · Matt 28.8% · guest 71.2%21:00 · Matt 28.8% · guest 71.2%24:00 · Matt 23.3% · guest 76.7%24:00 · Matt 23.3% · guest 76.7%27:00 · Matt 11.4% · guest 88.6%27:00 · Matt 11.4% · guest 88.6%30:00 · Matt 23.8% · guest 76.2%30:00 · Matt 23.8% · guest 76.2%33:00 · Matt 6.9% · guest 93.1%33:00 · Matt 6.9% · guest 93.1%36:00 · Matt 0% · guest 100%36:00 · Matt 0% · guest 100%39:00 · Matt 0% · guest 100%39:00 · Matt 0% · guest 100%42:00 · Matt 0.3% · guest 99.7%42:00 · Matt 0.3% · guest 99.7%45:00 · Matt 2.9% · guest 97.1%45:00 · Matt 2.9% · guest 97.1%48:00 · Matt 7.3% · guest 92.7%48:00 · Matt 7.3% · guest 92.7%
Sharpest disagreement ▶ 33:08 Dwight counters NY seed funding narrative

Dwight directly pushes back against Matt's suggestion that raising seed funding for NY enterprise startups is a major barrier, stating that perception has already changed.

Hardest push from Matt ▶ 8:33 Matt questions NoSQL replacement scope

Matt directly challenges NoSQL positioning by asking if it claims to be a total replacement for relational databases or if relational still makes sense.

Biggest teaching moment ▶ 3:56 Masterclass on hardware architecture and horizontal scaling

Dwight provides an in-depth explanation of modern processor architecture limitations and why commodity multi-server clusters replaced traditional vertical mainframe scaling.

Matt holds his own ▶ 45:06 Matt intervenes to contextualize Google Spanner

Matt actively manages the interview by interrupting an audience question to ensure Google Spanner's paper premise is explained to everyone before Dwight answers.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Understanding NoSQL and Horizontal Scaling 2610 Matt prompts Dwight with foundational questions on NoSQL and horizontal scaling. Dwight delivers an extended technical breakdown of hardware evolution, commodity cluster mechanics, and document-oriented data models.
Comparing NoSQL Categories and Relational Use Cases 3611 Matt probes whether NoSQL aims to completely replace relational databases and asks for categorization of NoSQL types. Dwight clarifies that one size no longer fits all and categorizes graph, key-value, and document stores.
MongoDB 2.2 Release and Future Product Roadmap 2500 Matt asks about the MongoDB 2.2 release and product roadmap. Dwight explains new technical features including the aggregation framework, concurrency lock improvements, and operational tool requirements.
Open-Source Strategy and Business Model 3411 Matt asks about open-source business strategies and community building. Dwight lightly reframes the timeline by noting MongoDB was open-source from inception, detailing market size and revenue compression dynamics.
Transitioning from Startups to Enterprise Customers 4401 Matt asks how MongoDB navigated transitioning from early web startup adoption to selling to enterprise clients. Dwight explains early adopter psychology and connecting enterprise needs to his past DoubleClick experience.
Startup Opportunities in Big Data and B2B 3400 Matt asks where untapped opportunities exist in B2B and big data. Dwight highlights Sequoia commentary on B2B underweighting and cites surrounding ecosystem opportunities like BSON and analytics tools.
Enterprise Tech Ecosystem and Hiring in New York 4312 Matt questions whether enterprise tech startups face fundraising barriers at the seed level in New York. Dwight reframes the premise, arguing that New York VC interest has caught up and engineering hiring is favorable.
Audience Q&A: Security Features and Government Adoption 1510 Audience members ask about security features for defense contracts and handling semi-structured vs unstructured data. Dwight re-defines unstructured data to semi-structured and shares client examples like Telefonica.
Audience Q&A: Aggregation Improvements and Hadoop Integration 0510 An audience member asks about batch aggregation and Hadoop integration. Dwight explains native MapReduce vs the aggregation framework and distinguishes Hadoop engine compatibility from HBase competition.
Audience Q&A: Monetizing Open Source and Subscription Conversion 0500 An audience question asks how to optimize free-to-paid conversion in open source. Dwight delineates lead generation challenges from buyer willingness to pay for mission-critical support and subscriber features.
Audience Q&A: Google Spanner and Distributed Transactions 2611 When an audience member asks about Google Spanner, Matt steps in to ask for a public definition before Dwight explains the fundamental trade-offs between distributed ACID transactions and horizontal scale.

Statements from this episode (15)

Assertion Supported
Turck: NYC had only two billion-dollar exits in a decade, both DoubleClick
“New York has had two billion dollar exits in the last 10 years, and they were both double-click, right?”
Matt Turck Dec 5, 2013 ▶ 1:07
Insight
Merriman: NoSQL databases omit distributed joins to enable horizontal scaling
“Distributed joins is a hard problem, so the solution that's been chosen by both us with Mongo and the other NoSQL products is to say, you know, we're not going to do joins”
Dwight Merriman Dec 5, 2013 ▶ 3:23
Prediction Not checkable as stated
Merriman: Developers will use a few specialized database tools within 5-10 years
“My belief, I'm a little biased, but my belief is that if we kind of project out Say five years, 10 years, is that what we're gonna see is you'll have a few tools in your toolbox for databases.”
Dwight Merriman Dec 5, 2013 ▶ 8:57
Prediction Held up
Merriman: JSON document models will emerge as the standard NoSQL data model
“My belief, which is opinion, is that in, in the end of these NoSQL data models, that the document-oriented model will be the one that kind of converges on for standard, and specifically JSON style, or JSON based.”
Dwight Merriman Dec 5, 2013 ▶ 13:07
Assertion Supported
Merriman: MongoDB 2.2 eliminates global read-write lock to improve concurrency
“Historically in Mongo, sort of the concurrency in a single MongoD process on a single server has been suboptimal. And, ah, between servers it's been good, but inside a single server it needed a lot of work. So there was a global reader-irelock, which has been …”
Dwight Merriman Dec 5, 2013 ▶ 17:34
Insight
Merriman: Open source causes revenue compression relative to proprietary software
“You are going to get some revenue compression, right, by being open source versus not open source if you have the same number of users, right?”
Dwight Merriman Dec 5, 2013 ▶ 20:53
Opinion
Merriman: Startup founders are underweighting B2B enterprise opportunities
“He's right. So I think there is an underweighting of activity on B to B because there's plenty of opportunity there.”
Dwight Merriman Dec 5, 2013 ▶ 27:23
Disclosure
Merriman: 10gen finds engineering hiring easier in NYC than Palo Alto
“I'm actually finding right now, because we're hiring engineers both locations, that we're having a little more luck hiring engineers in New York.”
Dwight Merriman Dec 5, 2013 ▶ 31:12
Assertion Supported
Merriman: In-Q-Tel recently invested in 10gen
“We announced about a week ago that In-Q-Tel invested in TenGen, right?”
Dwight Merriman Dec 5, 2013 ▶ 34:52
Opinion
Merriman: HBase is more directly competitive with MongoDB than Hadoop
“I think Hadoop, or HBase, which is a Hadoop subproject, that's more of a, that's something that's more of an alternative or competitive with Mongo, where you would look at A versus B”
Dwight Merriman Dec 5, 2013 ▶ 41:07
Disclosure
Merriman: SNMP support is MongoDB's only non-free subscriber feature currently
“Like, I think the only thing in the subscriber edition right now that, that, that is non-free is SNMP support, for example.”
Dwight Merriman Dec 5, 2013 ▶ 44:00
Insight
Merriman: Fast distributed transactions across hundreds of commodity servers are impossible
“If we have a thousand server cluster running on commodity hardware, on a commodity network, and you want to touch data and mutate it that's on 600 servers, and then you want to commit that all or nothing with isolation and option to roll it back, and you want …”
Dwight Merriman Dec 5, 2013 ▶ 46:02
Assertion Supported
Merriman: MongoDB only supports single-document ACID transactions
“What we do currently in Mongo is we do these little microtransactions, which sort of, say, give you ACID properties, but only on a single document at a time, right? Because a single JSON document lives on one and only shard in a Mongo cluster at a given point …”
Dwight Merriman Dec 5, 2013 ▶ 46:23
Assertion Supported
Merriman: MongoDB's name is derived from the word 'humongous'
“Right, so the name is the middle of the word humongous, so it was kind of a play on this concept of scale, right, so that's where the name came from.”
Dwight Merriman Dec 5, 2013 ▶ 47:55
Disclosure
Merriman: Modifying Postgres wouldn't deliver 10gen's desired architecture
“We looked at the traditional stuff and we thought, well, you know, could we take something, can we take Postgres and start changing, you know, changing code and make something we like? And we felt like we would never, Land where we wanted to get to if we did t…”
Dwight Merriman Dec 5, 2013 ▶ 48:47
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.