Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Lots of lots of good stuff. What's a code red look like?
A A code red is when Toby sends an email to the entire company saying this thing is the number one priority, and he means it. From an actual operational perspective, what that really means is if the team that's working on this code red asks you to help, please drop what you're doing and help. This is the number one priority. So that's how they work. But usually a code red is a symptom of Some other much more systemic problem that's gotten you to that point. At least in my case, checkout code red in 2020 was, hey, the checkout is failing in, I don't know, three, four, five different ways. And the checkout is a pretty important part of Shopify, obviously. So let's go fix that. So we, me and, you know, somewhere between, I think at its peak, it was probably two, 300 people working on various parts of the problem. So we all like, Scrambled for a year and like did what it took to, um, fix those issues. But then at the end of that year, Toby took a step back and said, okay, well, why did that happen? Like, how did we get to the point where those problems even happened in the first place? And then that led to some of the reorganization of the company around Less, like there used to be 12 to 15 of these kind of fairly small fractured business units, and now there is actually only like three or four, which is actually truer to what the product is, but of course, each of those units is big…
AI assessment note: “A code red is when Toby sends an email to the entire company”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q that you and Toby and the team has done is actually shipped a lot of that quickly when a lot of very large organizations are like, oh, that makes sense. The outputs are non-deterministic, um, by nature. We have a process for measuring and evaluating quality of our products. Generally that process does not apply. Like, How did you get to this is good enough and we can ship it?
A Well, one of the principles that we currently hold, this might change in the future if confidence intervals go high enough, but one of the principles that we have for all of the Shopify magic features today is that they're allowed to propose changes, but not commit changes, right? So they can like generate text, but they're not going to save it without you actually reading it and saving it. We might suggest a reply for the user that's, you know, someone writes in like, Hey, what's your shipping policy? We can like suggest the reply based on us knowing what's in store and like running that through an LLM, but you have to hit enter, right? And so one of the things that like human is in the loop essentially right now is one of the, um, one of the ways that we're kind of mitigating risk here. And obviously human in the loop is great because it gives you the feedback cycle. You actually get a three part signal. You get which suggestions are accepted clean, which suggestions were accepted, but then with minor edits and then which suggestions were outright rejected. And that's An amazing loop to be able to improve things through. So we're always getting better, but that sort of human in the loop human must click save thing is a, is a big part of the strategy.
AI assessment note: “human in the loop essentially right now is one of the ways that we're kind of mitigating risk”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Is there anything external to Shopify? Um, that you think is especially interesting right now in the AI world, be it startups or things that people are working on or projects or.
A This is literally every single person has probably said this, but I think the, like the rabbit thing at CES was pretty interesting. Um, I mean, I'm, I'm literally wearing the shirt right now. Like I'm a teenage engineering nerd. So like I was, I was hyped on the hardware, but I think one of the interesting problems with like LLM based Applications and agents in particular is like, what's the actual interface they're interacting with, right? Like in the case of Sidekick, right? There's a couple different places. Like, like what is Sidekick using? Right? Is Sidekick actually using the admin API under the hood, or is Sidekick actually reaching into the pixels of the web app and clicking around in the web app and doing stuff, right? Like, that's a pretty important question. Like, what is the interface that the agent is actually learning and interacting with? And, um, I thought the really interesting part of the rabbit presentation was that they decided to treat the actual GUI as the interface and to try and like have a model that became very good at interacting with essentially web apps. If it actually works, the strategic brilliance of that is they instantly have the world's largest app store because the world's largest app store is just the web, right?
AI assessment note: “I think the, like the rabbit thing at CES was pretty interesting.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do you think that Shopify has a different, like, risk orientation than other companies of its size? You still need guardrails, and even if they're suggestions, like, it's not, it's not perfect, right?
A Correct, yeah. I think Toby, Toby and Kaz and myself, we're all, like, fairly risk-hungry people. I mean, I think that's just a little bit of founder culture is, like, you want to take risks. Like, you don't really want to be safe. You get bored. You get annoyed when things get too safe, but there's always this balance of like, we like taking risks that are risks to us. We don't like taking risks that could hurt, like break someone's business. And so if there's a way to do a thing that's like very risky for us, but, but is the risk is mitigated for the actual merchant. Like we're very for that. And I think that's where that kind of human in the loop, human must click save thing is part of that. Right. And who knows, like maybe a year from now, maybe two years from now, like we get to like, you know, Hey, you're, you seem to be accepting our suggestions without edits. 99% of the time. Do you want to just go into like full auto mode? Like maybe we get there eventually. We're not there yet, but like you can sort of see that the hill climbing will eventually take us somewhere close to that. Right.
AI assessment note: “Toby and Kaz and myself, we're all, like, fairly risk-hungry people.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q so how do you all think about integrating and acquired companies and acquired technologies? But more generally, as you build and build and build and build, there's a big munch of stuff. How do you sort of consolidate that down in a performant way that's actually speedy to execute? And then I guess later on, maybe we could talk about whether AI has played a role in that or not.
A Yeah, it's, it's a really good question. It, the answer is, it depends, like most problems. Usually when you encounter these problems, you sort of have to look at the stack. You usually don't want duplicates at the same layer of the stack, right? Um, and the funny thing about this is sometimes you actually don't notice that two things are in fact at the same layer of the stack until much later on, right? So as an example, Inside of Shopify, there is a, ah, there's the checkout, which takes, you know, someone's shopping cart and says, okay, let's take the business rules of the store and convert that to an agreement for a sale with taxes and duties and shipping and all that stuff. So that's one engine. And then at different points in Shopify's history, other things that have been built that are sort of like adjacent to that, but not quite the same thing, right? So like a thing for creating like an invoice or a thing for like even editing an order after it's been placed. And a lot of these things kind of get built because, because you start with a small problem. Example, I just need to be able to add a discount to an order After it's already been placed because, you know, someone calls up and says, Hey, I forgot to put the discount code on. Can you please enter it? Right. And so someone starts a project over here, which is, oh, we're just going to support editing discounts onto or…
AI assessment note: “You usually don't want duplicates at the same layer of the stack, right?”
Partly raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q things can go right or wrong based on that. And so I was curious a little bit about the broader concept of integration. Number one, from a team and individual perspective as a founder coming in, and why is it such a good home for founders? And then number two, like once you're acquiring the, the, the tech and the, the product, how do you integrate that as well? Yeah.
A It's really notable. So like when I look at the, so the, the, the sort of exact team of Shopify contains a number of founders, some of whom were acquired in and some, some of whom just were founders. And when I look at like my management team, the, the people who report to me, who each of these sort of nine or 10 people run R and D orgs that are like a couple hundred people each and almost like a very high percentage of those people are also founders, right? And so you ask, okay, well, why are all these people ending up essentially running the main parts of the product of Shopify and Shopify is a product first company. So that's sort of the same thing as saying running Shopify. I think a lot of this comes from, uh, Toby himself who. Has extremely strong opinions and has, I think, learned for the better to no longer be shy about like saying, I want to aim in this direction, right? I think there was like a big culture of bottoms up decision making, power to the edges, like whatever you want to call that school of management philosophy, which is like, you know, delegate and empower. I think a lot of former founders are actually coming back to the place where they say, Hey, what makes, what am I uniquely Good at. I'm uniquely good at, like, having a point of view, being willing to smash my head against a wall and, like, do whatever it takes to manifest that point of view in the wor…
AI assessment note: “why are all these people ending up essentially running the main parts of the product”