The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Mikhail Parakhin no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I reached out was because you started, Promoting more sort of internal tooling, uh, primarily Tangled, but also a lot of people have seen and adopted Toby's QMD. Uh, and obviously I think, uh, Shopify has always been sort of leading in terms of engineering. I think more, it's just more recent that you guys have been more vocal about your sort of AI adoption. Is that, is that true?

A Well, I think AI tools in general are fairly recent development, uh, and, uh, we, Shopify, you know, at this stage of its development, we're developing AI in-house and building tools that use AI and, you know, interfacing with the wider AI community, you know, are on the sort of the runaway trajectory. So it just did by sort of natural byproduct. We talk about it more also. We just, uh, just even yesterday, Andrej Karpathy was famous in tweeting about, oh, there's some, uh, ways that you can organize your agents to store the data and then look up the data so that you don't have to research or lose context every time. And a little bit tongue-in-cheek, I tweeted that, hey, we've done it much earlier, and we even have different approaches, Toby and I. Toby, of course, is a big fan of QMD, and I'm More of a SQLite fan, but, yeah, very similar things that we've already done here. The point is, yeah, we're very dynamic, you know, explosively growing company, and we have to be at the forefront of AI adoption, obviously.

AI assessment note: “So it just did by sort of natural byproduct. We talk about it more also.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q think SimAI, I think Yongjun Park who did the Smallville thing. There's a very small cottage industry of people trying to do like the simulated customer thing. I think a lot of people maybe don't super trust this yet because they're like, well, obviously they would just do what you prompt them to do, right? But maybe just think, uh, tell us about the sort of inspiration or origin story.

A That's exactly actually the thing I wanted to cover because if you don't have the historical data, All you can do is prompt agents in the vacuum, and they will do exactly what you prompt them to do. In fact, when I first proposed it, and this is a bit of a, my brainchild initially, if I, I can boast. And then Toby said like, but wouldn't they, they just repeat what, what you tell them? And, uh, but I'm like, yes, except Shopify has decades of history of how people made changes and what there is, uh, what it resulted in terms of sales. So now what we can do is we can, we have this, it's not, it's a noisy data. These are small, usually websites, uh, you know, like things, things are never in isolation. It's almost never AB experiment. It's always AA experiment when there's, has two meanings, but basically, you know, in different time, you run two different things. But if you aggregate in general, uh, like everything together and you apply, uh, denoising and collaborative filtering like approach, you can extract a very clear signal and then you can optimize Your agents, and that's why it took so long. It took almost a year of that optimization of just us sitting and fiddling, and, and we had this internal goals of correlation of heating. Internal goal was to hit 0.7 correlation with, uh, add to cart events, for example, like that, that if we run real A, B test experiment, that it …

AI assessment note: “when I first proposed it, and this is a bit of a, my brainchild”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Uh, okay. Yeah, but, I mean, we, we can, we can discuss the, the, the release briefly, because we'll release this after the, after it's already announced, or whatever. There's a catalog that you guys are doing?

A Yeah, so we are, we are, We are bringing in capabilities of a whole Shopify catalog. Basically, you now, you can search for products. You can do lookups by specific ID. You can do bulk lookups when you need to bring multiple products. You don't need to know in advance what you're trying to show or to sell or check out. Like you can now, you can now have this decided at runtime and this big area for investment for us For both non-personalized and personalized searches, trying to provide basically a window into whole universe of products that are being sold everywhere in the world. And Shopify is really not exactly, but almost like a superset of anything being sold. Now we're bringing it into UCP and, uh, and, uh, identity linking is another big thing for us so that you, you can use, uh, Like Google or whatever, whatever identity you have, uh, they're minimizing a friction.

AI assessment note: “we are bringing in capabilities of a whole Shopify catalog. Basically, you now, you can search”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q worked on. I'm curious about the backlog, right? Like, the, the, the, I actually don't mind a pro-level model taking an hour, two hours to review my PR, because I've dealt with humans who take a week to review my PR, right? And I keep pinging them on Slack, hey, hey, review my PR. So, you know, I think there's some trade-off here where, like, it still doesn't make sense.

A Exactly. That's exactly my point, uh, that on one hand, you can tolerate longer latencies at PR. On the other hand, like right now, the real problem is not in Spending time waiting for PR is real problem is since there's so much more code, then, uh, probability of at least some tests failing, going up, and then you, like, keep failing, then you have to find the offending PR, evict it, retest it without that PR, and so deployment cycle becomes much longer. Uh, so it actually, in terms of the overall time to deploy, it's total time savings if you spend more time on a longer model, like, thinking for an hour, because then, then you, you don't have to spend all that time During testing and rolling, you know, rolling back the deployment.

AI assessment note: “total time savings if you spend more time on a longer model”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Is that a now only thing, or do you think you give Liquid a hundred billion dollars, and they will, is it, is it just more scale, or like, what, what is limiting it? You know, what prevents it from running into the same issues that SSM's had?

A Their scale is already much larger than the largest SSM I, I'm aware of. Uh, and so, yeah, so, uh, SSM was just not expressive enough, or in my opinion, like, um, again, I'm sure I'll get a lot of pushback and probably I'll share with this all, but in my opinion, SMs are not expressive enough, and, uh, liquid models are. I think, uh, especially in their hybrid form, when combined with, uh, Transformer, like in Mamba fashion, they probably the best architecture I'm aware of, like, period. But of course, Liquid AI is not at the scale of, uh, you know, Anthropic or Google or OpenAI in terms of compute. So I don't think they, I think if, if they, uh, if they had similar level of compute, they, they would be very competitive and maybe even beat the, uh, the largest models, at least from what I've seen. They don't have, uh, this level of, uh, investment, but they still have decent investment, and, and it's, uh, it's definitely for this scenario of smaller models and distilling into their second to none or very often. We're very, very omnivorous, and we own purely merit-based, so the moment they will start being competitive, like, we will switch to something else, and we constantly test, but, but so far, if you see, Progression, if I draw a graph of our workloads on Liquid versus our workloads on, I would say, Quen, which is another awesome model and probably another kind of standard …

AI assessment note: “if they had similar level of compute, they would be very competitive”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q has. Uh, but you know, I, I think some, some, Pretty cool stuff that you guys are working on is your ML experimentation, uh, and your, your sort of auto research training pipeline. Presumably you're much closer to this one because it's, it's a sort of personal hobby of yours. How would you explain them together? I thought we have a slide that like, uh, has this, the system diagram.

A Yeah. Tangle first and then tangent is a, is a thing on top of tangle. And, uh, Tangle is the third generation claim of, uh, systems of, uh, running any data processing, but a bit with a skew for ML experiments, but not necessarily any sort of data processing tasks where you need to iterate, share, and you have scale so that you want maximum efficiency. You know how, like, normally you would work. You would imagine you're a data scientist or an ML practitioner. You would get Jupyter notebooks, or, or maybe you would get, uh, you know, Python, your Python scripts, and you would munch the data, and you produce those TSV files, and you put them in some GFS or something, then you would notice that, oh, it has this, uh, weird missing values. You go and write another script that goes and replaces them with the dashes, and then, then you, then you run some, some, uh, oh, I need to filter bots, and so you run some, like, GBM model that, uh, removes the bots, and then, Then, like, and then you kind of, like, get into shape, and then you start experimenting, and you run multiple experiments, and then you're like, oh my god, like, this experiment is worse. You undo, and you cannot get to previous result, and like, what did I do? Like, and then, then you finally, like, get everything working, and then, like, start throwing it over the fence to production. You, you replicate it, those thing…

AI assessment note: “Tangle first and then tangent is a, is a thing on top of tangle.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.