The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

1,847exchanges match
1,797on raw tape
133redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q bearing in mind that, uh, you're a young startup and, uh, in a super fast moving industry, uh, how does that all work? Like those two pieces, how do they work together, including, uh, from a go to market perspective? Because GPD for all presumably is more of a developer audience. Is, is Atlas more of a kind of like enterprise audiences hugging face, like an enterprise customer for you?

A Yeah, so Atlas is definitely a B to B enterprise tool. Um, we make it, you know, very, we have very generous limits for individuals, like power users, academics as well, especially, um, to make it so that they can leverage it, but it is kind of the core growth engine behind the monetization of Gnomic. Uh, and I think that puts us in a really, really interesting position relative to a lot of other open source companies, because it means that we can keep GPT for all as this, like, Very pure, like, love letter to the community open source, uh, sort of project. Um, and we won't sort of be pressured to eventually find a way to kind of, like, squeeze it, squeeze it for dollars. Um, but I think the two interact also very well from a, from a funnel standpoint. So the kind of person that is interested in tooling to help them understand and curate massive unstructured data sets is the kind of person that's either one, using that to train models or two, generating a lot of that kind of data. And That's exactly the kind of person that is interested in, in tools like GPT for all. And so, you know, a lot of the, you know, first enterprise sales that we had with Atlas were inbound in the GPT for all discord. And that's, you know, maybe this is just what building a company in 20, 24 is, but our sales funnel is literally like Twitter to discord to enterprise sale. It's insane.

AI assessment note: “our sales funnel is literally like Twitter to discord to enterprise sale”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay, great. So maybe bring it to life for us. If I'm an agent and I interact with ASAPs technology, what do I do? What's happening?

A I'll give you an interesting example in the messaging side. For example, if you're chatting with a company, it takes on most industries, the agent roughly 20 seconds to type a response to whatever utterance you've sent their way. If I can predict what you should respond and instead of having to type that sentence, you just click on it, it takes, it goes from 20 seconds to roughly a second. And then the next logical question is that sounds wonderful for how many of my agent responses can you actually predict the right thing where they're going to go from clicking something that takes a second versus typing something that takes 20 seconds. And today in our most mature customers, roughly 80% of everything an agent does is click on those suggestions that the system's generating. So it's, I mean,

AI assessment note: “instead of having to type that sentence, you just click on it”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q Anything in the labs or roadmap you can talk about? I know you have this concept of sneaks where you can see some of the stuff that's brewing.

A Yes. We did show some of our generative video stuff and three D stuff. But one of the things I'm most excited about is we did a sneak of something that we call just the Firefly editor. Very creatively new. But the Firefly Editor is pretty wild because you have an image and there's, it's an image of a family, right? And so you could always select the five-year-old boy, but if you move him, well, there's going to be a blank spot behind him, right? But what this does is actually allows you to select objects and move them. And it generates all the pixels behind whatever you're moving in real time to fill in what would be there. And so it's like this crazy, you feel like everything is layers, but you captured every aspect of the image in endless layers because you can just do anything with it in real time. And it's one of those experiments where we were like, wow, this is going to change image editing forever. People are just going to just move their kids around in family photos, and it's going to be effortless. It's intensive from a generation's perspective, but it's the future.

AI assessment note: “we did a sneak of something that we call just the Firefly editor.”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q So moving on from the Adobe hat and putting your thinker and writer hat, I love all the content you've created and stuff you've written about the intersection of AI and creativity. Maybe give us some summary of some of the thoughts. What do you view as the future of how we work and how that gets impacted by AI?

A Yeah. Well, I've had a lot of discussions as you can imagine over the last year or so of people You know, asking what does this mean for the creative industry? What does this mean for advertising? And we're entering a world where, first of all, every brand's gonna absolutely flood the zone with content. That's gonna make the bar go up for the digital experiences that really engage us. Whenever you talk to any great creative and you ask them how they come up with a great solution to a problem, they always say it's a function of time. I mean, if you give them more cycles, they'll come up with more possibilities and then they'll find the better three to present to their client at the end and get a choice. So if you give them, you know, two X the time, they'll do two X the exploration of the surface area possibility, and they'll come up with better solutions. AI basically truncates the time. So what we're seeing in our products is that instead of an illustrator testing three color palettes for a packaging for a perfume product, they'll test, you know, 30, and then they'll have so many better choices to choose from. So I think that AI is going to raise the bar. Now, when you get more Capacity out of a creative professional. Does that mean you need fewer of them? Or does that mean you actually want more of them? Well, in the last 20 years, engineers have gotten a lot more productive …

AI assessment note: “AI basically truncates the time. So what we're seeing in our products is that”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So Glean, I read somewhere, one should think of it as build.com, but with a brain. Is that the idea? So meaning that you build both the sort of core workflow infrastructure of build.com, but also add AI.

A Yeah. We took a first principles approach to say, listen, if we're collecting all this data, what else can we be doing? So in addition to standard, what's called like AP automation, which is bill comes in, invoice gets extracted. And then you can sync it to your accounting system and get it approved and get it paid. We were like, all right, well, how about we do vendor intake? So if a team wants to bring on a new vendor, let's have a flow for that and, and, and support approvals for that. If you want to set up a budget for your vendors, let's set up monthly budgets. So when the bill comes in, we're comparing it to a budget, not just to last month, but we can get an alert that day that you were over budget. And every bill is like, you can conduct a variance analysis to any prior bill. To see what changed. So there's a lot of functionality that, again, took like a first principles approach to say, how do we just be much more strategic in terms of managing and optimizing our vendor relationships and spend?

AI assessment note: “Yeah. We took a first principles approach to say, listen, if we're collecting all this data”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yes. So how did you get started? What was the first thing that you built and when did you start the company?

A Yeah, so I had the idea in late 2019, we had our, like, pre-seed round in early 2020, like, right as COVID was kind of happening. We started building just like models to ingest documents. So from like a, an AI perspective, what we need to do is when we receive a document, like determine what type of document is it? Is it an invoice? Is it a receipt? Is it a billing statement? We can gather intelligence from those types of documents. If it's a contract, maybe it's like an NDA, maybe it's a, it's like a remittance slip. That's not something that Is going to get processed. So we have to determine what type of document it is first. Once it's a document, like we know it's a document that we can analyze, then who is the canonical vendor? So we have to build models and this is all done with like NLP models, like back in the day before like LLM existed. But like who, you know, who's the canonical vendor? We had to differentiate between Google ads versus like Google workspace versus Google cloud and like all the different taxonomy that exists with that. At the vendor level, then, you know, extracting all the elements off of a page. And it could be invoice date, due date, all the various fields. But the thing that made us very unique is we're extracting all the line items too. So what are you purchasing? For what dates? What's the unit price? What's the quantity ordering? And, you know, …

AI assessment note: “I had the idea in late 2019, we had our, like, pre-seed round”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And so you build those models initially, and then you added some LLM, commercial LLMs. I mean, I, I guess, how did you evolve the, your machine learning stack as this whole generic AI popped up?

A Yeah, initially we started with using like OCR vendors on the extraction piece. A lot of the mapping models were, were models that we developed in house. And then over time, We saw that the predictive value of the OCR vendors were no longer adding like the information value we can get versus like an ensemble model of like proprietary models were not worth the cost of maintaining those relationships. So it became like a hundred percent proprietary again, like an NLP ensemble model. But over the course of the last 12 to 18 months, we've done a hundred percent shift to LLM modeling for the complete stack. And we're using Vertex and OpenAI. And we're starting to experiment with Claude.

AI assessment note: “we've done a hundred percent shift to LLM modeling for the complete stack.”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay, great. So what would you say are the specific challenges when evaluating and monitoring LLMs? Obviously the category of evaluation and monitoring in the software world has yielded huge companies like Datadog in particular, but that's one world. Like how is the world of LLM different? What are the specific challenges and opportunities?

A Yeah. So machine learning differs from traditional software in a couple of key ways and then generative AI as a kind of subset of machine learning differs even further. So the first transition you go to when you go from kind of traditional code to machine learning is that it's no longer deterministic, right? We're used to, for software engineers, writing a program, you run it, you get the same results each time. You can write a deterministic test and people are doing performance monitoring with something like Datadog, but they're not expecting that when they run the code each time, they're going to get different outputs. Once you move to the world of machine learning, now it's stochastic. And not only that, but you're now specifying what the program does via a data set and a training process rather than, you know, deterministically in code. And so evaluation and machine learning focuses on accuracy metrics and things like that. And when we go to LLMs and generative AI, we go one step further where the use cases that people are applying these to are very general and very subjective often. To pick a couple of examples, you know, if you're helping someone draft a sales email, Um, or you're writing marketing copy. There isn't any longer a kind of ground truth answer that you can compare against for the model to know whether it's correct or not. Even if you're doing a question answe…

AI assessment note: “There isn't any longer a kind of ground truth answer that you can compare against”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q So how are you solving the problem? Maybe taking those three in turn, what human look do to address the issue.

A Yeah, so we provide the ability to get evaluation data, both from human feedback in various ways and also in an automated way. So I maybe take the human feedback version first, and this is actually the first version of human loops product started doing this, which is that because it's very subjective, there's some sense in which the only real ground truth is your customer satisfaction or opinion of the experience they get out of the product. And so we make it really easy for people to instrument their applications with the ability to capture feedback. So the simplest version of this is things like thumbs up, thumbs down that you might have seen in many applications. You see it in chat GPT, but people also tend to collect implicit signals of user satisfaction. So the actions that they take after interacting with a particular LLM app. And also if people are able to edit any generated texts or generated content, then also capturing those edits can be a very useful signal of how well things are working. And then in the human loop app, we triangulate those sources of feedback. Back against, okay, what model created it? What inputs created it? And we give product teams the ability to dig into that data, analyze it, understand what's working well and isn't and why, and then critically to also in the same application, take actions to change things, to make them better. So to edit promp…

AI assessment note: “we provide the ability to get evaluation data, both from human feedback in various ways”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q And hopefully not an unfair question because Datadog has built, we know a multi-billion dollar business, but let's assume human loop finds regression or, you know, different results from the same prompt. Then what can you do as an enterprise? Obviously knowing that this is an issue is essential, but is there a way you can fix that or you should just be aware?

A No, it's, I think this is one of the strongest arguments for why a new set of tools is needed and why using, you know, existing platforms for monitoring and observability is less productive. And it's because with generative AI, the speed with which you can make interventions is extremely high. And having that in one combined platform is actually one of the powers of this. So you're absolutely right. So within human loop I mentioned, we have this kind of interactive environment, both for, for prompt engineering And we also have the ability for people to fine tune models, which is where you do a little bit of extra training on a new data set. And so what will often happen is people will find a bug, you know, they'll basically like be exploring the, the log data within human loop. They'll find an issue. They'll reopen those data points back into that interactive environment where they can now run what if style analysis. So they can change the prompt or they can change the information retrieval system a little bit. And see what the impact was. If they're able to then fix that issue, they then run a regression test. They say, okay, is this new prompt still performing well on what worked before? And if the answer is yes, they can actually promote it straight to development or production from within that system. And so actually a product person or a domain expert who's able to go in a…

AI assessment note: “They'll reopen those data points back into that interactive environment where they can now run”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q own evolution. And as a result, it's a little, uh, a little blurry who does what. So any thought there would be very helpful to in particular, Frameworks like lane chain. Where, where does that fit? Are you a competitor? Are you a partner? And then a JGPD enterprise or what OpenDI does in terms of like getting further into the enterprise. Same thing is that a competitor, friend, foe.

A Yeah. So I think at base you have the foundation model providers, right? So here you would have OpenAI, Cohere, Anthropic, Mistral, whoever else it might be, the open source models, Llama. And then at the other end, you know, the other end of this spectrum maybe is the applications, which are the end user facing applications. And human loop is kind of a layer that sits in the middle. So we're model agnostic, we have close partnerships and we'd definitely be friends with all of the foundation model providers. We're keen to basically help their customers get to value. LangSmith and LangChain. LangChain is like a orchestration library. So this is basically just a set of utility tools for people who are writing the code around an, an AI application. And it has a whole bunch of helper functions built in that help them get started more quickly. So that wouldn't be sort of directly competitive. Their, their LangSmith product probably has a little bit of overlap with us, but it is not, you know, it has some overlap that is fundamentally, I think focused more around monitoring chains and agents. And then where I think we are focused is for enterprises who are building LM aspects where typically there's a lot more collaboration required. This becomes a team sport. And also where the need to have guardrails and evaluation starts to become much more significant because they're operating at…

AI assessment note: “we'd definitely be friends with all of the foundation model providers”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q know, some of the key debates in the industry right now. So certainly as we are recording this, there seems to be a tweet every second on open source, the closed source models and, you know, people that feel very strongly about preserving a very Free open ecosystem and others that are more concerned about security risk. Any, any, any thoughts on that debate? Where do you land? Just curious.

A My natural inclination is always to be in favor of open source, right? Just as a default knee jerk reaction, we've got so much benefit from open source in general. The software world is built on top of it that, that I always start from a position of optimism about open source. I understand though, some of the reasons why people have Safety concerns around larger models. There is an opportunity for misuse. And I've seen people like Jan LeCun say, oh, but we already have search engines. So like, you know, people can look up with a search engine, how to build a bomb or how to, you know, build a bio weapon. Like why do LLMs make it worse? But I think that really does downplay like how much better they are at synthesizing information and explaining steps to you. And there is a dramatic reduction in how hard it is to do certain forms of misuse. And that's before we get to the more safe, you know, the safety concerns about AI that actually is misaligned. But to me, the question is like, it's very difficult to know when to, like how to solve this or when to put restrictions in place. People, when GPT-II came out, they didn't release GPT-II for fear of misuse concerns. And then GPT-III, and now we're on GPT-IV. It would have been really sad for the world, I think, if at the point of GPT-II, we had decided, hey, you know what? This is too dangerous. No one can have access. Cause we would…

AI assessment note: “My natural inclination is always to be in favor of open source”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What was your journey into starting the company? I read somewhere that your dad is an oncologist, like you grew up in a health family. What was your journey?

A Yeah. Uh, so before DeepScribe, um, I was a research scientist at Bayer, which is Berkeley's AI research lab. Um, and our lab was very first principles, um, in terms of AI, because it was on the, um, on the border between the stats department and the engineering department, as well as my mentor, Jamie Murdock, who had actually written, uh, one of the first papers on interpretability when it comes to natural language processing. So, um, we were all about data labeling, data curation, Um, and it sucked for me as an undergrad student, because all I wanted to do was train large models on a lot of data. Uh, but, you know, my mentor took me aside and was like, we're gonna start small, and so that was my background. The way I actually got pulled into healthcare was through my dad, who's an oncologist, and I would say as a kid, I was actually desensitized by the problem of documentation, because I just assumed it was a way of life for my dad to spend the evenings catching up on notes, or on the weekends, missing important Life events because he had to, he was hitting that seven day mark in which health systems required you to complete your documentation by. So as a kid, couldn't really do anything about it. Just accepted it. All I could do is, um, get him onto the latest, uh, software, which at the time was Nuance and their dictation tools. Uh, so that was 10 years ago. Uh, and for tho…

AI assessment note: “The way I actually got pulled into healthcare was through my dad, who's an oncologist”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay, great. So that's the speech to text part. What's the next bit?

A Um, so for the actual summarization, um, we've, we've filled with this over time, but the, um, the one that gets us the highest accuracy is actually three separate, um, three separate models. So the first is our own classical models that have been trained on all of our data. And we have a little over two and a half million conversations right now. They're all labeled. And, um, so basically we have a stack of classical information extraction techniques, um, paired with our own in-house LLM that, um, we've fine-tuned and are currently in the process of pre-training. And then we have our, uh, we have GPT-IV that's also used. And so, um, depending on the, the task, we will either use one, two, or all three of them and see whether they agree or not. And by doing that, we have A way to validate the output, but then also, um, but then also leverage the non-deterministicness of language models that makes them so good.

AI assessment note: “for the actual summarization, um, we've, we've filled with this over time”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Um, you announced recently the beginning of this month, uh, the customization studio. Uh, what is that?

A Yes. Um, so customization studio is, um, probably the most important facet of DeepScribe's product. Uh, so in this AI age, I think it's fairly easy now to record a conversation And generate a node with GPT-IV. Um, but what really makes DeepScribe different from a lot of those solutions is the ability to conform to nuanced workflows for clinicians. Um, especially when it comes to the higher revenue generating folks like specialists, high patient volume, because for them, every single second of documentation time matters a lot. So customization studio gives clinicians about 35 different ways to, um, transform their note to how they like it. So we can natively fit into most of their workflows. We can collect discrete fields from their conversations. We can, um, change the style of the writing, and we put that all into clinicians hands. So previously, In healthcare, clinicians haven't really had a good way to configure and train their own models without it looking like a black box. So, uh, this interface now allows them to, with a few clicks of a button, uh, change how they like the note. And a lot of it is enabled by, uh, some of the advances we've seen in LLMs that allow you full control over the style of language, um, which has been, uh, the, the big breakthrough that enabled customization studio.

AI assessment note: “customization studio gives clinicians about 35 different ways to, um, transform their note”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay, super great. So to unpack, um, some of this, uh, so you mentioned retrieval augmented Fine tuning. Uh, is that, is that the same thing or is that different from, uh, retrieval augmented generation? Is that in terms or different?

A Uh, great question. They are two different things. So we, we have, you know, simple SDKs on top of our system for RAG, retrieval augmented generation, which is very common, um, these days. Um, retrieval augmented fine tuning is taking that and to the next, next level and to incorporate that into the training process. So retrieval augmented generation, RAG, which is very commonly used today, is a way to actually get information in, um, at inference time, um, during prompt engineering, essentially. Um, but for the model to learn new knowledge, retrieval augmented fine-tuning is actually incorporating that retrieval technology into the fine-tuning process. And something that I'm very, very excited about is These two very big communities kind of bring being moved together. So one community is, you know, the AI large language model community, and the other community has, you know, decades of research on information retrieval. Um, and I'm very excited to, you know, see these two communities really converge to make these models more powerful.

AI assessment note: “They are two different things.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And you went back to open AI upon graduation, right?

A I went back because I, I couldn't start any company after my PhD, um, either because the idea was too grand that nobody would want to come work with me or, uh, I had visa problems to start it all by myself. So I felt the best thing to do was like find a job and like continue to like keep exploring and learning more. But around the time in like maybe February or March in, like news started spreading that there were companies like Jasper and copy.ai that started making more money than even OpenAI at the time. And that was amazing. Okay. Like people are building products. This is real. And, uh, GitHub Copilot Uh, when, when, when they moved away from the waitlist to the paid version, they just had like hundreds of thousands of people paying from the first day. And so that all made it clear that, um, stuff that we were all thinking was just like, you know, new ideas and research was more about like strong execution, building teams and like shipping products. So I really wanted to be part of that too, and reached out to two prominent investors in Silicon Valley, uh, Elad Gill and Nat Friedman. And, uh, they both were willing to invest and people were like, you know, these two guys are, you know, backing you. You should just do it. Even if it fails, like you learn a lot. It's like, uh, getting MBA and getting paid to do it rather than paying Stanford or Harvard.

AI assessment note: “I went back because I, I couldn't start any company after my PhD”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So obviously we're going to talk, um, you know, all about you, but a quick word on, on OpenAI and, and, and your experience, uh, the second time you, you sort of came back after your internship, presumably there was no longer the 35, uh, uh, person nonprofit company. Well, what was, what was it like working at OpenAI, uh, in like 21 and 22?

A It was, it was cool. Like, you know, um, I think at that time, GPT three was already there. GPT 3.5 was being developed and, and, and, uh, GPT four was not even there. The biggest hit at the time was GitHub copilot and Dolly two. Both of them are really cool. And it was not clear Like, whether, like, you know, for a successful product, you needed the largest compute put on it. That was not clear at that time. Because, like, Jasper and, and, uh, you know, Copilot were all, like, tiny models making a lot of money. Similarly, like, Stable Diffusion and Midjourney were all, like, pretty competitive with Dolly too, and actually probably making more money than Dolly too. So, that part was unclear. So I, I wasn't very, like, um, Not just me, like most people were not sure, like, what is in it for GPT 3.5 or four. But, um, so obviously I was wrong, you know, like chat GPT and GPT four are like the reason why opening is so formidable today. So that nobody predicted and I didn't predict it either.

AI assessment note: “I think at that time, GPT three was already there. GPT 3.5 was being developed”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q time like we are right now, where new things seem to be appearing all the time? Do you have a series of like what discrete teams working on different parts, and then you're trying to put everything together? Um, and so again, I find it fascinating, this sort of intersection between like sort of a fundamental research with like turning this into products and, and how you land it all.

A Yeah, it's pretty straightforward. Honestly, we have a project, we have projects, not teams. Um, I don't like teams as a concept because they're like self, um, nevermind. We have projects, not teams, but one project is pre-training, uh, and fine tuning is kind of a little bit attached to that. Um, one project is agents, so getting agents that we can use. Um, and then a project, we have a project around infrastructure and we have a project around data collection. And they all feed into each other. So, uh, the purpose of data collection is to get good data for the pre-training, and that data can come from our agents, actually, and that data can come from everywhere else. Uh, the purpose of pre-training is to make the models better for our agents, uh, so that the agents actually work, and we have a lot of evaluations, so creating evaluations is kind of a big part of what we do, um, in order to figure out, like, are we actually improving things, or are we not improving things? Uh, the purpose of agents is to get agents that we are able to use every single day, so internally, every single day, that's across, you know, analyzing policy, recruiting, uh, code, writing code, Uh, and all sorts of other things. And, uh, that's kind of how we organize it. And so the agents, the serious use of these agents kind of drives improvements in everything else. Um, and then the purpose of infrastru…

AI assessment note: “we have projects, not teams”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q back now to the other end of the spectrum, like the very beginning of the company. So I was, as I was prepping for this actually, uh, read that the company started in 2017, which for, for the world of NLP was, was a big year, uh, but you were very much at the beginning of that, that, that big NLP wave. So like, how did it all come about?

A Sure. So, uh, it all started, uh, so we started a company, uh, Yoav and myself, um, I had a technical background. Um, this was my, uh, second company. The first one was, uh, analytics company, uh, in the networking space that was, uh, basically acquired by another Israeli company called Cellwise that, uh, was eventually acquired by Qualcomm. And, um, and I, I was, um, I was extremely curious about, uh, AI. I didn't have any background in AI, but I was, I was very curious about this space and actually I had a few ideas. Um, and, um, I kind of randomly, uh, it's an interesting story by itself, but I met, uh, Yoav, who's, uh, my partner and, um, and Yoav, his background, he was a professor at Stanford for almost, uh, 30 years. Um, he ran the AI lab and, um, Actually started. This is his fifth company and, um, and, and, uh, all of his previous, uh, companies were acquired. And when, um, when his last company, uh, acquired in 2015, they decided the whole family to move back to Israel. And, and that's how we met. And, um, and we joined forces to start a 21. And, uh, shortly after, uh, Amnon Shashua joined us as the third co-founder.

AI assessment note: “we joined forces to start a 21.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Was it as a consulting kind of way almost?

A Exactly. Exactly. We worked for companies like Airbus, um, many federal authorities in Europe, um, software providers, uh, in Germany. Building all kinds of systems, all time and material, as you say, um, it was consulting. We went in, we sculpt the problem. We, you know, looked into what do we find in open source communities? Which tools are around? What can we use? How can we assemble and build a solution for them? And then really we started building that solution. So it was great learning experience. And at the same time, also our financing strategy. And while we were doing it, um, end of 2018 already, actually, our hypothesis became somewhat reality, right? You are well aware about what happened. Google released BERT, the first of its kind of transformer. Uh, we were among the early contributors also into hugging face transformers. And this was when we felt, hey, this is, this is somehow, you know, this is a new level of technology maturity, and this is what we were waiting for. So we exposed ourselves a lot to this technology. We only built applications for customers with transformer models. And then end of 20 19, Um, we somehow, you know, reflected and tried to find a common denominator in the way we were building these transformer-based NLP systems. And this is how we came up with Haystack, because Haystack was, um, we understood that, look, there is somehow always a set…

AI assessment note: “Exactly. Exactly. We worked for companies like Airbus... it was consulting.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q to this, that, uh, Hippocratic is, is pre-launch. Um, you just looking at the websites, you, you, you, you, You cover a broad range of, um, uh, you know, things that go from healthcare administration and helping people just, uh, process, um, operations faster to having dieticians involved. So what, what, what is the space you're operating in and the general problem set that you're looking to solve with LMS?

A Yeah, you know, I saw, I lay out three different kind of, um, uses of LLMs to impact healthcare, you know, and that's why I call it a healthcare LLM because healthcare is broader than just diagnoses, right? It's, it's all the things going on, but there's three use cases. First, let's call it productivity in workflow. That means you're in the electronic medical record. You're helping the doctor write their answer or communication from a patient, something called the in basket, which is kind of their equivalent of an inbox effectively, um, or a doctor writing a note to an insurance company to escalate. A lot of people have come up with these ideas. They're interesting. They're helpful. They're actually probably better suited for the people who sell those software systems today to build in. Um, but they don't change healthcare that much. They maybe make you five percent more efficient, 10% more efficient. It's not clear if you make, um, somebody 10% more efficient that they see 10% more patients, right? In fact, if they're right now spending every evening answering their in basket, which most doctors are, and that's time they would rather have spent with their kids, when they get that 10% back from that time, they spend it with their kids, rightfully so. But that means the system didn't get any efficiency. That means we didn't see more patients and that's, you know, that's what ha…

AI assessment note: “I lay out three different kind of, um, uses of LLMs to impact healthcare”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I'd love to double click on some of what you just mentioned. Uh, how does that work to be, um, a, uh, a chat first search engine in particular in connection with the hallucination problem? When do you know when the LLM hallucinates and therefore you should be grabbing information from a live app or a live website? Um, how, how does that work practically?

A Yeah, it's a great question. It's actually, uh, pretty non-trivial, um, to, to know when is a person looking for something purely factual, and you want to have as many citations as possible, uh, and when are they just trying to jam on something novel? You know, you could have some people who want to talk about alien invasion of Berlin, uh, in 2030, because they're trying to write a short story about it, and other people write about alien invasion of Berlin in 2020, Because they think, you know, the politicians are all lizard people, and they're like, you know, think it's some kind of conspiracy thing going on. And so I think there is, ah, you know, a very subtle difference. In one case, you really want to be factual, and the other one, it's fine to jam. So what we often do is first build an intent classifier that understands sort of what is the intent that the user has, and those can get pretty fine-brained. And then as you go into it, you can essentially prime The large language model with a retrieval backend. And this is, I think, where the world is going. Every major LLM company, um, is asking to work with us, and we're actually going to start working with them now and supporting them, uh, in what's called RAG, Retrieval Augmented Generation. And the idea here is that you don't throw away everything from search and say, all right, the LLM now has to memorize everything and k…

AI assessment note: “build an intent classifier that understands sort of what is the intent that the user has”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And the step that's upstream from you guys, and I don't think that you do that because you focus on the storage part, but the conversion, how is that typically done? Um, are there, like, names that people should know in terms of models and companies that do the conversion?

A Yeah, exactly. So there are language models. I'm sure many of you, all of you, hopefully, have used ChatGPT and you're pretty familiar with what that is. Um, there are also something called embedding models. And the output of embedding model is not a paragraph of text that goes into a chatbot. The output of an embedding model is this list of numbers. So an embedding model is still a machine learning model. It's still trained. Uh, OpenAI has one called Ada-II that just got today, 75% cheaper. Um, there are also closed source embedding models from Cohere, Um, Google Palm has one, um, and then there is a plethora of open source embedding models as well that many of the people, many people in the community find to be very good actually, so.

AI assessment note: “OpenAI has one called Ada-II... closed source embedding models from Cohere, Google Palm”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And, uh, fast forward through that early, um, phase. When did data become math?

A Yeah, so there's a chapter on data's mathematical baptism where data takes on the, the sacredness of the academy, and in particular, this scientific way of knowing things by applying mathematics to it. That chapter opens up with a hot IPO, if I remember correctly. The, the, that chapter opens up with the hottest IPO in the late 19th century, which was Guinness. So Guinness, the beer company, IPO'd in late 1800, and like literally people were breaking the doors down to try to get on that, get in on that IPO. Guinness had, like, all of the money, and so they could afford the hottest tech of the day, and they hired the hottest nerds of the day who were the statisticians, except they called them brewers. Brewers was, like, the great title, like, chief data scientist of the late-

AI assessment note: “hottest IPO in the late 19th century, which was Guinness”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Great. All right. So to keep it, um, going, hopefully that gives, uh, a flavor for this, uh, really interesting, uh, book that I've, uh, enjoyed reading, uh, how data happened. Uh, I'd love to use the last few minutes to, um, zoom out and talk about the New York Times. So again, you are the chief data scientist. What, what does data science mean at the New York Times?

A So I started as chief data scientist at the New York Times in 10 years ago. Actually, this is, this summer is my 10 year anniversary there. Um, so data science still means sort of a more orthodox definition of, of data science from 10 years ago, which is developing and deploying machine learning. So the data science team is about a 22 person team that develops and deploys machine learning. For newsroom and business problems. Most of the projects are things that, um, are relevant to many different companies, certainly to many subscribe, uh, subscription companies like machine learning that actually controls the paywall that decides when you should be asked to become a paying subscriber. Recommendation engines, which is not just personalization, but also identifying what's trending and then serving it in a variety of different surfaces. Uh, fancy ad products, so we can, you know, create advertising that's useful to marketers, but, um, is also privacy forward. Um, marketing, so when we market on other advertising platforms, that is done not using guessing and pointing and clicking, but using Python and optimization. Uh, that's a variety of things. We, we have a couple of things that are editor facing to help editors understand the relationship between stories and how they're promoted and how people engage with the stories. Uh, but a variety of problems that are all about developin…

AI assessment note: “data science still means sort of a more orthodox definition of, of data science”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q To the extent you can, can you talk about, um, the tech stack for the data science team and data team in general at the New York Times? I don't know, data warehouse tools that one uses?

A Yeah, that's been a wonderful journey. Um, it was, so when I showed up at the New York Times in 2013, if you wanted to get your hands on data, you needed to write your own MapReduce jobs in Hive and hit buckets of unstructured JSON sitting in S three. Um, then we decided, sorry, then it was decided that we should build our own Hadoop on-premises, um, which was the style of the time. Then all of that went away, and, um, through a story that we don't have time to go into, uh, we started kicking the tires on GCP, Google Cloud Platform, and at this point we have fast, reliable SQL access via, um, Google, Google Cloud, um, and via BigQuery, which has made life so much less painful. That said, there is also a lot of work being done in AWS, and plenty of developer work happening on Amazon's Cloud, um, so the data stack is, in my team, the data stack is SQL and scikit, and occasionally Go, so it's scikit-learn is, is a particular module in Python where most of the machine learning you're going to want to do is already done, Um, a lot of containerization. Um, we still rely heavily on big table because a lot of big, big query, because a lot of times we want to score something and putting, put it in a table so that the analysts have fast, reliable SQL access, um, to the, to the output of those models. Um, I think that about it. Occasionally we code and go when things really need to be per…

AI assessment note: “in my team, the data stack is SQL and scikit, and occasionally Go”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And very clear. So is that, is that the beyond Victor? Is that the future of how LLMs get deployed in the world where the quality of answer matters?

A I, I think there is two ways to have, uh, GPT for your own data. The first way is to do what's called model fine tuning. That is here, Matt, go read all of these 10 books and relearn all of these concepts in your neural network. Uh, I think that way is susceptible to hallucination. That way is very expensive because retraining a neural network actually is very costly to retrain it. Uh, that way is very slow. A new fact coming in, a new book coming in, you have to retrain. It's not going to be available in the outputs until weeks later. Uh, the, the grounded, uh, generation approach is real time. A new fact comes in is showing up in the answers right away. There is no hallucination. I mean, you minimize it significantly and the cost is a hundred times cheaper. So I am biased because that's what I'm selling. But yes, I think, I think that is the new way. The new way is going to be the combination of excellent retrieval models that know how to fetch the facts and then domain specific models, whether that be legal or finance or health that know how to take these facts. And convert them into a response, uh, for the question the user is asking.

AI assessment note: “So I am biased because that's what I'm selling. But yes, I think”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So then clients license the data. How do, how do they consume those, those benchmarks?

A Yeah, clients can consume the data in a lot of different ways. So we have a SaaS application that they can do it. So we have tens of thousands of clients who consume it through a SaaS application. They can also license the data directly. And so we have distribution partnerships, um, either through systems integrators or directly to clients. And we're on some of the data exchanges, you know, like the AWS one, or the Snowflake one, for example. Um, and then we also have APIs, and so we have a lot of clients that will hit APIs, and then access data from us via APIs. Now all of that's permissioned by the consumers, but it's a, a really big thing that, that, that goes on.

AI assessment note: “clients can consume the data in a lot of different ways”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And what about, last question on the vendor, uh, front, like, I don't know, data quality, data lineage, um, any of the cool startups or more established companies in the mix?

A Yeah, so this is the one that, that gets me. So, uh, uh, Mitt Walia and I are good buddies now. Uh, we are, we went with Informatica, and the reason why we're on Informatica is very clear. Um, Not, as Spencer was saying, not everything is in the cloud. And unfortunately for startups, they're making decisions about where to build, and so they're building in the cloud. Informatica has built for the cloud. They've also built on-prem. And so they can help us with all the on-prem systems. We run data centers that are massive, right? And our data centers, we need to be able to connect to old sources. I need to connect to DB two. I need to connect to Oracle. I need to connect to other systems. And so Informatica solved our needs there.

AI assessment note: “we went with Informatica, and the reason why we're on Informatica is very clear.”

← previous page 6 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.