Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So speaking of that, so You know, what's, uh, what's, what's happened over the last couple of years? Like any, you know, metric, uh, including vanity metrics that you can share, fundraising history, number of customer, number of documents, whatever it is that you want to share to give a sense to people for the reality of the company as of today.
A I can't forget that you're a VC, so I'm not allowed to share any, any metrics over here. Uh, but there, there's some public ones. Uh, and, and I think things that I'm, I'm really proud of are really lasting or first and foremost, our team. So we built a team of a hundred amazing, uh, incredibly smart, uh, folks all five days a week, sometimes six in office in New York city. And we are now just starting to become a multinational corporation. Uh, and so we were opening SF and we've already opened a London office, uh, which is incredibly exciting with, with, uh, goals to end the year at 300 to 400 employees. Uh, lots of exciting growth. A lot of it here, uh, in Silicon alley. And unfortunately, Silicon Valley there, we have to, we have to move out there a little bit. But I think on, on like AI metrics, or a little bit more about the product and how it's used, one of our favorite things to track is the amount of unstructured data, the amount of pages that are processed by the platform. And a really interesting thing is, hey, last year, Hebbia and probably all of the other major consumer model providers processed around a hundred million pages. Probably around, whatever, hundreds of years, maybe thousands of years of, of, of reading. This year, we're already on track to process around four to five billion pages. Uh, so, you know, somewhere around 50,000 years of reading, uh, for, fo…
AI assessment note: “This year, we're already on track to process around four to five billion pages.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q in detail. What do you find Particularly interesting in AI today, AI research, open source, reasoning models, uh, that, you know, you may or may not, uh, sort of important to have you directly, but what do you find interesting in terms of like what's happening in research? You know, there seems to be a new model every other day, a new thing every other day. What catches your attention?
A It's a good question. There's the felt sense, I think, in communities like this one, and then maybe the larger AI community, that the scaling laws for training have slowed down a little bit. And, you know, we really haven't had a massive paradigm shift that has been released recently. At the same time, there's a lot of really interesting research direction in scaling laws for scaling during inference. And what that means is, hey, maybe we have a fundamental unit of compute, and that is a single inference, a single forward pass on a large language model. How do we actually now, you know, use that or, or run lots of inference like an agent, like OpenAI O three, O four, uh, and all of these kind of more agentic, more reasoning models to actually get better at doing these, these difficult tasks. And actually the, the whole scaling at inference paradigm was, was pioneered at Hebbia. So our early matrix product two years ago, we're one of the first people to say, hey, you get way better accuracy from using more large language model calls, um, at runtime. And so we built lots of infrastructure to scale that up. We, we actually built, uh, an agents team before it was even called agents, uh, to, to actually go out and run these larger jobs. And I think that's probably the most interesting research direction Uh, moving forward is like, hey, let's, let's go and say this current scaling la…
AI assessment note: “there's a lot of really interesting research direction in scaling laws for scaling during inference”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q If that's okay with you, a little bit of a technical deep dive into how that whole works behind the scenes. So, Presumably to start with, you need to connect with a bunch of sources of enterprise data and process the data. So how does that stage of ingestion work?
A Ingestion and indexing is like one of the fundamental pieces of like having any, any knowledge work application. There's many, there's different steps to it. So first it's just like collecting the data sources and hooking into as many providers as possible. And that's not fun or not technical. The, the more interesting thing is, actually, once you have the data, how you index it. We believe that you can do a lot of pre-processing, and I'm not talking about, like, keyword search and building a BM-TWO index or, like, having, like, you know, some sort of semantic search index with embeddings, even your super long embeddings that are coming out. That's not actually that interesting to our users. Again, that's good at doing searches. It's good at finding things in the data. But you want to start to process information before the user even asks a question. You want these agents to be doing work ahead of time. What we do And instead is, we actually ingest documents and pre-process, pre-populate, depending on the doc type, depending on the context of the document, a really rich schema and understanding of each document. That ideally would be, hey, we've pre-indexed and pre-done 90% of the work. At the same time, that's not enough for asking and running user questions on the fly. And for the 10% of the work, i.e., ok, this war just broke out here, or this crisis is happening here, Or th…
AI assessment note: “first it's just like collecting the data sources and hooking into as many providers”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And do you have pieces still of, I guess, almost classical rag architecture, like re-rankers and that type of stuff somewhere? Or you completely skip that?
A Very funny thing, and my first five employees will get a laugh of this and no one, laugh out of this and no one else, but we have, to this day, the most accurate re-ranker that has ever been released, and we do not use it. Um, we spent a lot, like the first year, year and a half, just training, embedding algorithms, training different versions of, of Colbert with like multi-embedding architectures for a single passage, and then training re-rankers. And we came up with a novel re-ranker architecture, which four years later, academia and industry have not beat, and we do not use it. And the reason we don't use it is because a lot of the time search isn't what's important. Right. If we were building a public web API LLF, like Rerankers would be important. Um, but we, we scrapped all of that for a really heavy infrastructure pay a play that uses tons of large language model calls. It is really expensive from a latency and just latent dollar cost perspective, but achieves really accurate answers for deep research, for deep diligence tasks that people can only do in the heavy apply. Great.
AI assessment note: “we do not use it”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And, you know, rewinding back to that chat from a couple of years ago, um, I remember that you had a very powerful kind of mission statement, which was to keep smart people from doing stupid tasks. And if you watch that back, and if you sort of fast forward to today, are you still doing exactly that? In what way has the mission evolved?
A I think when we, when we started Hebbia, I think that one of the fundamental insights was that you had a lot of the smartest people in the world doing the stupidest tasks. Uh, and that, that seemed like there was an arbitrage opportunity there to like, you know, save them some time or some effort or some sweat. And the, the goal of Hebbia was, was never to just stop at, you know, really highly paid knowledge professionals, you know, the investor or the finance guy, the, the, the lawyer. It was actually always to go much more broad. And our vision and mission have solidified really over the last couple of years into building capable AI platform for a billion people. And the idea is that there's lots of consumer AI products, products where you can go and, you know, talk about your sushi restaurant and wine in San Francisco, but there's actually not very horizontal general purpose workplace AI products. And what we mean when we talk about capable AI and AI that can do things is actually something much more horizontal. It's something that's much richer in its interaction. It's actually not as verticalized as you might think, but it still has the depth of what an expert platform would do. So if you think about how that vision and mission have evolved, we started on Wall Street. We started with lawyers. We started with investors and bankers. But now we've built an AI platform for exp…
AI assessment note: “how that vision and mission have evolved, we started on Wall Street”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q you just described, I wrote down, um, something about, uh, what you call the ISD architecture. So that's the RAG part after that. And you were saying that somewhere that RAG sort of doesn't cut it for very hard problems. So what is your approach to RAG? And I assume a lot of people here know, but maybe use the opportunity to define what RAG is in the first place.
A Sure. Uh, RAG is, is a architecture for using AI. That's retrieval augmented generation. Uh, it was first coined in a paper in March of 2020 by a bunch of Facebook researchers. Hebbia were actually the first to turn that into a product. So it's like a very close thing to my heart. So back in 2020, we were the first people to actually productionize it, roll it out. And the idea is that you could hook a search engine up to an LLM. It's basically what Perplexity uses. It's like a search engine to LLM. Except over the web, you always have an answer. Over offline and unstructured and private documents, you don't always have an answer. And what we say when we say RAG doesn't cut it or that heavy, actually, after, you know, making that as our baby, we had to kill the baby and we turned away from RAG to what we call ISD, is we said, well, we need an architecture that's closer to extending the context window versus just searching or calling an external tool. Uh, and ISD leverages that index. So a lot of work that has been done ahead of time, but then also, uh, it leverages and a way of recursively, uh, kind of reading sub documents. So still leveraging tools, sometimes leveraging that infinite effective context window, uh, to kind of bubble up an answer. And it's really a decomposition agent at its core. So you ask a question instead of it just pinging a search tool, it can ping a varie…
AI assessment note: “RAG is, is a architecture for using AI. That's retrieval augmented generation.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q All right. Last question from me, and then I want to open up to folks here. Um, looking forward, Hebbia, in a couple of years, where do you want to be, what do you want to be, uh, building and doing?
A You know, I, I often joke about internally that if we stay in financial services, I will have failed. Um, I think, you know, I really value the go to market of productivity tools that started in a vertical and, but built something and had a product vision all along that was generalizable. And a lot of the time I've actually made hard calls where I could have made a smaller cut, a simpler cut to just serve our ICP, our initial customer profile. I actually chose not to do that, you know, build something more broad. So we're starting to serve law firms, you know, we just started to serve the government in a bigger capacity. Um, I care that when my children are in sixth grade, and they're in computer science class, in their system tray, they have to learn how to use Google Chrome, how to use, you know, Microsoft Excel, and how to use Hebbia. And anything less than that would, yeah, I'd, I'd be remiss, so. Great.
AI assessment note: “if we stay in financial services, I will have failed.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 3 4.45
Q Are you finding that the best, um, deployment, uh, method for, for this is a copilot kind of, uh, analogy, or are you actually, uh, enabling people to just replace those unhappy analysts?
A So I think the thing that I, like, fundamentally very much believe, um, which may be kind of the last demo, you know, made me believe a little bit less, but, you know, it's a, is that humans are really, really important, and making humans better is actually the ultimate goal of what I do. And if, if you look at, you know, I'm gonna harken back to Excel once again, um, bookkeeping before Excel and the structured database was literally in books. You know, a spreadsheet, uh, is called that because the spreadsheet was spanned over the two different pages, and so you'd open up a big book and you'd see a spreadsheet. Um, and when this technology came out, and when it was easily available in these productivity tools, people's jobs changed, and their job actually became, hey, I'm not gonna use a slide rule and, you know, all kinds of, of calculators to, to do my job, and instead I'll use a computer. And I think Hebbia is, is, is very similar. Like, we are building a tool that is aligned with humans, uh, and aligned with creating value. And I think, you know, to date, not a single job has been replaced by Hebbia, you know, knock on wood.
AI assessment note: “to date, not a single job has been replaced by Hebbia”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q So fast forward a few years, what does that all look like? Whether you call it, you know, 2030, let's say, not 20 years out, 2030. So we all, what, AI managers. We spend half hour days interacting with agents. Do we still interact with humans? What does that look like?
A I think that Right now, we're already AI managers. The only difference is the AI takes a single step. So if you are an AI modern organization, you probably are using some sort of chat or rag application. And how good you are at prompting that AI is how good you are at managing that AI. And instead of, you know, kind of letting the AI go and prosecute a task over and over and over and self-correct, it can only take a single step. As agents are rolled out, you'll actually start to see, um, you know, people that are really good at prompting, really good at defining a process, be the best managers, and actually be the best at extending whatever their agenda is in the organization, or making their function the best. And so I think, I think that everyone will be prompting, and prompting is managing, and it will all blur pretty soon.
AI assessment note: “I think that everyone will be prompting, and prompting is managing, and it will all blur”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q Inevitable question about hallucinations, you know, high stakes professional contexts where people have paid a lot of money to provide very reliable results to the customers. Is that something that, you know, in 2025 is as much of a problem as it was?
A I think it's old news. And I think the only reason we still talk about it, it's like everyone talks about hallucinations and, you know, no one knows if it's happening or what's going on. It's like fugazi fugazi. Like, it was a problem back when the models were stupid, but they're obviously not stupid now. And I'd even say that they're way better than any human. So, you know, when we start to think about, like, hallucinations, it's like, well, where does that come up? Like, have we actually seen a lot of that happen? Probably not recently. But the place where it does come up is actually a limitation of rec. It's when you've got the wrong documents, or you're searching for something, and it finds something that's broken. And what you're realizing is that the models are actually pretty good at doing reasoning. They're really good meta learners, like the hallucination, like, whether it's a problem or not, like, no one really cares. What people care about is when these things fail, because they don't have the right context, or they're not reading the documents in the right way. And that's why we've done a lot of research on, you know, how to feed the right stuff at the right time. And if that's really, really expensive, we don't care because models are going to become, intelligence will become too cheap to meter. That's, that's what our industry is predicated upon. So yeah, we'll, w…
AI assessment note: “I think it's old news.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Do you come across actual cases where people will say, well, we're not going to hire as many junior consultants, analysts, bankers, because now we have this tool or similar kind of AI technology?
A I'll say one final story. I think we're running out of time. Um, but I think it's a really interesting story, and it kind of talks about, like, the future and whether or not we'll have juniors. I think we will have juniors. I've seen some people that talk about it. I haven't seen a lot of people that do it. I've seen a lot of third party expenses, like legal fees. I'm very short Accenture and consultancies, but, uh, and the big four. But in terms of juniors, Morgan Stanley, uh, and some of the folks there always claimed that they invented the analyst. And the story behind that is, you know, uh, one day they got a bunch of computers at a computer room and they said, Hey, you know, we don't know how to use these computers, all the bankers. And they hired a bunch of kids that were like nerds from Columbia, and they brought them down into the computer room, and they said, hey, we're going to go and have you use the computers. And the person that made that decision, you know, came back to Morgan Stanley, whatever, 2030, 40 years later, and said, hey, uh, you still haven't figured out how to use computers? And I think the intuition or, or, like, the underlying sentiment there is, When technology is created, like the computer or like AI, and it's a true revolution, you end up actually having lots of people come in and do those jobs, like do jobs related to the technology. So I firmly …
AI assessment note: “I've seen some people that talk about it. I haven't seen a lot of people that do it.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q And as a quick Reality check. I think you mentioned, um, somewhere that AI agents will contribute more to GDP than human workers, uh, within a decade. How backloaded is that? In other words, what's your sense as somebody who's in the proverbial trenches every day of the reality of agents in the enterprise industry, what they can do and what's overhyped?
A We will still be deploying AI. In its current form as a chatbot in 10 years. I think the way to think about that is there's actually still plenty of businesses in New York that don't take credit cards. They only take cash. And organizational change and technological change will take a lot of time and a lot of effort. At the same time, most people use credit cards. Uh, and when things work, and they're not just experiments, they actually happen very fast. You're starting to see that with certain AI companies, uh, that are massively penetrating markets that have gone from zero to some double digit percentage of their, their SAM. And we're fortunate that Hebea is one of them. But at the end of the day, the stragglers, the long tail of adopters will still take time. Um, and I think that to answer your question on the nose, everything is going to be back loaded, right? I think it'll happen in the decade. Uh, like the change to credit cards happened over, you know, five to 10 years. Uh, and now this change should just point and click, and credit cards happen even faster than that. Um, but there will still be stragglers. It'll still take time.
AI assessment note: “to answer your question on the nose, everything is going to be back loaded”
Answered raw tape
D 4 · C 3 · P 3 · Cm 3 3.30
Q So do you think that, to take the extreme of what you said, that research is actually slowing down? So there's Test Time Compute that we're talking about that sort of felt like the major innovation, you know, of the last six months. But after we're done with Test Time Compute, are we sort of running out of tricks?
A I think there, there will be more tricks. I do believe that AI will start to work on itself. I think that I don't mean to be relatively pessimistic, but I wouldn't start a company in AI. I mean that from the perspective of, like, the alpha is gone. Like, you know, when I was starting the company, there was, everyone was like working on crypto. And it was like, you know, the alpha was gone. I was like, okay, we're gonna work on AI. Uh, and so I don't know what it is, is like the next thing, but I, I do think that the models will continue to get better. They'll get way cheaper, they'll get way faster, and the user experience will improve. I think longer term, just as I'm talking about how generalization beats specialization, you saw over the last 20 years of enterprise SaaS, Excel get unbundled into a million specialization things. It only made Excel more powerful. But there will still be, you know, specialization as a second wave to create a lot of value. But I don't see lots of very large changes in the near future. Um, but maybe I've also gotten spoiled from how much stuff has changed.
AI assessment note: “I think there, there will be more tricks. I do believe that AI will start”
Redirected raw tape
D 2 · C 4 · P 3 · Cm 3 3.00
Q Okay, so how does that manifest then for your customers? What, what, uh, what do they, how do they deploy the product to do what? What are some of the use cases you observed?
A A really similar analog to what Hebbia does, um, actually harkens back to Excel in 1985. Where Excel took a technology, the SQL database, uh, and allowed just normal humans to wrangle it in an application and, you know, create value. And they started in a niche, uh, in fixed income, and now governments run on Excel, and in 2000 years, uh, archaeologists will be looking at Excel files, you know. Um, and in, to me, it's actually the most important software to ever have been made. On a similar vein, I think that large language models are this new, very powerful technology like SQL once was, and we're trying to build Hebbia to be just like Excel, honestly, a way to programmatically wrangle in a WYSIWYG way those large language models. And so what that means is it's an application that you can have in your browser or on your computer, and you can have any documents that are easily ingestible, and you could have any large language models that are kind of easily loadable over those documents, and you could run different workflows.
AI assessment note: “we're trying to build Hebbia to be just like Excel, honestly”