Huang: NVIDIA immediately pivoted into vision, robotics, and self-driving after AlexNet
“Almost right away, we started working on computer vision. Almost right away, we started working on robotics self-driving cars”
Wachs: Handwrytten Uses Computer Vision for Note Quality Assurance
“So we use QA or computer vision to QA the notes.”
Manning: Vision understanding stalled; language does 90% of work in VLMs
“I mean, I think it's fair to say that, you know, vision understanding sort of stalled out, right? You got to object recognition, and then progress just wasn't being made, right? If you look at any of these vision language models, it's the language that's doing…”
Vision Transformers benefited heavily from software infrastructure built for language models
“There was also benefit of doing that simply because the rest of the, that the machine learning field, which was working on, on, on language, they were using this, like architecture. So they were building infra for it, making it faster. And, you know, like the,…”
Huang: AI adoption increased overall demand for radiologists rather than eliminating jobs
“The surprising outcome is the number of radiologists actually went up And the demand for radiologists is skyrocketed.”
Kohn: Computer vision serves as the optical nerve for AI
“So computer vision is the optical nerve, the eyeballs of artificial intelligence. If AI is how a computer thinks, then computer vision is how the computer or the technology robots can interact with the physical world.”
Kohn: Meridian operates as a single-entity venture studio without spinning companies out
“We think about ourselves as this computer vision applied company, but really it's a venture studio model housed within a single entity. So we're not spinning out individual companies. We're not raising a fund. We are actively a single entity that is going to r…”
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Socher: Object detection in computer vision is practically solved
“For example, object detection and computer vision. It's actually kind of solved. We can classify most objects on the planet.”
Palafox: HappyRobot pivoted away from computer vision on YC Demo Day
“On demo day, I literally slag Diana, like Diana, we're pivoting. Like demo day, like we have the slides prepared, the presentation prepared, and like we're pivoting, like away completely from this computer vision platform thing that we've built, this auto labe…”
Isenberg: WeWork built internal AI to detect weapons and smoke via cameras
“I was head of product strategy at WeWork and not many people know this, but WeWork had one of the best AI ML teams in, in technology. And one of the most interesting use cases that we had, I don't think we ended up really productizing it, but it was this idea …”
Doshi: Vision AI can access infinite real-world data unlike internet text
“Whereas like with vision, at least you can like make a robot that just like travels down the street and like just keeps taking pictures of everything. You can get like infinite training data with vision, but it might be trickier to like sort of filter and clea…”
Pompliano: Startups will use computer vision to automate municipal fine enforcement
“So I think there's going to be an entire rise of businesses that just use computer vision to do the same thing. Like anything that can be automated will be automated rather than have humans with their lazy eyes walking around, just have computers that constant…”
Kleinerman: Computer vision democratization will follow LLMs into seamless multimodal AI
“What has happened for language models is coming for computer vision, for images. The democratization, democratization, how do you simplify it? Then there's the intersection of those true multimodal languages, which there are many of them out there, but how do …”
Machine learning engineering skills transfer broadly across vision, NLP, and audio domains
“One of the things about machine learning and AI in this industry is that there's a lot of crossover between these different data domains. You know, if you're able to solve a computer vision problem pretty well, a lot of those skills are transferable to natural…”
Parikh: Non-Visual AI Research Lacks Intuitive Feedback Compared to Vision
“I always thought that it was pretty cool that everybody gets to kind of look at the outputs of their algorithms and see what they're doing, whereas if it's kind of non-visual, then yeah, you see these metrics, but you don't really have a sense for what's, what…”
Delangue: NLP, vision, and audio are the top three Hugging Face tasks
“The three main tasks right now are NLP, so text, right?
From like information extraction, text generation, text classification.
the second one is text to image and computer vision, right?
So object detection, text to image, text image generation.
the third o…”
Zaharia: Model quality experiences diminishing returns from parameter scaling
“And also there's usually, there are usually diminishing returns from scale in, in terms of quality of models in general. And you can also kind of see it in other areas, like in computer vision, for example, we don't have, you know, trillion parameter models.”
Tamiste: Computer vision requires narrow niches because general vision AI does not exist
“You can only do it in a specific niche. And in computer vision, we're doing it in the CCTV camera angle because you can at least not yet make an AI algorithm that can, you know, look at everything.”
Singh: Self-driving vision algorithms can automate desktop keyboard and mouse workflows
“If you take that capability and apply it to your screen, it would know exactly what's going on, right, because the resolution is so high, and so it can understand the text really well, OCR accuracy has improved dramatically, and so if you can understand what's…”
Fortune 500s lack engineering talent to deploy computer vision
“One of the interesting things to note about computer vision right now is it's still very early on in the industry's adoption, and so a lot of these Fortune 500 companies or governments that we work with don't have the capacity or engineering talent To take thi…”
Selling enterprise computer vision is hard without established budgets
“I think that the hardest part about the enterprise sales cycle, especially with computer vision is it's a new purchase. It's a new line item and it's either going to marketing, it's going to security.”
Shum: Microsoft Research Beijing invented ResNet, the most popular vision network
“ResNet now is the most popular deep neural net, ah, in computer vision. Ah, we actually invented in our research lab in Beijing by my students, ah, using 152 layers of neural networks.”
UiPath uses computer vision to mimic human screen interactions
“At the core, what we've built is computer vision technology that allows us to see what's on person's screen, in addition to mimicking human interactions on the keyboard and mouse.”
Brie: Skydio bet on computer vision over heavy, expensive LiDAR sensors
“We did a lot of work with LiDAR, but the vehicles were super heavy, super expensive. So when we started Skydio, we made a big bet on vision because we felt like the progress that's happening in computer vision now especially with deep learning, but even in sor…”
Mike Curtis: Airbnb explores image embeddings to dynamically re-rank home listings
“Some of the areas that we're exploring now is like, how can we, you know, find other embeddings in those images that can take the unstructured data of the image, turn it into something that can actually be tagged and labeled and then used in that ranking algor…”
Lee Fan: Pinterest trains AI to identify visual cues that make photos inspirational
“We are trained computer to learn why this image of the same living room, you take a picture of this way and that way, that looks so different. One just looks so inspirational. The other, like maybe just boring and the computer will start to learn those cues an…”
Lee Fan: Computer vision progress in subjective domains is limited by data
“Right now, I will say a lot of domain is limited by the data. If you only have a limited data to teach, let's say fashion, how can we know this is a fashion that are high end and more for the runway instead of a daily? It's a lot of data because it is a subtle…”
Lasenby: Line-Based Vision Processing Is Classically Much Harder Than Point Clouds
“Lines are much more difficult classically in computer vision. Than points. A lot of reconstruction is done with points.”
Lasenby: Geometric algebra eliminates trial-and-error matrix hacks in computer vision
“If people have worked with computer vision, they will know that often things don't work. So instead of a rotation matrix R, they try R transpose. And instead of a translation vector T, they try R transpose T. And they mess around until it works because it's ki…”
Cohen: Reaching computer vision's final 10% accuracy requires massive scale
“With computer vision, it's easy to get 90% of the way there, but the last 10% is nearly impossible unless you're at huge scale.”
Computer Vision Alone Cannot Verify Poorly Lit User ID Photos
“So I know a lot about optics, and a computer vision system can do a lot, but it really has a lot of trouble if the image isn't perfect, and let's face it, if I said you can take a picture wherever, people take crappy pictures at Starbucks in bad lighting with …”
Hwang: Google DeepDream's barbell representation always included human arms
“It turns out that when you ask it to see, like, ask it to reveal what, like, it thinks a barbell looks like, you know, barbells always show up with human arms attached to them.”
Zaremba: ImageNet error dropped to 3%, achieving superhuman vision performance
“Within several years, people got down, I believe, to three percent error, and that's essentially superhuman performance.”
Hyatt: Rapid computer vision advances are unlocking autonomous tech
“Computer vision has, Made amazing leaps and strides in the last three or four years after being a relatively static field for the previous 20 years, and that is enabling a whole bunch of new categories of technology to exist, baseline of everything from some a…”
Dextro analyzes video strictly through computer vision, not metadata
“This is all just done with computer vision. We don't use any of the metadata whatsoever to identify what's actually happening.”