[00:07] Welcome back to the Azumuta podcast. This is part two in the series on dark factories. Last time, we looked at the academic side of everything to do with robotics and AI, and now we are moving from the lab to the factory floor to see what actually is being used or can be used on a shop floor. So we're focusing on the power of vision. We're talking about how giving machines eyes changes everything, and how that impacts the future of automation in manufacturing. Our guest today is Jonathan Berte, founder of Robovision, someone with a very long history and expertise in robotics, helping industries move towards a smarter, more autonomous future. So maybe Jonathan, before we delve into dark factories, can you explain a bit to the listeners and viewers your history and knowledge around robotics?
[01:03] Yeah. So I'm the founder of Robovision. Robovision was founded almost twenty years ago. We were quite pioneering on the level of how to steer those robots, and obviously we did that first with handmade algorithms. But while we were doing that, this whole deep learning revolution was there, and we decided to build a platform where algorithms could be created by people on the shop floor, supervisors, quality managers, automation experts, not needing any Python or programming knowledge. And that's what it's all about. It's all about giving the tools to the people on the shop floor to create their own intelligence, and most importantly, maintain it. Because we all know if you install something, things happen, some things change, and after a few years, you need to call the supplier, and that's where the friction starts. So we're in business because of the shop floor managers, the experts on the industrial floor, having their own control about AI.
[02:07] As an introduction, not many people know it, but I actually started with Robovision. I started my career at Robovision after my thesis, so it's really fun to have you on the podcast. And what I want to ask about that is, back in the days, this was 2010 or something. We programmed vision with all kinds of software, but it was quite a manual process. You needed engineers to program your vision algorithms. Are there any specific technology changes, discrete steps that have been made since then, why it's now fundamentally different than before?
[02:46] Absolutely, and it's a very good entry point, so I'm very happy to be in the same podcast. The whole deep learning revolution changed everything, because all of a sudden the concept of a program, especially a vision analysis tool, didn't need to be made by engineers nor software people. It was actually a basic deep learning pipeline that had to be trained on test data. And that's what fundamentally changed 15 years ago: you don't need to be there as a programmer to have this excellent pattern recognition tool available on the shop floor right here, right now. We had some big struggles back then, especially in horticulture with all these different plant types. But in the end, we were disruptors ourselves. We didn't need those programmers anymore to build the algorithms. We built a platform to create the algorithms, which is one abstraction layer further away from the real situation. That is the essence. That is the disruptive thing that happened between 2010 and now.
[04:02] Maybe switching over to dark factories. In the past, robotization or automation was typically something for high volume, low variety products, churning out tens of thousands of the same products. How is vision enablement, but also AI, enabling more high mix, low volume factories where robots or humanoids can adapt to product variations and be more flexible and generic?
[04:34] That's a very good question, and it requires a nuanced, sophisticated answer. I think AI gave the comfort to many of these humanoid manufacturers and advanced robotics companies of the new wave, the vision AI robotics wave, to already start building their tooling, their humanoids, producing large batches without the algorithms being fully ready yet. And of course, that rings a bell, because that is what happened with Tesla. Tesla built a whole network, a whole fleet, but they weren't ready with the self-driving, but it turned out to be a good bet. So what is happening in the industry is very similar. We will have sophisticated robotics on the floor and algorithms that are not yet working perfectly, just as with this whole Tesla history. And we know with the Tesla Taxi in Austin, that bet turned out really well. He will probably become the first trillionaire on this planet because of that bet. So many entrepreneurs, new humanoid factories, are taking that template from a more difficult situation, let's be honest. The outside world with children walking the street is much more challenging than the shop floor, and they just apply it to their own business. They say, "Let's already start building. Let's start selling the story, the narrative, and we'll see later." And the only reason that can happen is because AI has gone exponential. They're allowed to make that bet from a risk-taking perspective. And I also believe it will be the thing. It's all about the data, and the data can only be there when the robots are on the factory floor.
[06:37] So you're saying that the data capture precedes the actual AI capabilities. That's basically your thesis?
[06:46] Yes, it is.
[06:47] And as manufacturers, do you recommend them to capture data, even vision data, before they actually apply the algorithms?
[06:57] Yeah, absolutely. In Europe we have a special situation. There is not so much spare space anymore, so there are a lot of brownfield innovations, meaning legacy places that have been installed 15, 20, maybe more years ago. So you already have a lot of activity that you can start recording. Many innovations, especially in logistics, are not about just building something new from scratch. It's about innovating the old world, and that's where data can already be acquired, in hybrid conditions where humans are now doing the stuff and there is only a simple pick-and-place unit. We already have access to stellar 3D capturing devices that can record that data. Because let's be honest, one of the biggest investments of the last years was in data centers. The whole increase in the American economy is because of data center building and everything related to that. And that is mainly because of the disruptive power of the large language models. What they do is scraping everything, all that user data. But we in the manufacturing world don't have that power. We cannot scrape the internet to build the next generation model. So it's really important to already know that we're in this catch-22, and we have to record a lot of data out there to make the world model of the future.
[08:35] Do you agree with the thesis that large manufacturers will have a strategic advantage over smaller manufacturers because they have a huge dataset, and so they are actually stronger in terms of the hive mind capabilities, because they can learn from this huge amount of data that they have and smaller companies can't?
[09:02] The answer to that is more sophisticated. The problem is that you rely on the hype to have your first orders, meaning people need to buy your robots or advanced robotic systems, but they won't work perfectly yet because you need their data. So somewhere in between, there is this gap of desperation, and it's the good entrepreneur who will have the right kind of expectation management towards those early adopters, so that they don't feed the world with the narrative like, "We did this investment and it didn't work out." So we need to bridge that gap. With Tesla it worked. We had a passionate crowd of early adopters who knew they were onto an innovation and that it would not work the first year. But it kind of worked, because they got update after update, and they were even telling their friends, "Oh, I have a new update. I can do a New Year light show with my Tesla." So that gap was bridged almost in a gimmicky way. We cannot bridge that gap in a similar way in the industrial world, but it's up to the good entrepreneur to make the right scheme for that gray zone, because we will have that gray zone. The suppliers cannot just supply them at no cost. The buyers want something which creates return on investment. So maybe it can be gradual: in the beginning it will be very pinpointed, you'll have some Cognex or Keyence-like capability to just pick rectangular objects and put them in boxes, and then the next update will be more sophisticated, and so on. But it's up to the right entrepreneur company to create a scheme that allows an early return on investment, while at the same time using that data to create a more robust and versatile future.
[11:15] What are the most advanced or disruptive concrete applications that you've seen in an automated factory now, with Robovision or others? Things where you say, "Wow, this would not have been possible three to five years ago."
[11:31] The most sophisticated stuff are unpredictable situations where the robot can still do something meaningful, in a very similar way that a human would do that. What we see is that those situations mostly occur outside of the typical industrial world, more in ag tech and horticulture, because there you have very willing entrepreneurs, because they cannot find the labor force, with all the geopolitical conditions now in North America. So they're eager to push that forward and take the risk of writing that purchase order. In the industrial world, and you know that because you've been to the Detroit Fair, they're really conservative. They're already conservative about the vision algorithm. So it's not that these jumpy Boston Dynamics-like videos are anywhere to be seen in the factory. I haven't seen a super-duper humanoid real scalable app yet in an industrial setting. But everybody has seen the movies, the humanoid jumpy, dancing, fighting movies.
[12:53] And that's what we're trying to do, make a distinction between the hype and the nice guided video demos and what's really possible. When do you foresee that you would see these robots, humanoids working together with humans in a factory? Is that a timeline of a couple of months, five years, or 10 years?
[13:11] Couple of months. I would say six to nine months. I think 2026 will be the year of the breakthrough. But everybody thinks that breakthrough will be because of capabilities. I have to disappoint you. We're very boring in that world. It will be because of cost, just cost. If the bill of material of such a humanoid robot can allow that. My prediction, Jan, is that there will be some big scalable investment in humanoid robots by Foxconn or one of these early adopters, because they were always pioneering with their manufacturing chain. But they will do rather simple stuff with that. They will just negotiate a really good price and use the humanoid robot basically for the task of an industrial robot plus the movement of the location. So they will grab the best of both worlds, low cost, and use it just like cobots. Let's be honest, the cobot revolution didn't work out as many predicted. We certainly have humanoid robots, but we don't have large scale cobots, apart from Universal Robots. And that is now an opportunity.
[13:21] I think the cobot revolution didn't break through because it was really hard to program them. Our thesis at Azumuta is that robots are typically programmed in a hard-coded way by engineers, and that's a thing of the past. It's not scalable. In the future, they will be programmed in a human-understandable way, as we do with LLMs, as we prompt them. And everybody can prompt them, because we can already speak the language. So that's the future for robotics as well. You will task them with your procedure. You have to give the robot a work instruction, a documented work instruction with example videos, with example documentation. And the robot will understand the context and go step by step and be able to execute on that work instruction. Do you agree with this, or do you think there might be some additional things needed, some steps still lacking?
[15:45] I partly agree with that, because we may not create the impression that what we did with LLMs will be very similar to what we will do with humanoid robots, for a simple reason that humanoid robots can really disrupt reality in a more impactful way than some hallucinating LLMs. I predicted it will be more akin to autonomous driving. It will only allow the autonomous driving when the conditions are good, just as with Tesla. It's sunny, there is good visibility, and there are no roadworks it hasn't learned about. So what we've seen at Robovision is that it's like a highway, and the way to the highway needs to be managed by a human or a little bit assisted. And then when he's in a safe environment, he can do autonomous stuff and be unpredictable in a safe way. If there is a new situation, he will just move until it's safe to move, and that is promptable, and that will be AI-driven. But in the industrial world, we also like predictability, the predictability of an update. This is why we pay a lot more for an industrial PC than for a gaming PC, because we want it to be predictable. We want to be able to buy it off the shelf in five years, because a machine is depending on it. The same will happen with this humanoid robots revolution. It will need to be compatible with the expectations of the industrial world, and that is not necessarily a high degree of flexibility. That's also not what we want with many of these operators in a car manufacturing line. We don't want them to be creative about inventing new work procedures. Once the work procedure is approved by the manager, it needs to be followed in a very rigid way. So very often that powerful AI component is not entirely needed in a stationary regime. And rightfully so. You don't want to buy a lot of GPUs just because you want the ability to be creative, but you don't need it 99% of your time. So it's also a cost factor. To conclude this diversion, I would say we will use it to the extent that it's healthy to use it in an industrial environment. That means less programming, more prompting. But in the end, what you will see are safe procedures, protocols, and that will not change very much in the next years, because we want this kind of predictability also from a safety point of view. Imagine that a humanoid robot is suddenly going to surprise you and get a coffee and is in an area where it's not supposed to be. It's very gimmicky that he brings you coffee, but you will fire him anyway.
[18:53] How do you think about the human in the loop, going to your example of self-driving cars? Waymo, famously, if it's reaching the edge cases and it's stuck because the algorithm wasn't trained for that case, it halts, and then a human can take over to solve the problem. Is this something you see in the industry being feasible as well, or do you think there are other conditions there?
[19:20] No, I think it's already happening. We know the stories about pick-and-place applications in logistics centers where people in Southeast Asia or Asia are remotely controlling the robot. That's a thing. Some of these things are also not really well advertised, because the stakeholders involved don't want to make a story out of it. They're just happy to pay less operational costs. So I think that's already happening at a large scale, in a very similar way that Facebook in the early days was using the Philippines to make sure unwanted content was not on your timeline. So some of these schemes will just be copy-pasted from the past, like low cost labor doing the edge cases. But in the end, the industrial world is another beast than the real world, and the industrial world has much more responsibilities from a corporate point of view. It doesn't want the thing to hallucinate. So I don't think these racist, hallucinating AI engines that we've seen some years ago will be a thing in the industrial world.
[19:20] Does Robovision also have a human-in-the-loop option? I can imagine that, for example, in agriculture, there are a lot of varieties, a lot of differences, and then suddenly the algorithm gets stuck. Is there some way that you also implemented that humans can course correct or do some extra labeling?
[20:57] Yes. We've done it from the early days, around 2015. It's a principle of a not-completely-dark factory, but somebody responsible for N amount of production lines, where before that there was a whole team of humans needed. That person will intervene in the edge cases, and in our applications it means extra labeling or even creating a new AI algorithm, which in horticulture took a couple of minutes. So that was entirely feasible in the setup that we sold and made value for stakeholders. But the real remote-controlled robot, we haven't done that yet, because we haven't banked so much on the cobots revolution with Robovision. To steer industrial robots remotely is still a bit dangerous, because you need a lot of backup plans with connectivity, otherwise you can be stuck in an awkward situation.
[22:03] So the labeling means there are a lot of cameras doing their analysis, and only the edge cases get highlighted, and so you have one supervisor managing a lot of cameras, and intervening when something goes wrong.
[22:03] Yes. You also have AI drift. Working with a large corporate with such an application where either the product changes over time or the illumination conditions change, so the initial AI model is not performing at the initial standards. And then you have the right dashboard to alert the supervisor that some extra labeling is needed or a new deployment needs to be planned.
[22:48] So dark, fully autonomous factories will never really happen, because you will always have that human in the loop or the human backup needed?
[23:00] There are production situations where this will absolutely happen, because it's simple enough, but it will not happen for all cases. It's a nuanced answer. There are already pretty dark factories. For instance, if you have a plastic extrusion process, I can think of Nico or something, having very little people in the weekend and still very high production output. So it depends on the situation, and very often the company is also very pragmatic about it. Suppose you have four components and three have potential quality issues, and one component can be pretty autonomous in its production. Then in the weekend you could make the calculation: my salary cost is too high, but I still need this one product in great volumes, and then with a very minimum amount of people, you can still do that. Then what constitutes a dark factory? That is maybe also a matter of definition.
[24:10] Do you like the term dark factory? I can imagine, as a vision expert, cameras need a lot of light. It's perhaps a bad name?
[24:17] I don't think it's created by some genius marketing director, because dark, especially in the movie industry, is not positive. It's the dark side, a dark factory.
[24:32] How would you call it then? What would be a better term in your opinion?
[24:37] Something referring to fully autonomous in a more playful way, that it's more attractive to non-industrial people so they have some affinity with it. But I cannot find something on the spot.
[24:55] Any other thoughts about, well, not dark factories, but fully autonomous factories?
[25:01] We're at the beginning of the year still. One question could be what was the key thing in 2025 from this perspective, and what will be 2026? I think 2025 was the year of agentic AI, the breakthrough of the first MVPs. And 2026 will be agentic AI on an industrial level. What does that mean for dark factories, for the industrial world? I had a WhatsApp discussion with a group of five hundred AI like-minded people yesterday. It's about involving the humanoid or the robot or the agentic AI agent as almost a fully capable team member. And what we have with team members is that they have roles and responsibilities. They have access in terms of IT, they cannot go through that door because it's the manager department. And what we have now is the Wild West: "Oh, let's buy this humanoid robot, and he will just walk around in the office and say hi, and that's cool." But then the CFO comes by and says, "Wait a minute. We bought a fifty thousand euro humanoid robot. What for?" And the marketing manager says, "It's a good marketing story for LinkedIn." And the CFO will say, "That's not good enough. If we buy more of them, there needs to be a business case." And from the moment that happens, we will need roles and responsibilities. We need this agentic AI loop for the industrial world to be really safe, to be guided, coached. In 2026, there will be the first startups fully focusing on agentic AI in the industrial world, having frameworks to really work with that in a way that the IT managers and general managers know that the liabilities and the risks are as low as possible.
[27:17] Do you think the compute is there to do something in robotics? LLMs require a lot of compute, but robots based on vision are a whole other level of compute. Do you think we are there already, that it is cost efficient and we can actually deploy something?
[27:33] Very good question. Not entirely. The best thing would be to stream that visual data to something super powerful on-prem. You don't want these very consuming GPUs inside humanoid robots, only for basic tasks or security purposes. But then you also need the networks. You need the next generation of 5G. We only have so much bandwidth in the air to stream all that data. So I think large scale deployments are still to be seen how that will work. I'm really interested how Amazon and their logistic centers would have solved that, because it's a very crowded EM space there. And also, depending on the country, I can assume that the humanoid in Vietnam would consume too much GPU power compared to the cost of a human. But then in Belgium it's another game. It's depending on the local availability of cheap energy, cheap GPUs, cheap political conditions. We know that GPUs have gone through a price hike in the last weeks alone, 15%. So will we have a compute crisis in 2026? I think so. Because some of the things that are happening make sense, but from a pricing point of view, it's very sketchy.
[29:19] You were an early believer in Nvidia, from before people knew Nvidia or knew it only for gaming PCs, you were already using them for vision algorithms. What's the current state today? Do you think Nvidia is still the winner, the default go-to, or do you think eventually there will be competitors? You have, for example, the TPUs of Google. Are there any others?
[29:19] I think 2026 will be the year of the first sizable competitor of Nvidia, but it will not be because of Nvidia's product portfolio or pricing, because people are still happy to buy very expensive GPUs because the business cases make sense. It will be because of geopolitical reasons. China cannot afford to be depending on the latest and greatest compute from the US. And even if they made a 180-degree turn and are now allowing more Nvidia sales into China, the damage has been done. The engineers, the architects are replacing that compute. At some point, compute from China and others will flood the market. It will first start with a very high compatibility with some really good model. For instance, DeepSeek will be compatible with this cheap chipset, and then people will start testing their application with that cheaper pipeline, and then the word will spread, in the same way that the original DeepSeek announcement a year ago led to a price breakdown of a trillion dollars of market cap of all these AI-related companies. That will happen again.
[31:19] Isn't it an inherent property of a GPU that it can be parallelized, and so it makes sense to have less capable, cheaper units, and you just buy more? I think that's exactly what China is doing right now. They are making less capable, but they can produce them cheaper and in way higher quantities.
[31:40] Yeah. I really want to tell you one anecdote. You know Ralph Wiggum? Ralph Wiggum is this figure from The Simpsons, this nerd with few hairs who is very ambitious, very hardworking, but not so smart. Well, after this podcast, go on x.com and search for the Ralph Wiggum Loop. It's the hype in AI. The Ralph Wiggum Loop is something which was discovered by surprise. Somehow, these extremely powerful models like Opus 4.5 of Anthropic internally have their loops and their tools, so they can be fundamentally a little bit more dumb than you'd think, because they're solving that by autocorrecting and giving us the good stuff in a conversation. Now, Ralph Wiggum makes up for many of the shortcomings of the cheap AI. So I made a test last week, and I compared for the same application, a very challenging application, Opus 4.5 with DeepSeek, the latest 3.2. I want you to guess what the price difference was for a challenging application if you just used the Ralph Wiggum Loop. What factor difference was there?
[33:00] Half the cost.
[33:02] Yeah, I was going to say that also.
[33:03] Just go on. I will send you the screenshot afterwards.
[33:09] One third of the cost.
[33:09] More. The difference is bigger.
[33:15] One-tenth?
[33:19] 500 times. Can you imagine that there are people at McKinsey, at EY telling their boss, "Hey, look, works great, but it costs a lot of money." Yeah, whatever. You also cost a lot of money. So you can imagine how many people make mediocre decisions about, "Hey, this is good enough because it's only half price of a consultant," but there is so much engineering to be done to make that one-hundredth of the cost. I will really send you the screenshot. You will be flabbergasted. And that is very important for the industrial world, because we want low cost bill of material. Operational costs are very important. So that is also something for 2026 to be done. It has to bring value, but it may not break the bank.
[34:10] Perhaps one last question. It's a bit of an ethical question. If we assume, I think we are both techno optimists, and so we believe this is really revolutionary, eventually there will be abundance, abundance of goods, abundance of everything. How do we make sure that it gets evenly distributed? Because if it's not evenly distributed, we probably get revolutions of people that are not coming along. Do you believe in things like basic income, or how would you approach it?
[34:44] I think a binary simple solution will not work in a patchwork of nations, because of internal strife and competition. So if one decides to put a tax on the robot, it will basically lead to a non-level playing field. I think that in the next 24 months, governments will discover the power of AI to make new regulation, which is already happening in a very naive way, like staffers using ChatGPT for the text. But I think that once a world model is there, and once there are more generic models about human nature, the GDP will be an AI coefficient, which will just optimize for the best of a nation. And future regulation will have a condition that it needs to increase that new indicator, and there will maybe be a tax, but it will depend on the delta on that new indicator, which will no longer be GDP. Almost a little bit like the Bhutan philosophy of the national happiness factor, but then AI generated. That is my belief.
[35:57] Thank you very much.