[00:07] Hello, everybody. Welcome to the Azumuta podcast. This is the first episode in a series of three where we focus on one of the topics that has been widely discussed in modern manufacturing and also often misunderstood: dark factories, or non-human factories. We will have discussions with business leaders, experts, and labor unions to look at this concept from all different angles. Today, we want to start with the science of it—looking at the technological future—and that's why we've invited Professor Francis Wyffels and Andreas Verleysen, who I personally know as the winners of the cloth folding contest for robots some time ago.
[00:49] But I guess you're doing more than just that. Today in our lab, we focus on all the technologies necessary to create robot helpers that can do everyday tasks in your home or in your company. That goes from mechanics and electrical engineering up to artificial intelligence. Research in AI and robotics—what does that actually mean? What are the things that keep you busy during the day and at night? In one sentence: what keeps us busy is research—every research technology towards a robot helper in your home.
[01:25] So it means a focus on electrical engineering, mechanical engineering, and AI. We want to bring those aspects together so that robots can do more tasks—giving tasks that humans are currently doing to robots so that humans can focus on things we like to do, like creative things. To translate that to a factory environment, there's been a lot of talk for decades around lights-out or dark factories.
[02:00] Do you see a difference in how we used to talk about or define dark factories, say, 10 years ago versus now? From our perspective, our focus is on lower volumes but high mix of goods, because this generalization aspect is really important. The machine or the robot has to become smarter. While in the past, in the '80s and '90s, the focus was more on high volumes and lesser goods variability.
[02:40] So this becomes more important, and hence our research becomes more important. We used to look at how we can change the environment and the task so that it suits robot technology in order to solve the task, while now we are looking more towards how we can use AI to keep the environment and the task the same, but have the robot adapt to the task. Does that mean the definition of a dark factory is different today than it was, say, five years ago?
[03:13] Good question. Honestly, I don't know. From our perspective, dark means you can be energy-saving by cutting the lights and heat. But from our research perspective, it doesn't have to be that case. Dark doesn't need to be dark. It can also mean adopting existing lines while people are still around, working together with these robots.
[03:49] That's an important part. Some people would include a definition like no humans at all close to the robots—hence changing the environment so that no humans will be around. For us, we generally look at humans being around, which explains the current popularity of collaborative robots that share their workspace with humans. And we never look at automation per se.
[04:18] It's a broader shift—from automation to general-purpose robotics in a human-driven context. Where do you think people are at this moment overestimating or underestimating what robots can or can't do? You always hear the story about folding a towel that takes ages for a robot to do—that's what was said a year ago. Are there other, more general things where people expect too much or too little nowadays?
[04:56] Let's first take the towel example. Folding, say, 10 years ago: Pieter Abbeel was one of the first to do cloth folding, and it took approximately 20 minutes. Then three years ago, we won a cloth folding competition with our lab, folding towels and T-shirts in under two minutes. Today I think we can even push it towards 30, 40 seconds. But there is an important condition there.
[05:31] While our case three years ago was focusing on variable T-shirts and towels, you can push it faster, but then the diversity of towels and clothing types becomes less. With imitation learning, for example, you can very well train a robot to fold one T-shirt really well. But the question, of course, is: can it then fold any T-shirt that's out there in the world? That's a good example of how people now think, "Okay, the robots can do it faster," but an important aspect still lacks—and that's generalization capabilities.
[06:12] In that sense, I don't think we have to expect robot butlers or robot helpers everywhere in the world in the next year or two, as is promised by many big companies. The general idea that lives right now among humanoid companies is that we can just plug and play a robot where a human is currently doing something—hence the big dream of making a robot look like a human, because we have structured the world around us, for humans, for our form factor, and they assume that if we build this humanoid robot, it can just take over these tasks.
[06:58] But I do think that in general we are highly overestimating what these robots are currently capable of doing, because of the generalization properties that Francis was mentioning. For us as researchers, this puts a lot of pressure on us, because we are known as a textile cloth folding lab—textile is a really difficult task, a holy grail task for robots to solve, because if you can solve cloth folding, you can do a lot of other tasks too.
[07:29] That's the general idea. But we get a lot of phone calls when people see the Tesla Optimus robot in a commercial folding clothes, and they're like, "Hey, but they can already do this quite well. What's the status? How is your progress going?" Then we say: if you look closer, you can see that this robot is being operated from a distance by a human in those commercials. So you basically say robots probably won't be able to do every task, especially if they are more complex, and they will need work instructions to know what to do.
[08:05] Yeah. That's at least our premise at Azumuta. They need work instructions. They also need a way to understand those work instructions. And they need training data to learn how to do the task—or other approaches. It doesn't always have to be imitation learning with large behavioral models. It can also be old-school but good computer vision together with task planning and motion planning, for example.
[08:38] Are there specific things that have changed in the last, say, three years? As so many things changed with LLMs, are there technology or hardware shifts worth noting? Many exciting things, of course. To name one on the mechanical side: for a long time, mechanics hadn't made that much progress, but in the last two years we see a lot of anthropomorphic hands at actually affordable prices.
[09:14] Which enables us as researchers, and in the longer term the automation industry, to handle more variable goods. Then, with the advent of vision language models and large behavioral models, we can give robots something like common sense—because that's still lacking today. A robot can do very good computer vision and recognize objects, but having common sense—this glass is full and that one is empty—is way harder to deal with.
[09:52] That's also something we try to research, because that's what humans definitely have: common sense. If something comes up, the human adapts—"I cannot do it because I'm lacking this component." A robot, if it hasn't seen it in the past, will not act upon it and will just try to do the task. You have public models—I mean VLMs and VLAs from Groot, Pi, and what have you.
[10:26] Do you use those in your research group, or do you develop your own things? How do you decide what's built in-house versus off the shelf, when NVIDIA's budget is almost unlimited? First of all, our lab tries to identify niches where we can be strong, and one point where we are strong is interdisciplinary work—combining sensing, electronics, customized fingertips, hands, and skins, and leveraging that with state-of-the-art artificial intelligence.
[11:10] On that side, we often start from LBMs or large VLAs, but then we fine-tune them with local data, often in collaboration with industry or other stakeholders. It might be difficult for a research lab to find the resources to train these very big models. But in truth, for large language models we knew we needed language as input to get language as output.
[11:46] For robots, to be honest, it's still an open question: which inputs do we need? Work instructions will be one of them, but what other inputs do we need to get robots doing actions? It would also be very costly for us to just try around, train very big models for months, and collect the data for it. So we're more focused on researching what type of data we need to collect. That's why we were focusing on making these fingertips feel multiple types of things, just as a human does.
[12:20] And then we can adapt these LBMs so they can also understand tactility, or maybe temperature or other signals from the hands and skin. What are LBMs? Large behavior models—the ChatGPT models for robots. They have their own LLMs: vision, actions, and hopefully tactility and other information too.
[12:53] What is your view on using training data to make those models? I hear the assumption that a humanoid modality is useful because we have so much training data on YouTube, and we can put on a mock-up suit or whatever to generate example data—so a humanoid robot modality makes most sense. Do you agree, or is that not necessarily the case?
[13:32] I would say I tend to disagree, because collecting all that data is still very time-consuming. For large language models, there is a huge amount of data available on the web—digital libraries, internal company data. For robots, what we still see is someone has to teach the robot by executing the task remotely, together with the robot.
[14:06] That's the dominating trend today, and it's time-consuming because it has to happen in real time. The question for us as a lab is: how can we reduce the training time? How can we do that more efficiently by adding other sources of inputs, as Andreas said, apart from the video data? Can we add tactile information, for example? And can we use other representations than pixels to accelerate the process?
[14:39] Concerning data collection: we know large language models have very good performance right now, but they still make mistakes—and we can ask whether we can afford those mistakes on robots. Putting that aside, if we compare the dataset size of something like ChatGPT to the largest robot dataset we have right now, and take into account the data collection farms currently running around the world—all cubicles of people next to each other with a VR headset and remotes operating a robot—
[15:20] if we take that data collection rate into account and want to reach a dataset size like ChatGPT uses, it's more than 100 years of collecting data like this. If we take the current tech and just keep doing what we're doing now, we won't have general-purpose robots in the first 100 years. One hundred human years—you can parallelize it.
[15:43] The current data collection rate, if you keep the current baseline of what we're collecting at, is more than 100 years of collection. On the other hand, it's very tempting, because this becomes very accessible. You see many companies fine-tuning or training their own small VLA to do a task.
[16:12] But then again, the question is: how does it generalize? You show the robot it has to pour water in a cup—that's easy, and you can train a robot within 10, 20 hours to do that. Even an 18-year-old can do that. But the question is: can it pour into any glass in the world? We have experiments from medical sciences where they paralyze people's fingertips and see how well they can take a match out of a box and light it—success rates drop to about 25%, because you cannot feel anything with your fingertips.
[16:51] But you still rely on the senses in your arms and muscles to do the task. These are things we are starting to take into consideration in the robot world, which is why interdisciplinary work is very important for robots. The question is how much these fundamental research topics are taken into account in the current humanoid trend, which often just tries to use the same trick as large language models.
[17:17] Do you see that as part of your teleoperation—feeding tactile information back to the teleoperator? Ideally, yes—so a human can feel what the robot is feeling and vice versa. Something that's really difficult for robots to do is opening doors. It's such a tactile thing. We can feel where the hinges are just by the friction we feel in our whole body when opening a door.
[17:47] It would be nice if, when teleoperating this robot, we could communicate what the robot is feeling on a joint level—the resistance on the door. But it's hard to develop those devices. I had a question on human-robot collaboration. For training purposes it's obvious, but once in production, is it a necessity because the world is structured that way and you still need that collaboration?
[18:21] Or is it more by design that we still want humans somewhere in control or in support of robots? I think it depends a little bit on the use case. If you have high volumes and simple, repeating objects, in the end the robot can do it autonomously—that can be the goal. If you have high task variability and high goods variability, you might want to be able to show the robot a new task without having to go back to the company that developed the robot.
[19:01] One big question is still: a robot can do 100 tasks—how can you teach the 101st? For me, it's a technical question. I'm not an ethics expert, so I don't know whether it's a design choice—there are more people suited to answer that part. But on the technical side, the general premise is that we won't need humans anymore if we have a robot that can do the same things humans can—our hands are magical for dexterous manipulation.
[19:35] That's what we're trying to achieve: the same dexterity and force, and the knowledge on how to operate it, as humans have. We're still very far away from that, and we're not going to see it in the next couple of years. So we will need humans in the loop to collaborate—for example, the robot does the heavy lifting of an object but hands it over to a human on the table, or directly to the human to do a more dexterous task with their hands.
[20:09] So human-robot collaboration will remain an important part. Is there a difference in approach among the companies developing humanoids—Unitree, Figure, Neo, Optimus? Do they approach it differently, and do you have a favorite? We have favorites, of course. What they all have in common is usually a focus on humanoids, which I think is a little bit strange, because in most places you don't need legs.
[20:56] It's way cheaper to have a wheeled base. Another focus they always have is imitation learning—leveraging VLAs, LBMs, and trying to push towards that. What differs sometimes is how they collect data. Some use their repos—I think Google has YouTube, so it has access to a lot of human videos. We don't know actually, but we might think they do it like that.
[21:32] Others just use VR controllers and a VR headset to remotely control the robot and collect the data. Some companies—for example, I was in China in August at a conference on robotics for industry—they even sell a humanoid robot together with a VR headset and remote controllers, so you can train the robot in your own context.
[21:59] For example, you have a small supermarket—you can train the robot to do that task. And that's their business. Another example is Sunday Robotics. They recently launched very impressive videos of a wheeled robot helper—or humanoid robot, you can debate that. What they use is humans collecting data with the same gripper as the robot does.
[22:30] They don't use the hands of the human, but they have a gripper glove with sensors, and the vision data is collected together with the robot. That makes sense, because the closer you are to the morphology of the robot, the easier the robot can execute the task. And they also have nice performance in some household contexts.
[23:01] A company that used to be a little different, but has been around for a very long time, is Boston Dynamics. They built the Spot robot, the Atlas robot, and many more. They make great technology. Now they have an AI institute around Boston Dynamics, so they are incorporating imitation learning and reinforcement learning.
[23:26] A couple of years before this trend, they were more like: "We don't believe in cheap robots—robots need to be very expensive, and we'll put crazy sensors and hardware in there." They relied more on traditional control theory to control their robots and let them do tasks. Now they are incorporating modern AI techniques. But to compare: some humanoid companies try to let their humanoid walk just by doing reinforcement learning, which is different from how Atlas runs around.
[24:07] They are probably converting now—we think—but once again, it's proprietary, we don't know. They probably use a combination. The same is true for Toyota Robotics Research. There's a lab we actually really like because they have an interdisciplinary approach, rooted in control theory and mechanical engineering, and expanded that to human and social sciences, electrical engineering, and computer science—to combine that in a robot helper that can do tasks together with humans.
[24:49] That's something I find more impressive than yet another company just copy-pasting tasks. When do you expect us to have a robot helper in our house to do random tasks—the things we don't like to do, so we have time for the things we do like? Well, actually we already have robot helpers—we just have many tiny robot helpers.
[25:27] But one generic robot that can combine all these tasks and put things in the dishwasher? It's always hard to put a date on that, but I would expect—earliest 10 years, at least a decade. At home we will probably need legs to walk around, but at this point in time, look at the videos: you don't want to get close to these humanoid walking robots that weigh 100 kilograms. They keep their distance, except when there's a big table in between. In general, they try to stay two to three meters away, because if these things fall, they first try to recover by pumping a lot of energy into their legs. If a 100-kilogram robot puts that energy into your feet for a couple of milliseconds, you'd need a medic at home.
[25:53] And apart from that, task generalization remains the challenge.
[26:17] Take the example of pouring drinks in a glass. In our lab we are confident we have a very good system that has a 99.3% success rate. But there's still a 0.7% success rate where the glass breaks. And then what? We are confident in that number because we can deal with any kinds of glasses used in houses and elsewhere.
[26:49] Most companies show just one type of glass, or maybe two or three. And what's the success rate? What's the generalization capability? They are very quiet about that. And when something breaks, then what? The robot cannot recover from that. If you would buy a humanoid robot now for your house, you have a lot of cleaning to do.
[27:16] Not to pin us on the number—we think a decade. But as a society, I think we will slowly change the definition of a humanoid. Given the marketing budgets right now, we'll still call it a humanoid, but they will probably give it legs and wheels, and still call it a humanoid. The current form—legs and humanoid hands that they claim can do the same thing as human hands, or at least try to achieve—not in the first decade.
[27:45] And we don't want to be pessimistic. There are a lot of opportunities in what we like to call semi-structured environments, because the house is total chaos. But there are many environments where there is some structure—SMEs, small companies, healthcare, hospitals, elderly care centers. They have very similar rooms and similar layouts.
[28:17] So I truly believe you can mean something in those locations first, and that makes sense. But the household robot is still very hard. So you expect a more dedicated laundry-folding robot sooner than a general cleaning robot in your house?
[28:51] Yeah. As a lab, we expect that. Today you have wheeled carts that drive goods around—even in hotels in China, and there are some in Belgium in restaurants already. I truly expect they will get arms or manipulators and do some basic tasks, like maybe cleaning up the table—but not in your house, in a restaurant or another standardized environment. And that's already challenging enough.
[29:19] How do you see open source versus closed source approaches? I know there are projects like Le Robot on Hugging Face that collect huge datasets because they are open source and cheap. Do you think that's an approach that can win, or will it be the private players with larger budgets?
[29:49] What I like about these open source approaches—the same with some YOLO implementations for vision—is that they make it very accessible. Some very creative bottom-up approaches might pop up, because now you see even school kids working with that. I know kids of 16 who bought a Le Robot, an AZERO 101, used the imitation pipeline, and do creative tasks with that. That's really interesting because it gives creativity to people, and we might get ideas from that.
[30:24] But on the other hand, accessibility also comes with over-expectations: "My 16-year-old kid can train a robot to do this task—next year we'll have much more in companies." That's the counter side. We saw the same in computer vision with deep learning models that became very accessible for basic tasks, but not all vision problems are solved.
[30:56] The difficult thing with closed source is it's hard to find a moat around your business model—to protect your business from people copying what you're doing. Same with humanoids now: we have so many humanoid companies because you can copy-paste the thing, collect the data, and use open source models. We've seen this with language models: open source was always a couple of months behind the closed source models—not as good, but there's a business case if you can run these models free locally, privacy-wise more interesting. It has a place in the landscape.
[31:30] Open source is always catching up, so it's still to be seen whether this will be true for robotics models. If it's also the case, I think it forces companies to make things open source, which as a society is a net benefit. We're academics, so that makes sense.
[31:57] Do you think there is a difference? We are on the SME side, working with factories and manufacturing environments. Our thesis is that large enterprise manufacturers have a lot of people and a lot of data—they will actually have a moat compared to up-and-coming companies, because they have large amounts of training data. Do you think this is true, and will this be the case going forward, or will foundational models become good enough that smaller enterprises can also start using them?
[32:32] One of the problems is that data and automation are a bit similar to robotics. There is plenty of data, but nevertheless it's still not that much data—it's a niche. If you look at foundation models, the proportion of data from the manufacturing industry is a very small portion. So the generalization capabilities to use it in a small company will be limited.
[33:06] That's why I think it's really important for companies to value their data and their data sources and to leverage that. I think that's also where we can play a role as a regional entity—as Belgium or even Europe. We can put efforts into leveraging our automation industry by combining data sources together and having models trained on that.
[33:41] We cannot win on general LLMs, but by combining efforts on niche data, we can maybe at least have a good foundation model for the automation industry. The same is true for language: we cannot make a large language model for English and reasoning capabilities, but we can maybe make one that handles dialects very well for basic conversations.
[34:31] And that's really useful for our local industry and hospitality. For instance, people in West Flanders—if you put a robot in elderly care or a hospital, it doesn't need to speak English. It needs to speak the local dialect. That's what people want to hear, and it doesn't need to talk in very complex ways. It needs to do basic conversations, and that's feasible. The same is true for the automation industry: maybe not catch it all, but a smaller niche—and I think we have high-quality data available, because we have a very good industry here.
[35:30] So we have high-quality data, and maybe we should be leveraging that. If you had a manufacturing factory yourself, would you today already start collecting data, or do you think that's not necessary—the robots are not yet good enough, it can wait? I personally would have started 10 years ago with collecting data. Of course, that's a bit of a biased opinion. But I would feel very late if I would start collecting data now. I think they should be collecting data, and if they are not yet doing it, they should start definitely now, because there's a lot of value there.
[36:04] Thank you very much—very interesting. For the people watching or listening: with what questions can they approach you? How would you be able to help them? Is there anything you would want to say to them?
[36:31] As a lab, we are always interested in collaborations with industry. Any company can come to us, especially on problems in robotics and AI, with a focus on the manipulation and handling of challenging objects—deformables, glass, shiny objects. We're always willing to help. We also actively try to put common sense into the current inflated expectations around robotics—to give a very clear view on what's currently possible, because there's a lot of new possibilities out there, but also what is currently impossible.
[37:11] Great. Thank you very much. Thank you. See you. Thanks. All right.