Reasons robotics is hard(secondthoughts.ai) |
Reasons robotics is hard(secondthoughts.ai) |
I'm impressed with how far legged locomotion has come. But as yet, nobody seems to be using legged robots for any commercial purpose beyond the demo level.
Is Tesla still going to produce vast numbers of humanoid robots by the end of 2026?
There's been a lot of progress on the hardware side. Motor technology from drones has produced much better robot motors. The sweet spot on gear reduction seems to have been found. (Too much reduction, and you can't back drive. Too little, and the motors have to be too big.) The volumes are now large enough to justify making robot-specific components. Robot arms are much better than a decade ago. So are robot legs. Control is better, too. It looks like a humanoid robot will cost about as much as a car.
But they're still not quite good enough to be useful.
We'll know they are real when an Amazon Prime truck drives up and a robot does the last 100 meters of the delivery.
Their robots do movement using sliding scale in 3D spaces (think poles that are left right, up and down, and the "picker" being able to glide and move. They currently have a ceiling slider that goes down and suctions things into a pneumatic tube to then end up in my grocery bag. IMHO it works pretty well - especially considering that delivery is ~$8 for me.
Ultimately we're going to end up with several different types of robots, and not with a human centric vision. The question is if bipedal is a long term dead-end, and merely a short term method to fit into the world we currently have designed.
That being said, I don't know what exact state the industrial automation technology is there and I can only extrapolate (or do websearch, which didn't lead to enough details; I only found things like https://www.youtube.com/watch?v=JnUGgc8R3ng).
When something like https://www.allegrohand.com is mass produced and used industrially, I bet the last mile (meter?) would change a bit, and full automation would be easier and less finicky.
Current generation tactile sensors cost a couple thousand $ PER FINGER, and have a real world MTBF of hours. The cost can be solved with economy of scale. The fragility is harder.
>> It will be difficult to match this scale of breadth and depth of data for physical tasks. There’s no straightforward equivalent of “just Efficient learning, generalization, and adaptability / on-the-job learning seem like requirements.
Isn’t this the idea of NVIDIA’s Isaac? Model based adaptive learning in virtual environments for robotic systems? Or is this oversold?
Are we talking about experimental laboratory ones here? What happens when the Alibaba players start getting into the game? They have plenty of humanoid robots.
I think it'll be more capable appliances at first.
Like a lawn mowing device that also spots weeds and can spray them.
Next iteration has arms to rip weeds out of the garden.
Next has attachments so you can direct it to do pruning.
Next it can figure out the pruning itself and move the outcome into the woodchipper.
And so on and so on.
It's not going to be one day a humanoid robot comes into the house and does everything.
That's available as a tractor-pulled implement for farms. Deere and some others make such things.
Which doesn't mean I disagree with you. I also think that this progression is the most likely. But it implies we're decades away from broad adoption rates.
Think fifty years to hit mass adoption, not five. (Because that much more closely aligns with other structurally disruptive tech like automobiles or computers) Which is definitely not the story being pitched to investors at the moment.
Products that start life as "for the rich, first adopters" and work their way down the economic classes are a thing. Whether it is the thing in this case I don't know.
> Which is definitely not the story being pitched to investors at the moment.
I would think that what is pitched is what investors want to hear....
We have self driving cars because what are the control inputs? Pedal, brake, steering wheel. This already took many many years.
Now for a humanoid robot: An action space that is metaphorically Hilbert. (Physically, yes, obviously)
Also, IMO, LLM's can aid the development of robots, but do little beyond a planning, human control interface. Below that it's the domain of control and the solution will be the correct combination of classical, neural, and real time optimization based control.
All the bad-ass biped robots that actually look natural? It's PID controls wrapped with control barrier functions constraining the QPs that are being solved in real time.
But that's annoying to derive per-application. So we'll need neural methods which can be learned (while being constrained by a priori knowledge of dynamics). My hunch is that the Yann LeCunn type of jepa models will be how tasks can be learned.
That's not entirely true. Locomotion is well addressed by RL in sim. It's true that there is still a PD layer, and the RL policy produces setpoints for it.
Data is a problem. LLMs had the advantage of the whole internet to train on. Robots don’t have that corpus of information. And real time learning seems to be something that everyone in AI is studiously ignoring.
Also there’s imitating humans, via a suitable mapping from the human sensor, control and configuration space to the robot’s. Some groups have gathered video and other data from humans doing tasks, for example with a VR headset.
Maybe in another 2 decades I could see it possibly starting to change, but even then I wouldn't bet the horse on it until I saw it. Cars only have three degrees of freedom and even that we are barely able to get working well enough to put it into limited practice. And yet one single human finger has atleast 3 degrees of freedom, and is covered in what is the equivalent of a million tiny ultra sensitive tactile sensors.
Wrong. Try hitting the brakes of your self driving car on highway at 65mph or during unprotected left turn with oncoming vehicles.
Or have a glitched self-driving car hit its brakes and block the road, for emergency vehicles, and endangering other people.
Self-driving cars can also suffer a glitch without knowing they suffered a glitch, like Waymo cars driving into flooded roads.
I do. I don’t want robot vacuuming or making noise at night or doing something potentially dangerous unmonitored while people are asleep.
Not folding the laundry, though.
I feel this is actually somewhat straightforward. I assume deep water on roadways is not commonly in the training set, because frankly it isn't common in real life, and when it is common people do not drive and do not gather that training data. As a result the proper response has not adequately been beaten into the models. There are probably also challenges of world-sensing, since water can act as a mirror, and maybe other complications. So waymos are bad at handling deep water on roadways. However, deep water on roadways is also not common in the areas where waymos are deployed. As a result, waymo's have a great safety record, and at the same time they make mistakes that are obvious to a human.
A common criticism of AI discourse is that people act as if LLM's "think". I don't want to be a vocabulary purist, but I suspect that's related to the astonishment here -- the Waymo doesn't know what flooding is, it doesn't fear drowning, it doesn't think. So unless it's been repeatedly trained, or a special case has been hard coded by manual effort, it doesn't know that flooded roadways are dangerous.
I have made a lot of assumptions here, and I don't truthfully know what the training data looks like. Feel free to push back if you think my assumptions are wrong. I'd especially be interested if somebody can show that water on roadways _is_ in the training data
https://www.youtube.com/watch?v=FUUzmRH5Yi4
https://www.construction-physics.com/p/robot-dexterity-still...
The videos I've seen of humanoid robot applications are basically that it can do dishes and fold laundry, but I think if household chore robots ever come to market, they would probably not look humanoid at all and probably look like semi dishwashers/washing machines with wheels and a gripper arm.
The form factor has been solved basically everyone is building humanoid robots and hoping a transformer with a big enough dataset is going to do the rest.
Perhaps nanobots will be able to carry the chemical makeup of a cheeseburger and rebuild a bite directly in our mouths, no cooking necessary!
Or order it from Amazon, in which case there was likely a robot in the pipeline.
Robots are very widely deployed, but almost entirely invisibly to the customer yet.
Roomba is the main exception.
> Once they have context, robots will need to reason, plan, and exercise judgement and common sense. LLM-based systems like ChatGPT and Claude are making great strides in these areas
But why would I want to make AI more powerful - and disruptive - than it already is? I don't see this as a benefit but as a disadvantage. Let's also not forget that e. g. Google deliberately ruined its search engine. Now if you search something, by default, you get AI slop results that are often not truthful or only partially truthful. This is a private web. Google wants to control information.
really anything that involves a bade near your body or where body contact is the point.
The Chinese government is also pretty good at finding things for people to do, so I doubt that's a factor. Go to any big park in a city and see how many people are sweeping up leaves, or guarding a building/area that isn't particularly secure.
That's why the human operator needs a gun. :-)