Getting out of the way: my robotics crash course(thisismypersonalblog.com) |
Getting out of the way: my robotics crash course(thisismypersonalblog.com) |
Finding the table surface is pretty useless using a top-down view, even with April tags, because the range error to an april tag is much more than the bearing (pixel) error to the april tag. You basically have trouble observing the thing you're trying to measure.
If you do this again, ask your agent to conduct this analysis and make sure your desired calibration variables are observable with small error. A second camera from a 45deg angle or even on the table would go a lot further, but then of course other things become unobservable.
Nice workaround using a proxy for force sensing to get touch info, however. And neat project overall!
Can you elaborate on the calibration variables comment? This sounds useful but how would I apply these observations?
You're effectively trying to understand how motor inputs change the end hand position, and in particular, you want to know where the table top is so you can position the hand close to it to pick up/ put down.
This means you have some tuning to "learn" before you can apply a control policy / algorithm - and you should be careful how you phrase this so claude/ai can pick up the right vocabulary and bias towards good solutions.
Adding multiple views helps as follows:
0. Measure from multiple views the table top - Keep cameras stead and rigid, and ask claude to use opencv to do multi-view registration so the plane of the table is known precisely. Keep the cameras steady throughout this process - if they wiggle, you can do multi-view registration each measurement...
Paint the "finger tips" bright orange. Not kidding. Use a very flat chess board under the arm for your "workspace". Also not kidding.
1. Move arm to known position, the multiple cameras will measure the april tags movement AND THE FINGER TIP LOCATIONS. The chess board provides a very nice texture. Or a big flat texture of any kind helps here. Since we know the cameras and table positions, we're getting closer to knowing how the arm movements move the hand w.r.t. the table.
2. If your arm has encoders, then manually touch the table at several points, and claude will record the april tags + joint angles + finger locations.
2b. If your arm does not have encoders, then manually touch the table at several points, and record the april tags + finger locations only, but SPECIFY that the arm is now touching the table. --> Claude can use more opencv code to actually locate the touch point on the table. The measurement is noisy, but you will use many movements to figure it out.
Repeat many times, 10-20, multiple touch points. Then, ask claude to do an error analysis and suggest more touch points. Tell it to use system identification techniques / camera/arm calibration techniques. tell it to research these techniques and report errors.
When you're done, you have to do repeatability experiments - the key is this is now automatic. Claude picks joint angles for the arm, the fingers move, the april tags move, the multiple camers calibrate and record positions, and claude repeats. The manual steps are just the bootstrap - you should be automated now.
What you want is claude to output and ORDF of the whole system - bang you can now control the arm using off the shelf software, which claude is happy to set up for you.
When it comes time to pick things up, the multiple views will locate the object precisely and a control network or algorithm can plug in to control it via the ORDF.
On a technical note, your setup is not far from where professional setups are headed [1]. Astra is really doing an end-run around (for now) research Robotics setups.
[1]: https://x.com/ihorbeaver/status/2104646736652447854?s=20
This post has given me some inspiration though!