Robot Special: Hardware and Algorithms Behind Dexterous Hands
Summary
The real milestone is not getting a robot to open a Coke on camera, but having it autonomously perceive a can in any orientation, rotate it in-hand, align both hands, and apply fine force control. 齐浩之判断,目前“没有这样的公司”。A single task might technically be solvable in a few months with enough resources, but the GPT moment will come when the algorithmic framework can generalize to opening doors, turning screws, and other tasks; his cautious estimate is 3–5 years, while 陶一伟 thinks hardware products at this level could appear within the year.
Robot demos should be discounted based on how constrained the setting is, whether they are teleoperated, and their continuous success rate—not judged on a single successful attempt. In a fixed environment, getting a robot to autonomously load plates into a dishwasher at an 80%–90% success rate and film the result is not difficult; household use may require something close to 100%, because “if you have 10 plates and it shatters one,” users will not accept the robot. Pouring a drink merely involves pressing down on a handle; opening a Coke requires two hands to apply opposing forces to the same object, putting the tasks in entirely different leagues.
Dexterous-hand hardware has yet to converge on a single architecture, but the field is gravitating toward direct drive and one-way tendon drive. Direct drive offers more degrees of freedom and makes control and simulation more straightforward, but comes with costs including a weight approaching or exceeding 1 kilogram, high reduction ratios, and potential problems with service life and impact resistance; one-way tendon drive is less sensitive to tendon creep and makes force control easier, but relies on springs to extend the fingers, sacrificing grip strength and lacking active outward force. 齐浩之 expects “several more rounds of iteration” before the industry converges on a form factor, as humanoid robots have begun to do.
Data, not model slogans, remains the core constraint on scaling dexterous manipulation. Based on conversations with Physical Intelligence researchers, 齐浩之 says π0.5 “seems to” have more than 10,000 hours of data, potentially the largest real-world dataset in robotics; the industry’s “pyramid” has teleoperation data at the top—small in volume but most effective—video data at the base—vast but with no settled answer on how best to use it—and simulation, robot-collected data, and specialized equipment in between. Companies seeking the best absolute performance today still prioritize teleoperation; human video may eventually overtake it through scale, whether in a few months, a year, or longer, but “probably not right now.”
Tactile sensing is the feedback that moves robots from “grabbing by sight” to fast, stable force control. The hand blocks the contact point, making it difficult for vision to determine whether a grip is secure; opening a Coke requires one hand to hold the can firmly without crushing it while the other applies force at the right angle without tearing off the tab. Surface tactile sensors, vision-based tactile cameras, and actuator current together correspond to human skin and muscle sensation: “with touch, it has more information.”
Tesla’s Optimus case shows that the dexterous-hand race is not just about degrees of freedom, but also assembly and production capacity. When 陶一伟 joined Tesla in July 2023, the internal third-generation hand had 6 active degrees of freedom and 11 joints in an underactuated tendon-driven design; after encoders and tactile sensing were added, “you could assemble one hand from morning to night and still not finish it.” After a redesign into what Tesla internally called version 3.1, the production line could build more than 100 units a week by the time he left, rather than one or two a day in the early phase.
Today’s expensive dexterous hands look more like R&D platforms and customer filters, while the commercial path remains split. Sharpa costs about $50,000 per hand, or $100,000 for a pair; Shadow Hand costs about $150,000. 齐浩之 believes vendors may not currently be making money from hardware, but are using it to identify major companies and government-funded universities and validate architectures. 泓君 goes further, describing the strategy as attracting developers and building an ecosystem. TetherIA is taking a different route: accepting lower degrees of freedom and less “advanced” tactile sensing in favor of stable, reliable, low-cost products that can be deployed quickly, then using application feedback to build an ecosystem.
Deep dive
1. Demos Can Be “Made to Work,” but That Does Not Prove the Robot Knows How
齐浩之 first draws the capability boundary: with a human teleoperator and no need for fine finger coordination, many demonstrations are not difficult. In Optimus’s pouring demo, the fingers barely interact; the hand simply reaches the dispensing handle and presses it down. Using scissors or a screwdriver is a different matter: multiple fingers must coordinate precisely, and the control challenge rises sharply.
Once a task expands from one household to thousands of households with different tools, layouts, and environments, the difficulty “rises exponentially.” The gap today is therefore not only fine motor control, but also the ability to transfer the same skill to unfamiliar objects and settings.
陶一伟 adds another hardware constraint: current hands still cannot operate continuously and reliably for long periods in real environments without breaking after repeated contact with natural objects. More degrees of freedom and tactile sensing enable higher-precision tasks, but also raise system complexity and make reliability harder to guarantee.
2. A Randomly Placed Coke Can Is a Comprehensive Test of Two-Hand Dexterity
陶一伟 breaks opening a can into a complete chain: the robot must first perceive the orientation of the can and tab, pick it up from any pose, and adjust its angle in one hand; the other hand must then precisely catch the tab in midair and pull it in the correct direction with the right amount of force. The hand holding the can must counter the pulling force without squeezing hard enough to crush the can.
齐浩之 emphasizes that “this example can be made very hard, or it can be made easy.” Restricting the can’s orientation or adding teleoperation can sharply reduce the difficulty of filming a successful demo; technically, a company could probably complete this single task in a few months with enough resources.
When 泓君 asks whether a can placed “any which way” can be opened autonomously, 齐浩之 gives a clear answer: “There is no such company right now.” This is not technically impossible in perpetuity; companies simply prefer improving algorithms so that future tasks can be adapted more quickly, rather than spending months fine-tuning for a single video.
The dishwasher task mainly involves grasping plates, opening the door, and placing them on the rack—simple grasping and handle-pulling operations. 陶一伟’s comparison is decisive: it lacks the closed-loop control required when two hands act on the same object, oppose each other, and must simultaneously limit force, so it is “not in the same league.”
3. The Household Bar Is Not 80% Success, but Almost No Errors
On dishwasher videos from companies such as Figure AI, 齐浩之 does not speculate about whether the footage was secretly teleoperated. His more useful judgment is that, in a fixed setting, getting current algorithms to achieve an 80%–90% autonomous success rate and filming it may not be particularly hard.
泓君 shifts the question from whether a video is real to whether the product is usable: one successful attempt does not demonstrate capability. 齐浩之 points out that when processing 10 plates in a row, a 90% single-attempt success rate could mean one plate gets smashed—“when people see that, they won’t want to use the robot.”
The key metrics for consumer household robots are therefore cross-environment generalization and a near-100% continuous success rate, not merely the number of actions on the demo list. When investors hear that a robot can pour drinks or clear dishes, they should ask how much the setting can vary, how many consecutive attempts were completed, and what happens when the robot fails.
4. A Dexterous Hand Is Not a Standalone Component; It Links the Hard Problems Across the Robot
齐浩之’s systems view is that one dexterous hand needs at least two arms to form the smallest workable robot system; generalization in real environments also requires a mobile base. Once the robot faces complex terrain and vertical motion, the question becomes whether a fully humanoid form is necessary. Making a dexterous hand genuinely valuable cannot be solved by a single hardware module.
Software complexity also rises with the number of contact points. A gripper may need to plan only two contact points, while a multi-finger hand can have 10 contact points at once; every joint may interact with an object or the environment, with some contact forces helping each other and others opposing one another. The computational and control burden rises substantially.
Research can pair a dexterous hand with a robotic arm and focus on fine manipulation and generalization in tabletop tasks. Or it can mount the hand on a humanoid robot and study manipulation while moving, such as opening a door or completing a task while walking.
Hardware supply in China and the United States has advanced substantially over the past year or two, but companies still differ in shape, actuation, and degrees of freedom, forcing algorithm teams to adapt to each hand individually. 齐浩之 expects several more rounds of iteration before the industry arrives at a “gradually converging architecture,” as it has begun to do with robots from Unitree.
5. Linkage Hands Are Mature and Intuitive, but Low-DOF Designs Struggle to Beat Grippers
Traditional low-DOF linkage hands commonly have 6 degrees of freedom: one for each of four fingers and two for the thumb—bending and lateral movement. Each fingertip closes along a fixed one-dimensional path, while the thumb aligns only with the index or middle finger. 陶一伟 believes that despite their human-hand appearance, their actual use differs little from a gripper.
High-DOF linkage hands do exist. 陶一伟 cites the Korean ILDA paper, in which each finger base has 3 active linear actuators connected through complex linkages to produce 3 degrees of freedom. The trade-offs are greater bulk, rigidly connected components, less compliant grasping, and a higher risk of damage in collisions.
Linkage structures also add burden in simulation. 齐浩之’s correction is important: researchers favor direct drive not because it is inherently more human-like, but because complex linkages are difficult for physics engines to simulate accurately.
6. Direct Drive Simplifies Control and Simulation, but Puts Weight, Cost, and Durability Front and Center
Improvements in the size and power density of miniature motors have made it possible to place an independent actuator at each joint. Direct-drive hands can therefore achieve high degrees of freedom, with a one-to-one mapping between joints and motors that makes control relationships clear. For reinforcement learning, multiple joints are simply multiple direct-drive motors connected in series, so simulation difficulty does not rise as it does with complex linkages.
The costs are concrete: miniature systems often require high reduction ratios, reducing transmission transparency, while the service life and impact resistance of precision gears may become problems. High-strength metal structures also push the typical hand toward 1 kilogram or more, a significant load for a robot’s end effector.
Sharpa impressed the guests for two reasons: it achieved direct drive and high integration within a human-sized form factor, and its level of completion and industrial design were high. Its dual-arm system has demonstrated taking photographs with a camera and dealing playing cards. The latter requires precise control of contact points and force in a tightly packed deck; otherwise, the robot pulls out multiple cards or scatters the entire stack.
泓君 initially interpreted human-hand similarity as a simulation requirement. 齐浩之 clarified that the main reason for a human-sized form factor is that real-world tools are designed around human hands. An oversized robotic hand cannot pick up small objects; the ideal platform should be both easy to simulate and compatible with the scale of human-designed objects.
7. Tendon Drive Resembles Muscles, but Creep, Routing, and Active Extension Become Engineering Costs
The leading examples of bidirectional tendon drive are the 26-DOF Shadow Hand, priced at about $150,000, and the open-source ORCA Hand developed at ETH Zurich. Each joint uses two tendons to control motion in both directions, but the material creeps and loosens with use, reducing precision. ORCA Hand uses ratchets to make tensioning easier, but still requires periodic adjustment.
Another challenge for high-DOF tendon drive is space utilization. Mechanical designs must leave room for tendon paths that change continuously, making it difficult to pack components as tightly as gears and motors. Evan Tao then cites Shadow Hand, ORCA Hand, Tesla, and Yuansheng Intelligent, noting that their actuators are integrated into the palm at the cost of a larger palm volume.
Tesla uses one-way tendons and springs to extend the fingers. The design is less sensitive to tendon creep and easier to compensate for algorithmically, but it lacks active extension force; the stronger the spring, the more force the hand must first overcome when gripping. 陶一伟 summarizes the state of play: the industry is “mostly still solving the gripping problem,” while opening motions have not yet been used as extensively.
Active outward force is not useless. 泓君 uses the example of searching for something inside a backpack with one’s eyes closed: sometimes the robot must push other objects aside before grasping the target. TetherIA uses a hybrid of one-way tendon drive and direct drive, though its exact structure remains confidential. 陶一伟 observes that the industry is currently converging mainly on direct drive and one-way tendon drive.
8. The Training Route Shapes Hardware Choices: Teleoperation Favors Data, Sim-to-Real Favors Modelability
齐浩之 divides mainstream robot learning into two routes. The first collects data through human teleoperation and trains a neural network, with ALOHA and Physical Intelligence as representative examples. The second trains with reinforcement learning in a physics simulator and transfers the policy to the real world; quadruped and biped robots commonly use this route for walking and dancing.
For dexterous manipulation that requires high degrees of freedom and reinforcement learning, teams first assess whether the hand is easy to simulate. 齐浩之 has used linkage and direct-drive designs but not tendon drive; his experience is that direct-drive joint structures are easier to express in a simulator and therefore better suited to current Sim-to-Real research.
This also explains the difference between modern robotics and ASIMO-style achievements 20 years ago. In the past, making a robot run could require a top-tier team to iterate for months or even years. Today, once an algorithm can produce running, modest adjustments may transfer it to dancing. The real advance is not just the action itself, but the possibility of dramatically shortening the time required to learn it.
9. Expensive Dexterous Hands Are First Screening Customers and Architectures, Not Proving Mass-Production Economics
泓君 initially recalled Sharpa costing about $100,000. 齐浩之 corrected that to roughly $50,000 per hand, or $100,000 for two; Shadow Hand costs about $150,000. The buyers are likely R&D departments at large companies and universities backed by major government funding, not ordinary application customers.
齐浩之 believes these vendors may not currently depend on hardware sales for profit. The price acts as a filter, limiting distribution to customers with strong demand who can provide feedback. He cites OpenAI in 2017–2018 as one of Shadow Hand’s main customers, using it to solve a Rubik’s Cube, and adds that at the time there was “not enough financial capacity to support work like this,” though the original remark did not clearly identify what that referred to.
The immediate goal of the high-end route is to identify what still needs improvement in the current architecture and go through more rounds of iteration. 泓君 further characterizes the strategy as attracting developers and building an ecosystem, but that is the host’s interpretation, not something 齐浩之 explicitly states here.
TetherIA is attacking the market from the opposite direction. Its product can temporarily have lower degrees of freedom, performance, and tactile sophistication, but it must be stable, reliable, inexpensive, and quick for application customers to deploy. 陶一伟 emphasizes that cheap “does not mean it lacks technical content, nor does it mean it lacks commercial value”; feedback from scaled deployments is itself an input for further hardware development.
10. Optimus’s Third-Generation Hand Improved First on the Assembly Line, Not in the Demo
When 陶一伟 joined Tesla in July 2023, the Optimus mechanical hardware team had roughly a dozen people, with the hand project mainly handled by him and one other person. The internal third-generation design used worm-gear-and-tendon drive, with 6 active degrees of freedom and 11 joints in an underactuated hand.
The main upgrade at the time was adding joint encoders and tactile sensing to capture spatial pose and contact information. What looked like a simple electrical-function upgrade had to fit inside the inherited first-generation architecture, forcing assemblers to compress springs, thread components, and manage cable routing simultaneously. “One hand—we could work on it from morning to night and still not finish it.”
Musk was dissatisfied when he saw the third generation: it still looked like a laboratory prototype, and production was only one or two units a day. His concern was not just a flashy exterior, but whether the entire design could be manufactured. 陶一伟 then worked with the industrial design team to rebuild the hand from the inside out, producing what Tesla internally called version 3.1.
Version 3.1 later became the version described on the program as still being used “to this day” for installation and mass production. By the time 陶一伟 left, technicians and the production line could assemble more than 100 units a week. The change shows that encoders, tactile sensing, and degrees of freedom become product capabilities only when integrated with process design, maintainability, and throughput.
11. Biomimicry Offers Structural Clues, but Mass Production Must Turn Tendons into Standard Process
Tesla studies biomimicry through “first principles,” with the team reading papers on hand anatomy and tendon forces. They also took advantage of a colleague’s mother, a hand surgeon, to observe the structure of a real human hand in person. 陶一伟’s reaction is direct: “The human body is incredibly ingenious,” but it also makes one realize that “humans are extremely fragile.”
From the previous generation to the one 陶一伟 says is “about to be released,” one of Tesla’s biggest changes was moving the actuators from the palm to the forearm, borrowing the location of some human finger-flexor muscles. He adds that not all human muscles are in the forearm: the forearm contributes more to gross gripping force, while muscles inside the palm control fine manipulation. This is one point of difference between TetherIA’s design and Tesla’s.
齐浩之 asks about the production-capacity difference between tendon drive and direct drive. 陶一伟 believes direct drive is closer to a mature mechanical system, with precision and efficiency achievable through conventional processes such as screw fastening and welding. The industry is still exploring how to connect both ends of a tendon quickly and reliably to the actuator and end effector.
陶一伟 draws a clear boundary, however: tendon fixation, tensioning, and rapid assembly “are ultimately an engineering problem, not a fundamental science problem.” As the supply chain and manufacturing processes advance, he believes these issues can be overcome.
12. Human Video Offers Scale, but Teleoperation Still Delivers the Best Capabilities Today
Meta does not have just one robotics route. A project launched early this year aims to build household humanoid robots across hardware, data, and algorithms, while the robotics group under FAIR focuses more on basic research and algorithms that can deliver the strongest dexterity. 齐浩之 is focused there on learning dexterous manipulation from human video.
Teleoperation requires one operator per robot, so the number and production capacity of robots limit the scale of data collection and make it difficult to reach the volume of language data. Existing cooking and cleaning videos contain abundant hand movements; the long-term goal is for robots to learn skills by “watching” those actions.
齐浩之 does not overstate the current state of play. If the goal is the best possible performance today, directly collecting robot teleoperation data remains superior; learning from video is still in the research phase. It may eventually displace teleoperation through sample scale, but the industry has not yet accumulated enough data or found the most effective way to use it.
1X NEO brings this tension into the home: a human operator still controls the robot behind the scenes, generating real-task data while raising privacy and ethical concerns. 齐浩之 compares it with Tesla collecting data from users’ driving, except that household robots cannot be driven directly by users and must be operated by company staff.
13. The Data Pyramid Sets the Capability Ceiling; World Models Can Add Signals but Cannot Replace Reality
Dexterous-hand data is harder to collect than ordinary robot-arm data. Simple pick-and-place actions can be repeated until an operator gets tired, but teleoperating a robot to cut paper patterns with scissors or fold origami is so difficult that “even collecting one example is extremely hard.” Algorithms must therefore reduce their dependence on scarce, high-quality demonstrations.
齐浩之 says he is “not 100% certain,” but believes Physical Intelligence very likely has the largest volume of effective teleoperation data. From conversations with its researchers, he learned that π0.5 “seems to” have more than 10,000 hours of data. Other companies may have collected more raw hours, but the real question is “what kind of data is useful.”
The industry uses a “pyramid” to describe the data landscape. Teleoperation sits at the tip: low volume, high impact. Video forms the base: massive in scale but not necessarily the most effective at improving robot performance. Autonomous robot collection, simulation, and specialized equipment such as Sunday and Generalist sit in the middle. 泓君 summarizes the distinction as teleoperation favoring precision and video favoring generalization, a framing 齐浩之 accepts.
On world models such as Genie 3, 齐浩之 rejects both extremes. Saying foundation models are completely useless is “one-sided,” but believing that training a huge video-generation model will solve robotics is “unrealistic.” Video generation has not fully solved physical realism; if data alone were enough to learn the laws of the world, language models should already have eliminated hallucinations, and they have not.
14. The GPT Moment for Dexterous Hands Depends on Transfer Speed, Not the First Coke-Opening Demo
齐浩之 declines to rank robotics teams overall because research and products are still “each strong in their own way, with many approaches flourishing.” On the research front are Amazon’s frontier AI and robotics institute, Nvidia GEAR, and universities. On the product side are Tesla, Physical Intelligence, and 1X, along with Sunday and Generalist, which have expanded the vision for data collection, as well as Chinese companies.
Berkeley’s advantage is the density of talent and research directions. Pieter Abbeel founded Covariant; Sergey Levine contributed to Google Brain, Google Robotics’ PaLM-E, RT-1, and RT-2, and later co-founded Physical Intelligence. The university has also built depth in vision, robotics, and systems through projects such as vLLM and SGLang from Sky Lab.
齐浩之 describes Berkeley’s AI researchers as “all sitting on the same floor, like a small startup.” Finding a collaborator in vision or systems often takes only a few steps. Collaboration is usually initiated bottom-up by students, with advisors rarely blocking it. Professors can generally spend 20% of their time on outside activities and can also start companies through a leave of absence or academic leave.
The main disagreement is about timing, not direction. 齐浩之 believes the real GPT moment will come when a reusable algorithm can learn to open a Coke, open a door, and turn a screw in quick succession; he says “I always feel the prediction will come back to bite me,” but still estimates 3–5 years. 陶一伟 believes hardware products could reach that level within the year. 泓君’s summary is that hardware is the foundation, while software and models do more to determine the ceiling.