Research Engineer, Robot Learning

Palo Alto, CA (On-site)

Make human demonstration data train robots. Cross-embodiment transfer, VLA training, and the retargeting that turns a captured hand into a gripper trajectory.

About Orbifold AI

Orbifold AI is building the infrastructure layer for Physical AI. As intelligent systems move beyond language into the physical world, they require a fundamentally new understanding of physics, action, and interaction.

We partner with leading robotics and world model research teams to advance the foundations of embodied intelligence, enabling intelligent systems to perceive, understand, and operate in the real world.

The standards we set, and the infrastructure we build to scale them, will define the next frontier of robotics and Physical AI.

Role Overview

A person driving a robot arm through a task moves nothing like a person doing that task with their own hands. The timing, the reach, the corrections, the recovery from a slip, all of it belongs to the teleoperation rig, and a policy trained on it learns the operator rather than the world. Teleoperation is also the least scalable source there is, gated by the cost of the arm, the operator and the room.

Human demonstration captured directly is the scalable alternative, and it works: once the pose quality clears a threshold, adding human video measurably improves policy performance on settings the robot never directly observed. But only once somebody has solved the transfer. That is your job. You own everything between a captured human episode and a trained policy on a specific robot, retargeting, action-space design, cross-embodiment transfer, and the training recipes that prove the data is worth what we charge for it.

Where the Policy Evaluation seat measures whether a partner’s model works, you make the data itself trainable. The two roles sit next to each other and argue productively.

What You Will Work On

  • Own cross-embodiment transfer: retargeting captured human motion onto robot kinematics, action-space design, and the correctness of the mapping at the moments that matter.
  • Train manipulation policies on our data to establish what it is worth, reference implementations that demonstrate, rather than assert, that a dataset moves a partner’s metric.
  • Run scaling studies on data mixture, quality and volume: how much human demonstration is worth how much teleoperation, and where the curve bends.
  • Turn model behavior into data requirements. Where a policy fails for reasons rooted in the data rather than the architecture, say so precisely and specify the fix.
  • Define how our data should be structured for training, formats, conventions, action representations, so it drops into a partner’s stack without a week of integration.
  • Work directly with partner research teams on their training recipes, and bring back what you learn.

What We Are Looking For

  • PhD or equivalent research experience in robot learning, imitation learning, reinforcement learning, or a closely related field, with first-author publications or shipped work.
  • You have trained manipulation policies that ran on real hardware, and you have debugged the gap between a good validation loss and a robot that still drops the object.
  • Depth in imitation learning and VLA architectures, and real familiarity with the human-to-robot transfer literature.
  • Strong PyTorch, comfortable at scale on Ray, and able to reason about how data quality, mixture and structure impact model performance.
  • Self-driven and high agency, with experience in fast-paced applied research or startup environments.
  • We index on the quality of the work rather than years served. A recent PhD with strong first-author publications in a directly relevant area is exactly who we want to talk to.

Nice to Have

  • Cross-embodiment or multi-embodiment work: humanoids, bimanual, dexterous hands, varied grippers.
  • Egocentric or wearable human data, UMI-style rigs, or sensorized gloves.
  • Reinforcement learning from real-world interaction, or post-training for robot policies.
  • Large-scale dataset construction, Open-X, DROID, Ego-Exo4D-style efforts.
  • Tactile and force sensing in the training loop.
  • Frontier lab, autonomous vehicle program, or humanoid company.

Why This Role

  • You settle the question the field is arguing about. Whether human demonstration can replace teleoperation at scale is an open empirical question, and you are positioned to answer it with better data than anyone else has.
  • Every embodiment, not one. Our partners run different hardware and different architectures, so your work generalizes in a way it could not inside a single robot company.
  • Your results are the commercial argument. The number you produce is what a partner decides on.
  • Founding scope and direct partner exposure from week one.

How to Apply

Please send your resume and any relevant work, papers, projects, repos, to careers@orbifold.ai.

Join the Fold