Design the harnesses that tell a frontier lab where its robot policy actually breaks. Robot learning or evaluation research background; PyTorch and Ray.
About Orbifold AI
Orbifold AI is building the infrastructure layer for Physical AI. As intelligent systems move beyond language into the physical world, they require a fundamentally new understanding of physics, action, and interaction.
We partner with leading robotics and world model research teams to advance the foundations of embodied intelligence, enabling intelligent systems to perceive, understand, and operate in the real world.
The standards we set, and the infrastructure we build to scale them, will define the next frontier of robotics and Physical AI.
Role Overview
Language models converged in part because everyone agreed how to measure progress. Embodied AI has no equivalent. A robot policy today is graded on a few dozen real-world trials and operator intuition, because nothing can run and check a physical policy a million times. Half a dozen well-funded architectural bets are running in parallel with no shared way to settle which is working.
You will build the thing the field is missing. You take a partner’s trained policy, run it against thousands of real scenes drawn from our corpus, and return not a score but a ranked, repeatable list of the scenes, objects and grasps that break it. Then you close the loop: every failure you characterize becomes a data specification, and the data we collect against it becomes the next evaluation.
This is highly applied work. You will be measured on whether a partner’s real-world policy improved, not on a benchmark number in isolation.
What You Will Work On
- Design evaluation methodology end to end: task suites, held-out slices, scoring, and statistical rigor that survives a skeptical research lead’s scrutiny.
- Build fine-grained failure taxonomies: occlusion, transparency, deformables, long-horizon, contact-rich, distribution shift, and the edge-case discovery and long-tail probing that populate them.
- Train and fine-tune reference policies (π0-class, OpenVLA, ACT, diffusion policy) on our curated data, so we can demonstrate rather than assert that a dataset moves a metric.
- Build automated critics and judges that approximate human evaluation at scale and correlate with downstream model behavior, the only way this works at corpus scale.
- Run evaluations on real hardware where it matters, and know precisely when a real-robot trial is worth its cost versus a corpus replay.
- Work directly with partner research teams: reproduce their setup, run their model against our data, and be the person who tells them the truth about what you find.
- Close the loop: translate every evaluation finding into a concrete collection specification, and back again.
What We Are Looking For
- PhD or equivalent research experience in robot learning, imitation learning, or a closely related field, with first-author publications or shipped work.
- You have trained and debugged real policies on real robots, and you know the difference between a policy that is failing and a robot that is miscalibrated.
- Real opinions about how embodied models should be evaluated, and the statistical literacy to defend them.
- Strong PyTorch, comfortable at scale on Ray, and fast inside a codebase that is not yours.
- Self-driven and high agency, with experience in fast-paced applied research or startup environments.
- We index on the quality of the work rather than years served. A recent PhD with strong first-author publications in a directly relevant area is exactly who we want to talk to.
Nice to Have
- You have authored an evaluation suite or benchmark that other people use.
- Multi-embodiment work: humanoids, bimanual, dexterous hands.
- Experience training or evaluating multimodal LLMs as critics, judges, or reward models.
- Published work on generalization, scaling laws, or data quality in robot learning.
- Teleoperation stacks and real-robot infrastructure.
- Frontier lab, autonomous vehicle program, or humanoid company.