Research Engineer, World Models & Simulation

Palo Alto, CA (On-site)

Train learned world models on real captured experience, and turn them into environments a robot policy can be graded in. Generative video, world models, or simulation background.

About Orbifold AI

Orbifold AI is building the infrastructure layer for Physical AI. As intelligent systems move beyond language into the physical world, they require a fundamentally new understanding of physics, action, and interaction.

We partner with leading robotics and world model research teams to advance the foundations of embodied intelligence, enabling intelligent systems to perceive, understand, and operate in the real world.

The standards we set, and the infrastructure we build to scale them, will define the next frontier of robotics and Physical AI.

Role Overview

To grade a robot policy you need an environment that is both varied and accurate, and nothing available today is both. Testing on a real robot is accurate and covers almost nothing, a few dozen trials, a broken arm, a room that never resets the same way twice. Hand-built simulators run cheaply but can only show you what somebody already modeled by hand, so the variety stops where the asset library stops. Web video has endless variety and no measured physics at all.

You will fill the empty quadrant. Learned world models, trained on a corpus of real captured experience that is reconstructed rather than estimated, good enough that a partner’s policy can be rolled out thousands of times against the same failure taxonomy we apply to real data. And the same models run the other way: amplification, turning one hour of captured experience into far more coverage of the long tail than an hour of filming can buy.

This is a founding seat. The generative side of Orbifold is what you decide it is.

What You Will Work On

  • Train generative world models on our multi-view real-world corpus, action-conditioned video prediction, learned simulators, and make them good enough to evaluate a policy in.
  • Generate targeted synthetic data aimed at specific measured deficits rather than volume for its own sake, and prove whether it helps a downstream policy.
  • Build the rollout and evaluation harness around those models: scenario suites, scoring, reproducible runs, and reporting a skeptical research lead will accept as evidence.
  • Measure the reality gap rigorously. Quantify where the model diverges from the world, publish it internally, and let it drive what we go and capture.
  • Work with classical simulation where it earns its place: Isaac Sim/Lab, MuJoCo, Genesis, ManiSkill, and be clear-eyed about where hand-built physics beats a learned model and where it does not.
  • Reconstruct real sites as environments a partner’s policy can be tested inside before it is allowed to run there.
  • Scale the training. Multi-node PyTorch on Ray, with throughput and memory treated as part of the research problem rather than someone else’s job.

What We Are Looking For

  • PhD or equivalent research experience, with first-author publications or shipped work in one of: world models and action-conditioned video generation, video diffusion or autoregressive video models, model-based RL, neural rendering and reconstruction, robot learning in simulation.
  • Strong ML engineering, PyTorch, distributed training, and at least one project where infrastructure was genuinely part of the research problem.
  • A specific, clear-eyed view of where generated data stops being useful. We would rather hire someone who can say precisely what their model gets wrong.
  • Comfortable being measured on whether a partner’s real-world policy improved, not on a generation benchmark alone.
  • Self-driven and high agency, with experience in fast-paced applied research or startup environments.
  • We index on the quality of the work rather than years served. A recent PhD with strong first-author publications in a directly relevant area is exactly who we want to talk to.

Nice to Have

  • Large-scale video model training; 3D and 4D generative modeling.
  • Hands-on with a physics simulator, and an honest account of where sim-to-real broke for you.
  • You have built an evaluation benchmark that other people use.
  • GPU performance engineering, kernels, parallelism strategy, throughput.
  • Open-source contributions to a world-model, video-generation, or simulation repo.
  • Frontier lab, autonomous vehicle program, or humanoid company.

Why This Role

  • A training set almost nobody has access to. Curated, verified, multi-view real-world capture collected against known failure modes, not scraped video.
  • The empty quadrant is the whole opportunity. Accurate and varied at the same time is what the field is blocked on, and it is what this seat exists to build.
  • Your evals change what the field captures. The gaps you find become collection programs across multiple frontier partners.
  • Founding scope. No playbook to inherit, and the architecture decisions are yours.

How to Apply

Please send your resume and any relevant work, papers, projects, repos, to careers@orbifold.ai.

Join the Fold