Member of Technical Staff, ML Engineer (Inference & Performance)

Palo Alto, CA (On-site)

Own how fast, how cheaply and how reliably models run on our infrastructure. Throughput, latency and cost per unit of work, from the serving layer down to the kernel.

About Orbifold AI

Orbifold AI is building the infrastructure layer for Physical AI. As intelligent systems move beyond language into the physical world, they require a fundamentally new understanding of physics, action, and interaction.

We partner with leading robotics and world model research teams to advance the foundations of embodied intelligence, enabling intelligent systems to perceive, understand, and operate in the real world.

The standards we set, and the infrastructure we build to scale them, will define the next frontier of robotics and Physical AI.

Role Overview

Everything we deliver is produced by a model running on our infrastructure. Perception models, evaluation models, verification models, and increasingly our partners’ own models. The inputs are video and sensor data rather than short text, the volume grows every month, and the economics of the entire platform run through how efficiently that work executes.

You own that execution layer. Serving architecture on Ray and PyTorch, batching and scheduling, memory fit, quantization, compiler and kernel level work where it pays, and the profiling discipline that tells you which of those is actually the bottleneck today rather than which one is the most interesting.

This is a performance role with a direct product consequence. Every improvement you land buys more thorough evaluation, more verification passes, and more partner workloads on the same hardware. It is also a founding seat: the inference function at Orbifold is what you decide it is.

What You Will Work On

  • Own model serving end to end. Design and run the serving layer for our production models on Ray Serve and PyTorch, across a heterogeneous GPU fleet with very uneven workload shapes.
  • Drive throughput and utilization. Batching strategies, scheduling, concurrency, queueing, and the unglamorous work of finding out why a GPU is sitting at forty percent.
  • Make models fit. Quantization, precision selection, memory layout, activation and cache management, together with the measurement that proves output quality did not move when you did.
  • Go down a level when it pays. Compiler paths, operator fusion, custom CUDA or Triton kernels where the off-the-shelf version is leaving real performance on the table, and the judgment to know when it is not.
  • Serve the awkward workloads. High-volume video and multimodal inference, evaluation harnesses, and reinforcement learning environments that need many fast rollouts rather than one large request.
  • Build the benchmarking harness that makes performance claims reproducible instead of anecdotal, and keeps regressions from shipping quietly.
  • Run it in production. Autoscaling, fault tolerance, graceful degradation, observability, and cost per workload that anyone in the company can look up.
  • Bring partner models onto our infrastructure: packaging, throughput and memory fit, validation that what we return matches what they get locally, and an honest account of where it does not.

What We Are Looking For

  • 3+ years building production machine learning systems, with deep Python and PyTorch.
  • You have owned inference performance for a real workload, and you can describe the profile before and after in numbers rather than adjectives.
  • A working mental model of GPU execution: memory hierarchy and bandwidth, kernel launch overhead, occupancy, and where the time actually goes.
  • Experience operating a distributed serving or compute framework such as Ray or Kubernetes under real production load.
  • Comfortable being measured on throughput, latency and cost, including the weeks when the number does not move.
  • Judgment about when to optimize and when to leave something alone. Most of the value is in choosing correctly.

Nice to Have

  • CUDA, Triton, or compiler level work (torch.compile, TensorRT, XLA, or similar).
  • Video or multimodal inference at scale: hardware decode, preprocessing, batching across variable-length inputs.
  • Quantization, distillation or speculative methods taken all the way into production.
  • Ray Serve, vLLM, TensorRT-LLM, or comparable serving stacks.
  • Serving diffusion, video generation, or vision language action models.
  • Open-source contributions to a serving or performance stack that other engineers rely on.

Why This Role

  • Performance is the business. Compute is the dominant cost of what we do, so the work you own is visible at the level of what the company can afford to attempt.
  • The playbook is not written. Serving video and embodied models behaves almost nothing like serving text, and very few people have solved it at scale yet.
  • Full vertical ownership, from the serving API down to the kernel, without an abstraction layer separating you from the hardware.
  • Frontier workloads and real hardware, at a size where a single good decision is measurable within a week.

How to Apply

Please send your resume and any relevant work, papers, projects, repos, to careers@orbifold.ai.

Join the Fold