About the company
Humanoid is building the world's most capable, commercially-scalable, and safe humanoid robots. Their platform HMND-01 Alpha is running in real industrial pilots.
Responsibilities
- Train language-vision conditioned manipulation policies via reinforcement learning (RL) in simulation and in the real world.
- Construct challenging and diverse suites of manipulation tasks in simulation.
- Partner with teleoperations to collect trajectories in simulation for behavior cloning.
- Partner with testing and operations to establish real-world RL training pipelines.
- Experiment with various ways of bringing policies trained in simulation to the real world.
Requirements
- 3+ years building deep-learning systems (industry or research) with shipped models or published artifacts.
- Hands-on with at least one of: LLMs, VLMs, or image/video generative models — architecture, training, and inference.
- Experience solving real problems using reinforcement learning with deep neural networks.
- Strong Python + PyTorch/JAX; ability to profile, debug numerics, and write maintainable research code.
- Self-driven, pro-active, efficient communication, clear documentation of experiments.
Nice to have
- Experience with simulators for robotics (Isaac Sim, MuJoCo etc.)
- Experience in RL for robotics.
- Experience building infrastructure for large-scale RL (e.g. using ray).
- Publications at ICLR/ICML/NeurIPS or equivalent open-source contributions.
- Familiarity with OpenVLA, Physical Intelligence (π) models, or similar open VLA frameworks.
Conditions
- Competitive equity: stock options.
- 30+ paid days off, including 23 days of annual leave, all UK bank holidays, and additional company closure days.
- Private healthcare, including virtual and in-person care.
- Pension scheme with 8% total contribution (5% employee, 3% employer).
- Free daily breakfast, catered lunch, and snacks in-office.