About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer, AI/ML Platform based in the United States.
Responsibilities
- Design and implement the ML platform that orchestrates the complete AI lifecycle, including data processing, training, evaluation, deployment, and monitoring.
- Develop reliable and scalable workflows across cloud infrastructure, Kubernetes, and continuous automation environments.
- Build foundational ML infrastructure components such as model registries, feature stores, experiment tracking systems, and model management tooling.
- Create developer-facing APIs, command-line tools, and reusable infrastructure that make machine learning workflows simple, reproducible, and accessible to engineering and research teams.
- Implement CI/CD capabilities for ML workflows, supporting continuous retraining, automated testing, standardized model packaging, and reliable production delivery.
- Apply MLOps best practices covering reproducibility, data and model lineage, rollback, monitoring, governance, and operational reliability.
- Partner with ML researchers, robotics engineers, data platform engineers, and other technical stakeholders to translate requirements into scalable infrastructure solutions.
- Integrate ML orchestration and metadata tracking capabilities with existing data lakes, pipelines, and broader platform infrastructure.
- Mentor junior engineers and contribute to architectural decisions, technical standards, and the long-term roadmap for cloud and ML infrastructure.
- Help establish centralized experiment tracking, performance visualization, standardized deployment practices, and post-deployment model monitoring.
Requirements
- 5+ years of professional software engineering experience, including at least 2+ years building or operating production ML infrastructure, data platforms, or MLOps systems.
- Hands-on experience developing modern ML platform components such as experiment tracking, model registries, training pipelines, deployment systems, or related infrastructure.
- Familiarity with ML orchestration and experiment management technologies such as MLflow, Weights & Biases, Airflow, Kubeflow, or comparable tools.
- Strong experience with cloud-native platforms such as AWS, Google Cloud, or Azure, along with containers and Infrastructure as Code technologies such as Terraform or CDK.
- Experience processing or modeling multimodal datasets, including sensor logs, camera streams, behavioral traces, or similar robotics and machine-learning data.
- Strong software engineering fundamentals and the ability to build reliable, maintainable, production-grade systems.
- Experience collaborating with research scientists, data engineers, robotics teams, or autonomy engineers to deliver infrastructure used by technical teams.
- Strong understanding of CI/CD, automation, observability, reproducibility, and scalable platform architecture.
- Excellent communication and collaboration skills, with the ability to work across disciplines and influence technical decisions.
- Experience with robotics, autonomous vehicles, drones, or embedded machine learning is a plus.
- Contributions to open-source ML infrastructure or MLOps projects are also advantageous.
- Must have current authorization to work in the United States.
Conditions
- Anticipated base salary of $197,000–$307,000 USD, with final compensation determined by factors such as location, experience, knowledge, and skills.
- Competitive total rewards package for full-time employees.
- 401(k) plan with a 6% company match.
- Company stock options/equity.
- 100% company-paid medical, dental, and vision insurance for employees.
- Company-paid short- and long-term disability insurance.
- Benefits eligibility beginning on the first day of employment.
- Employee Assistance Program (EAP) and well-being support.
- Flexible, unlimited PTO for exempt employees, plus 12 company holidays including a winter shutdown.
- Vacation and paid sick leave for non-exempt employees, plus 12 company holidays including a winter shutdown.
- Generous paid parental leave.
- Flexible work arrangements within a remote-friendly, distributed engineering environment.
- Professional development and tuition reimbursement opportunities.
- Relocation assistance for eligible roles.
- Annual discretionary bonus for eligible roles.
- Catered lunches and snacks at designated office locations.
- Opportunity to work on technically challenging AI, ML, cloud, and robotics infrastructure with significant real-world impact.