About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI/ML Engineer SME based in the United States.
Responsibilities
- Define, document, and maintain scalable, modular AI/ML architectures aligned with enterprise cloud strategies and product requirements.
- Architect and implement end-to-end machine learning pipelines covering data ingestion, feature engineering, model training, deployment, monitoring, and retraining.
- Apply MLOps best practices to enable continuous integration, delivery, deployment, and lifecycle management of ML models.
- Design multi-tenant and multi-region AI/ML workloads that are elastic, highly available, secure, and cost-efficient.
- Use Infrastructure-as-Code tools to provision and manage cloud-based AI/ML infrastructure in accordance with applicable federal security and compliance standards.
- Design and support reusable, secure, and high-performance AI/ML services and APIs for integration with enterprise applications.
- Conduct model validation, optimization, and performance assessments in cloud environments while promoting responsible AI, fairness, and transparency.
- Maintain comprehensive AI/ML architecture documentation and update technical artifacts throughout Agile sprint cycles and as solutions evolve.
- Incorporate AI/ML metrics, resource utilization measures, and performance KPIs into technical dashboards and reporting.
- Guide development teams in evaluating and securely integrating third-party ML tools, frameworks, platforms, and SaaS solutions.
- Collaborate with Agile development teams and technical stakeholders to support enterprise application modernization initiatives.
Requirements
- 15+ years of specialized experience in information systems or a related technical discipline, with substantial experience in AI/ML engineering and architecture.
- Bachelor’s degree or equivalent required; a master’s degree in a related field is preferred. Equivalent professional experience may be considered in lieu of a degree.
- Demonstrated success architecting and deploying machine learning workflows in cloud environments such as AWS SageMaker, Azure Machine Learning, or Google Cloud Vertex AI.
- Hands-on experience with machine learning and deep learning frameworks including TensorFlow, PyTorch, scikit-learn, XGBoost, or Keras.
- Strong Python programming skills and experience using Docker and Kubernetes for machine learning workloads.
- Experience with MLOps technologies such as MLflow, Kubeflow, TFX, or Airflow and integrating them with CI/CD workflows.
- Familiarity with federal data governance, security, and privacy standards, including JISF, NIST 800-53, and FedRAMP.
- Proficiency with Infrastructure-as-Code tools such as Terraform, AWS CDK, or CloudFormation.
- Experience designing and supporting multi-tenant, distributed, and cloud-native systems.
- Strong architectural, analytical, problem-solving, and technical communication skills.
- Ability to work effectively within Agile development environments and provide technical guidance to cross-functional teams.
- Relevant certifications such as AWS Certified Machine Learning – Specialty, Google Professional Machine Learning Engineer, or equivalent are preferred but not required.
- Must be able to successfully complete a background investigation and obtain a position of Public Trust.
Conditions
- $195,240–$264,148 estimated annual salary range, with actual compensation determined by factors including experience, geographic location, and contractual requirements.
- Fully remote work environment within the United States.
- 40-hour standard workweek.
- Less than 10% travel required.
- Comprehensive medical plan options, including plans with Health Savings Accounts.
- Dental and vision coverage.
- 401(k) plan with company match and pre-tax and post-tax contribution options.
- Paid time off, including vacation, sick, personal, holiday, parental, military, bereavement, and jury duty leave.
- Typically 15 days of paid leave per calendar year for new employees, plus 10 paid holidays, subject to applicable policies and prorating.
- Up to 160 hours of paid family leave over a rolling 12-month period for eligible employees.
- Short- and long-term disability benefits, life insurance, accidental death and dismemberment coverage, and other supplemental insurance options.
- Flexible work arrangements and full-flex work weeks where applicable.
- Wellness and employee support programs.
- Opportunities for professional development, internal mobility, and career growth in AI, cloud, data science, and engineering.
- Access to an AI-powered career development tool designed to identify potential career paths and learning opportunities.
- Opportunity to work on complex, mission-focused technology initiatives alongside experienced technical professionals.