About the Company
Raft is a customer-obsessed non-traditional defense tech company dedicated to empowering U.S. military and government agencies with cutting-edge AI/ML and data solutions. We are a leader in autonomous data fusion and Agentic AI, with a purposeful focus on Distributed Data Systems, Platforms at Scale, and Complex Application Development. With headquarters in McLean, VA, our range of clients includes innovative federal and public agencies leveraging design thinking, cutting-edge tech stack, and cloud-native ecosystem. We build digital solutions that impact the lives of millions of Americans.
Responsibilities
- Build and evaluate machine learning models for mission-relevant use cases working directly with government researchers and program stakeholders to understand requirements and translate them into executable technical solutions
- Develop and maintain model training, fine-tuning, and benchmarking workflows that are reproducible, well-documented, and usable by teammates without hand-holding
- Build and improve evaluation pipelines for repeatable, rigorous performance measurement across model architectures, datasets, and operational scenarios
- Integrate models into production-ready [R]AIMS platform infrastructure, working with platform engineers to ensure deployments are containerized, observable, and operationally sustainable
- Support experimentation across model architectures and datasets, maintaining clear records of results and surfacing actionable findings to AI leadership and mission stakeholders
Requirements
- 3 to 6 years of hands-on experience building and shipping production software or AI/ML systems
- Strong Python software engineering skills; writes clean, maintainable, production-quality code rather than notebook-only scripts
- Demonstrated experience developing and evaluating machine learning models, with a clear understanding of what makes an evaluation rigorous versus misleading
- Hands-on familiarity with modern ML frameworks such as PyTorch, TensorFlow, JAX, or Hugging Face
- Experience building and managing model training pipelines and experimentation workflows at a level beyond tutorial projects
- Experience working with distributed systems or cloud-native environments; comfortable in infrastructure that isn’t fully managed for you
- Strong debugging instincts; able to diagnose failure modes in complex pipelines and explain findings clearly to both technical and non-technical audiences
- Ability to work independently and manage workstreams without close supervision while staying well-integrated with a distributed team
- Strong written and verbal communication skills; able to produce clear technical documentation, status updates, and evaluation summaries
- Ability to obtain Security+ certification within the first 90 days of employment
- U.S. citizenship required; ability to obtain and maintain a Top Secret/SCI clearance
Conditions