About the company
FourKites, the leader in AI-driven supply chain transformation for global enterprises and pioneer of real-time visibility, turns supply chain data into automated action. FourKites Intelligent Control Tower® breaks down enterprise silos by creating a real-time digital twin of orders, shipments, inventory and assets. This comprehensive view, combined with AI-powered digital workers, enables companies to prevent disruptions, automate routine tasks and optimize performance across their supply chain. FourKites processes over 3.2 million supply chain events daily — from purchase orders to final delivery — helping 1,600-plus global brands prevent disruptions, make faster decisions and move from reactive tracking to proactive supply chain orchestration.
Responsibilities
- Design, build, and productionize ML models for problems like ETA/ATA prediction, using regression, classification, and time-series forecasting techniques
- Develop NLP/LLM-based extraction pipelines for message-based ETA and status updates (text extraction, entity recognition)
- Own models end-to-end: data pipeline → training → deployment → monitoring → retraining
- Work with noisy, real-world logistics and supply chain data (GPS pings, check calls, carrier data) rather than clean, pre-processed datasets
- Diagnose gaps between offline evaluation performance and live production accuracy, and drive fixes
- Build and maintain automated training/retraining pipelines using orchestration tools such as Airflow
- Set up and maintain model monitoring and observability (e.g., Grafana) to catch drift and degradation proactively
- Replace manual or rule-based processes with ML-driven automation (e.g., automating manual check calls)
- Translate model performance improvements into business impact — operational savings, efficiency gains, and deal-relevant outcomes
- Mentor and guide other data scientists/engineers on technical approach and best practices
- Make build-vs-buy and architecture tradeoff decisions independently
Requirements
- Strong ML fundamentals across regression, classification, and time-series forecasting
- NLP experience — text extraction, entity recognition, or LLM-based extraction
- Production ML experience — you've shipped models serving real traffic, not just built POCs or notebooks
- Strong Python and SQL skills — pandas, scikit-learn, and comfort querying large datasets (Redshift/Snowflake a plus)
- Experience with cloud and data infrastructure — AWS (S3, EC2), and orchestration tools like Airflow for training/retraining pipelines
- Experience setting up or working with model monitoring and observability tooling (Grafana or similar)
- Comfortable working with noisy, real-world data rather than clean, curated datasets
- Experience diagnosing and closing the gap between offline evaluation results and live production performance
- A track record of replacing manual/rule-based processes with ML solutions
- Ability to translate model output into business value and communicate that impact to non-technical stakeholders
- Experience collaborating cross-functionally with product, engineering, and operations teams
- Experience mentoring or guiding other data scientists or engineers
- Ability to make build-vs-buy and architecture tradeoffs independently
- A track record of reducing manual intervention or turnaround time through automation
- Excellent oral and written communication skills
Conditions
- Competitive compensation with stock options, outstanding benefits and a collaborative culture
- 5 global recharge days, in addition to generous PTO and standard holidays
- Parental leave for all parents, annual wellness stipend and volunteer days
- Medical benefits start on first day of employment
- 36 PTO days (Sick, Casual and Earned), 5 recharge days, 2 volunteer days
- Home Office set ups and Technology reimbursement
- Lifestyle & Family benefits
- Mental Wellness support and guidance
- Ongoing learning & development opportunities