About the Role
As a Software Engineer on ML Platform at Cursor, you'll build the infrastructure that turns real product usage into better models — and keeps research moving fast on large GPU fleets. ML Platform is organized into four teams. Depending on your background, you may join any of them:
- Telemetry — Own the collection and serving path that turns real product use into a record research can trust; without slowing the product, and under a small, explicit policy. Client-side or high-volume ingestion experience is a plus.
- ML Data Platform — Build the shared environments and pipeline substrate researchers extend, so new experiments don’t fork their own stack.
- Observability — Make it easy for researchers to start, watch, and debug their own runs.
- ML DevX and Systems — Shorten the path from idea to a trusted run on the research fleet.
We're looking for strong distributed-systems and infrastructure engineers who want to sit next to research and ship platform primitives that move the product.
We're in-person with cozy offices in North Beach, San Francisco, Palo Alto, and Manhattan, New York, complete with well-stocked libraries.
Responsibilities
- Design, build, and operate core platform systems used daily by ML researchers and product engineers.
- Partner closely with research to turn recurring pain into durable infrastructure.
- Own reliability, performance, and developer experience for the systems in your lane.
- Ship iteratively in a flat, high-ownership environment. Measure impact, then raise the bar.
Requirements
- Strong background in systems / infrastructure software engineering and enjoy building platforms other engineers depend on.
- Owned production distributed systems at meaningful scale (ingestion, data pipelines, scheduling/orchestration, or similar).
- Comfortable across Linux, cloud and/or bare metal, and modern orchestration (Kubernetes, Ray, or equivalent).
- Like working closely with ML researchers and product engineers.
- Thrive where ownership is high and the feedback loop is short.
Especially strong backgrounds by team:
- Telemetry: event ingestion, product analytics pipelines, OpenTelemetry / tracing, reliable data APIs.
- Product Data Platform: data frameworks, Spark / Flink / Ray, ML dataset and training-data infrastructure.
- Observability: experiment / run monitoring, debug and eval tooling, agent-friendly observability UX.
- ML DevX and Systems: GPU / cluster scheduling, job queues, node health, research compute developer experience.