About the company
Yugabyte is the company behind YugabyteDB, the AI-ready, multi-modal, distributed PostgreSQL database for cloud-native apps. Trusted by industry leaders including Shopify, Paramount+, GM, Kroger, Fiserv, and NPCI, YugabyteDB has been deployed in over 100 countries and powers more than 5 million clusters worldwide.
Our Yugabeings (distributed, like our database) span 12+ countries and multiple time zones, sharing expertise from diverse backgrounds and industries.
Role Overview
As a Staff Engineer on the Distributed Storage & Transactions (DST) team, you will play a critical role in designing, building, and scaling the core distributed storage, replication, and transaction foundations of YugabyteDB.
You will work on some of the most challenging problems in distributed systems — including consistency, durability, fault tolerance, performance, and scalability — while collaborating closely with query layer, platform, and cloud teams.
Responsibilities
- Lead the design, development, testing, and delivery of core storage and replication features in YugabyteDB.
- Write high-quality C/C++ code with comprehensive automated tests; actively participate in design discussions and code reviews.
- Troubleshoot and resolve correctness, stability, and performance issues in complex distributed storage and transactional subsystems.
- Improve database scalability and throughput as cluster sizes, data volumes, and transaction rates continue to grow.
- Build and streamline database management operations, including horizontal cluster scale-out, incremental and point-in-time backups, online schema and index operations, rolling upgrades and blue-green deployments.
- Identify and implement performance improvements across the storage engine, transaction processing, and replication layers.
- Contribute to the open-source YugabyteDB project, helping evolve its storage architecture and operational reliability.
- Mentor and technically influence other engineers in distributed systems design, performance engineering, and systems-level debugging.
Requirements
- 8+ years of professional software engineering experience, with a strong foundation in systems programming using C/C++.
- Bachelor’s, Master’s, or PhD in Computer Science (or related field), or equivalent practical experience.
- Deep understanding of distributed systems fundamentals, including replication and consensus, transactions and consistency models, fault tolerance and recovery.
- Experience working on storage engines, databases, or other infrastructure-level systems.
- Strong problem-solving skills and the ability to operate effectively in a collaborative, distributed team environment.
Preferred Skills
- Hands-on experience with distributed storage systems, transactional engines, or consensus protocols.
- Familiarity with LSM-tree based storage engines, WALs, snapshots, or compaction strategies.
- Experience with PostgreSQL internals or other relational database engines is a strong plus.
- Prior contributions to open-source systems or database projects.