Site Reliability Engineer jobs in San Francisco – Browse 5,209 openings on RoboApply Jobs

Site Reliability Engineer jobs in San Francisco

Open roles matching “Site Reliability Engineer” with location signals for San Francisco. 5,209 active listings on RoboApply Jobs.

5,209 jobs found

1 - 20 of 5,209 Jobs
Apply
companyDrata logo
Full-time|$166.9K/yr - $225.9K/yr|Hybrid|Hybrid - San Francisco

Drata helps organizations demonstrate their commitment to security and integrity. The platform supports companies as they build and maintain trust with users, customers, partners, and prospects. Values Built on Trust: Consistency shapes decisions and actions. Integrity: Choosing to do what is right, every time. Customer-Obsessed: Prioritizing customer needs above all else. Competitive Fire: Striving for higher standards and greater achievements. Diversity: Welcoming different perspectives to encourage creative solutions. Automation First: Pursuing efficiency by saving time and resources wherever possible. How the Team Works Drata blends high standards with a supportive environment focused on growth. Team members are encouraged to own their work, improve continuously, and deliver meaningful results. The company values quick, informed decisions that drive immediate impact, while always keeping the mission and customer needs at the center. The San Francisco-based team uses a hybrid work model. Colleagues collaborate in the office Tuesday through Thursday, focusing on alignment and innovation. Mondays and Fridays offer flexibility for deep work or personal needs. Growth and Culture Drata has expanded to over 600 professionals worldwide, recognized for a culture that values trust, speed, and continuous learning. The environment supports both personal and professional development. See the Speed: CEO Adam Markowitz discusses Drata’s rapid journey to $100M ARR in four years. Hear the Voice of the Team: Employee stories highlight collaboration and growth at Drata.

Apr 27, 2026
Apply
companyCognition logo
Full-time|On-site|San Francisco Bay Area

Join Our TeamAt Cognition, we are at the forefront of applied AI innovation, developing cutting-edge software agents that redefine the engineering landscape. Our flagship products, Devin, the pioneering AI software engineer, and Windsurf, an AI-native IDE, embody our commitment to creating AI that collaborates with engineers as a true partner.Our team is composed of elite talent including competitive programming champions, visionary founders, and researchers from top AI institutions such as Scale AI, Palantir, Cursor, Google DeepMind, and more.Your MissionAs a Site Reliability Engineer, you will play a crucial role in ensuring the reliability of our user-focused products, which are utilized by hundreds of thousands of developers daily. Your mission is to preemptively address potential issues and swiftly resolve any incidents that may arise, maintaining a seamless experience for our users.You will be responsible for overseeing production reliability and enhancing our platform engineering practices, encompassing SLOs, incident response, and on-call duties, alongside CI/CD pipelines, deployment infrastructure, and developer tools. At Cognition, we believe in integrating reliability into our systems rather than treating it as an afterthought, and we strive to cultivate a culture that reflects this philosophy.Your AchievementsProduction Reliability: Establish and manage SLOs, SLIs, and error budgets for our products. Develop robust monitoring, alerting, and observability systems to maintain a transparent view of service health.Incident Management: Spearhead incident response with precision and promptness. Conduct blameless postmortems to derive actionable insights from outages, and create effective runbooks and tools to enhance on-call sustainability.Platform Engineering: Oversee deployment pipelines and internal developer tools, ensuring rapid, reliable shipping of code while minimizing unnecessary toil for engineers.Infrastructure as Code: Manage cloud infrastructure via code, creating reproducible, auditable environments that can scale with product demands and mitigate configuration drift.Capacity Planning: Analyze growth trends, anticipate resource requirements, and ensure our infrastructure is always ahead of user demand, optimizing system performance proactively.Security and Reliability: Integrate security protocols with reliability practices to create a robust framework that safeguards our infrastructure.

Oct 13, 2025
Apply
companyfal logo
Full-time|On-site|San Francisco

Join our dynamic team at fal as a Senior/Staff Site Reliability Engineer. In this key role, you will leverage your expertise to enhance our systems' reliability and performance. If you are passionate about building scalable systems and enjoy working in a collaborative environment, we want to hear from you!

Feb 23, 2026
Apply
companyBaseten logo
Full-time|On-site|San Francisco Office

ABOUT BASETENBaseten is at the forefront of powering mission-critical AI inference for some of the most innovative companies globally, including Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. We integrate cutting-edge applied AI research with a flexible infrastructure and intuitive developer tools to empower companies at the leading edge of AI to deploy sophisticated models effectively. With our recent $300M Series E funding round—supported by prominent investors such as BOND, IVP, Spark Capital, Greylock, and Conviction—we are rapidly expanding. Join our dynamic team and contribute to creating an essential platform for engineers to launch AI products with ease.THE ROLEAs a Site Reliability Engineer, you will design and implement resilient systems and processes that ensure our infrastructure is scalable, reliable, and efficient. Your responsibilities will encompass everything from automating deployments and monitoring systems to enhancing performance and managing incidents effectively.Collaboration is key; you will work closely with our users to understand their challenges in operationalizing machine learning, facilitating their onboarding onto our platform, and leveraging these insights to inform improvements to Baseten.EXAMPLE INITIATIVESAs part of our Infrastructure team, you will engage in exciting projects such as:Innovative multi-cloud capacity managementOptimizing inference on B200 GPUsImplementing multi-node inferenceUtilizing fractional H100 GPUs for efficient model servingRESPONSIBILITIESDesign and maintain scalable infrastructures to support the deployment and operational needs of machine learning models.Establish standards and best practices to enhance reliability and performance across the infrastructure.Proactively identify and resolve reliability issues using monitoring and alerting systems.Collaborate with cross-functional teams to apply best practices in infrastructure management and incident response.Create automation scripts to streamline processes and reduce manual intervention.

Oct 9, 2025
Apply
companyHive logo
Full-time|On-site|San Francisco

About HiveHive stands at the forefront of cloud-based AI innovation, providing cutting-edge solutions that enable organizations to understand, search, and generate content. Our platform is relied upon by some of the world's most prestigious and forward-thinking companies. We empower developers with an extensive suite of state-of-the-art, pre-trained AI models that handle billions of API requests each month. In addition to our robust model offerings, we deliver comprehensive software applications backed by proprietary AI models and datasets, unlocking transformative applications in various sectors such as content moderation, brand protection, sponsorship measurement, and context-based advertising.With over $120 million in funding from esteemed investors like General Catalyst, 8VC, Glynn Capital, Bain & Company, and Visa Ventures, Hive has cultivated a vibrant global team of over 250 employees across our San Francisco, Seattle, and Delhi offices. If you’re passionate about shaping the future of AI, we invite you to join our dynamic team!DevOps and Systems TeamIn response to our distinctive machine learning demands, we have developed our own data centers focusing on distributed high-performance computing with GPU integration. While we harness the power of these data centers, our infrastructure remains hybrid, leveraging public cloud solutions when advantageous. As we scale our machine learning models for commercial use, we are expanding our DevOps and Site Reliability team to ensure the reliability of our enterprise SaaS offerings. Our ideal candidate thrives in dynamic environments, embraces automation, and believes that every task can be automated and every server can scale. You take pride in enhancing performance across all layers of our stack and are committed to never performing the same task manually twice.

Apr 20, 2022
Apply
companyHyperbolic Labs logo
Full-time|On-site|San Francisco, CA

Who We AreAt Hyperbolic Labs, we are committed to democratizing AI by removing barriers to computing power with our Open-Access AI Cloud. By aggregating global computing resources, we provide an innovative GPU marketplace and AI inference service that ensures both affordability and accessibility. As trailblazers at the convergence of AI and open-source technology, we envision a future where AI innovation is only limited by creativity, not by resource availability. We invite forward-thinking individuals who share our dedication to making AI universally accessible, secure, and affordable. Join us in crafting a platform that empowers innovators worldwide to realize their visionary AI projects.In anticipation of our growth following our Series A funding, our team — guided by co-founders with advanced degrees in AI, Mathematics, and Computer Science — is set to transform the computing landscape.About the RoleWe are in search of a skilled Site Reliability Engineer to guarantee that Hyperbolic's GPU marketplace and AI infrastructure function with outstanding reliability, performance, and security. As an aggregator of computational resources from numerous global providers, our service level objectives (SLOs), trust, and economic efficiency are critical to our product. Your key responsibilities will include defining and maintaining service level objectives, developing resilient incident response protocols, managing capacity across our extensive GPU network, and implementing secure rollout and rollback mechanisms to ensure uninterrupted platform operation around the clock.In this influential role, you'll set the reliability benchmarks that foster customer trust in our platform, design comprehensive monitoring and alerting systems for enhanced infrastructure visibility, automate capacity management and resource allocation processes, lead incident response and post-mortem evaluations, and collaborate closely with engineering teams to bolster system resilience. Security and infrastructure hardening will be paramount, necessitating strong isolation protocols between tenants and suppliers, the implementation of effective key management systems, and the establishment of compliance frameworks. This high-impact position will directly affect our ability to deliver on our commitment to providing affordable, accessible AI compute at scale.

Mar 26, 2026
Apply
companySierra logo
Full-time|On-site|San Francisco, CA

About UsAt Sierra, we are pioneering a transformative platform that empowers businesses to forge authentic customer experiences through AI technology. Headquartered in the vibrant city of San Francisco, we also boast a dynamic presence in Atlanta, New York, London, France, Singapore, and Japan.Our operations are anchored in core values that shape our culture: Trust, Customer Obsession, Craftsmanship, Intensity, and Family. These principles guide our actions and are integral to our mission.Our visionary founders, Bret Taylor and Clay Bavor, bring unparalleled expertise. Bret, currently the Board Chair of OpenAI, previously co-led Salesforce and served as CTO at Facebook, while Clay led numerous initiatives at Google, including AR/VR projects and Google Workspace.Your RoleIn your capacity as a Software Engineer on the Site Reliability team, you will play a crucial role in establishing and enhancing the reliability, observability, and scalability of Sierra’s AI-centric infrastructure. Collaborating closely with our engineering and product teams, your goal is to ensure our systems remain highly available, efficient, and primed for growth.Lead the development of Sierra’s observability stack—including monitoring, alerting, logging, and tracing—to provide engineers with critical insights into system health and performance.Collaborate with product and platform engineers to architect systems that prioritize reliability and scalability from the outset, not as an afterthought.Design and implement robust, scalable, and secure cloud infrastructure on AWS, employing Terraform and cutting-edge DevOps tools.Enhance the reliability and scalability of our LLM deployments, ensuring they operate efficiently and cost-effectively.Drive improvements in deployment pipelines, CI/CD tooling, and incident management processes to minimize downtime and accelerate response times.Define and cultivate SRE practices within Sierra, shaping culture, tooling, and best practices across the engineering organization.QualificationsBachelor's degree in Computer Science or a related field, or equivalent experience.Proven experience in Site Reliability Engineering or a similar role, with a strong understanding of cloud infrastructure (AWS).Proficiency in Terraform and modern DevOps practices.Experience with observability tools and techniques—monitoring, alerting, logging, and tracing.Strong problem-solving skills with a focus on scalability and performance optimization.Excellent collaboration and communication skills, with the ability to work effectively in a team environment.

Oct 21, 2025
Apply
companyalembic logo
Full-time|On-site|San Francisco HQ

About the RoleJoin alembic as a Senior Site Reliability Engineer (SRE) and become an integral part of our mission to enhance platform reliability, observability, and operational excellence. In this pivotal role, you will collaborate with engineers and data scientists to architect, automate, and maintain the robust infrastructure that drives our platform, including data pipelines, machine learning workloads, and real-time analytics systems.This hands-on position offers significant visibility across the technology stack and provides you with the opportunity to shape the future of our infrastructure and operations.

Dec 22, 2025
Apply
companyVeeam Software logo
Full-time|On-site|San Francisco Bay, CA, USA

Join Veeam Software as a Site Reliability Engineer III, where you'll be at the forefront of ensuring the reliability, scalability, and performance of our software solutions. You will leverage your expertise in system administration and programming to improve our infrastructure and automate processes, making Veeam a leader in cloud data management.

Mar 22, 2026
Apply
companyAndromeda Cluster logo
Full-time|Remote|Global Remote / San Francisco, CA

Site Reliability Engineer - AI InfrastructureLocation: Global Remote / San Francisco · Full-TimeAbout AndromedaAndromeda Cluster, established by Nat Friedman and Daniel Gross, aims to democratize access to advanced AI infrastructure for early-stage startups, previously exclusive to hyperscalers. Our journey began with a single managed cluster that quickly reached capacity, propelling us to develop robust systems, networking, and orchestration layers to make AI infrastructure more accessible than ever.Today, we collaborate with top AI laboratories, data centers, and cloud service providers to deliver compute resources precisely when and where they're needed the most. Our platform efficiently manages the routing of training and inference jobs across a global supply chain, facilitating flexibility and efficiency in one of the most rapidly expanding markets worldwide.Our vision is to create a liquidity layer for global AI compute — a marketplace that dynamically moves the infrastructure and workloads essential for AGI, akin to the capital flows in global financial markets.We are on the lookout for talented individuals who excel in AI infrastructure, research, and engineering to join our pioneering team.Your ResponsibilitiesProvision, configure, and manage Kubernetes clusters for clients across various service providers.Develop automation tools to enhance the deployment and integration of clusters.Troubleshoot customer issues related to networking, storage, scheduling, and system layers.Enhance the reliability and scalability of training and inference infrastructures.Design and implement monitoring, alerting, and observability solutions for critical systems.Work collaboratively with engineering and product teams to strategize and deliver infrastructure for new services.Engage in on-call duties and incident response, leading postmortems and reliability enhancements.Ideal Candidate ProfileA minimum of 5 years of experience in Site Reliability Engineering (SRE), DevOps, or infrastructure engineering roles.Solid foundation in Linux systems and networking principles.Extensive expertise in Kubernetes and container orchestration at scale.Proficient in Infrastructure-as-Code methodologies (Terraform, Helm, etc.).

Nov 6, 2025
Apply
companyOkta, Inc. logo
Full-time|$194K/yr - $267K/yr|On-site|San Francisco, California

Discover OktaOkta is recognized as The World’s Identity Company, empowering individuals to securely leverage any technology across various devices and applications. Our versatile Okta Platform and Auth0 Platform provide reliable access, authentication, and automation, placing identity at the forefront of business security and expansion.At Okta, we value diverse perspectives and experiences. We seek continuous learners and individuals who can enhance our team with their distinct backgrounds.Join us as we create a world where identity is truly yours.We are in search of a highly skilled Observability Site Reliability Engineer specializing in Google Cloud, to take charge of and elevate our Observability ecosystem within GCP. In this position, you will progress beyond basic monitoring to develop a world-class, comprehensive, and scalable Observability Platform that supports our SRE teams and business collaborators. You will implement infrastructure as code by employing Terraform and demonstrating strong coding skills in Go, Python, or Ruby to automate the deployment of agents and collectors across intricate distributed systems.Key ResponsibilitiesAutomated Infrastructure: Design, build, and maintain scalable observability infrastructure utilizing tools such as Terraform.GCP Observability Engineering: Enhance the collection, processing, and storage of Observability data to guarantee high reliability and low latency for our Splunk and Grafana services.Incident Response: Engage in on-call rotations and conduct post-incident reviews to foster systemic improvements and promote 'observability-driven development.'Automation: Minimize 'toil' by automating the deployment and scaling of observability agents and collectors.

Mar 11, 2026
Apply
companyTubi TV logo
Full-time|$227.2K/yr - $324.5K/yr|Hybrid|San Francisco, CA (Hybrid)

About the Role: At Tubi, our Site Reliability Engineering (SRE) team transcends traditional operations. We embody a software engineering ethos, leveraging a developer's toolkit to tackle the complexities of large-scale, distributed systems. Our core mission focuses on building resilience from the ground up, empowering our product teams to innovate swiftly while delivering an exceptional user experience. We oversee the availability, latency, performance, and capacity of our platform, driven by a culture of data-informed decision-making, blameless learning, and relentless automation. We are on the lookout for a seasoned and visionary Senior Manager of SRE to lead and expand our newly formed Site Reliability Engineering team. You will be more than just a people manager or tech lead; you will be the strategic architect behind our reliability roadmap. Your role will involve building and mentoring a team of skilled engineers, cultivating an environment of blameless learning and continuous improvement, while advocating for the engineering practices that balance rapid innovation with unwavering stability. You will play a pivotal role within our engineering leadership, collaborating with peers across the organization to embed reliability as a shared responsibility and a fundamental principle of our engineering culture.

Mar 17, 2026
Apply
companyOpenAI logo
Full-time|On-site|San Francisco

About Our TeamThe Frontier Systems team at OpenAI is at the forefront of technological innovation, responsible for designing, deploying, and maintaining state-of-the-art supercomputers that power our most advanced model training initiatives. We transform innovative data center designs into fully functional systems and develop the necessary software to support extensive frontier model training.Our mission is to ensure the stability and efficiency of these hyperscale supercomputers, providing an uninterrupted environment for the training of frontier models.About the OpportunityWe are seeking passionate engineers to manage the next generation of compute clusters that fuel OpenAI’s leading-edge research. This role merges distributed systems engineering with practical infrastructure expertise across our expansive data centers. You will be tasked with scaling Kubernetes clusters to unprecedented levels, automating bare-metal deployments, and creating software solutions that simplify interactions across a multitude of nodes in various data centers.You will operate at the confluence of hardware and software, where speed and reliability are of utmost importance. Prepare to oversee dynamic operations, swiftly diagnose and resolve critical issues, and continuously enhance automation and system uptime.Key Responsibilities:Deploy and scale substantial Kubernetes clusters, implementing automation for provisioning, bootstrapping, and lifecycle management.Create software abstractions that integrate multiple clusters, delivering a seamless interface for training workloads.Oversee node deployment from bare metal to firmware upgrades, ensuring swift and repeatable processes at scale.Enhance operational metrics, striving to minimize cluster restart times (e.g., reducing from hours to minutes) and expedite firmware or OS upgrades.Integrate networking and hardware health systems to ensure comprehensive reliability across servers, switches, and data center infrastructure.Develop monitoring and observability systems that proactively identify issues and maintain cluster stability under peak loads.Be prepared to perform at the level of a software engineer in execution and problem-solving.You May Be a Great Fit If You:Possess extensive experience in operating or scaling Kubernetes clusters or similar container orchestration systems.

Nov 3, 2025
Apply
companyCarta logo
Full-time|On-site|San Francisco, California; Santa Clara, California; Seattle, WA

Join Carta as a Senior Site Reliability Engineer, where you will play a pivotal role in enhancing our infrastructure and ensuring the reliability of our platforms. You will work collaboratively with cross-functional teams to implement innovative solutions that drive operational excellence and scalability.

Apr 3, 2026
Apply
companyMercor logo
Full-time|On-site|San Francisco

Join the Mercor TeamAt Mercor, we stand at the dynamic intersection of labor markets and AI research. Collaborating with premier AI labs and enterprises, we empower the human intelligence that is crucial for AI's evolution.Our expansive talent network plays a vital role in training cutting-edge AI models, akin to the way educators impart knowledge to their students—by sharing insights, experiences, and contextual understanding that code alone cannot convey. Currently, our network of over 30,000 experts generates more than $2 million daily.We are pioneering a novel category of work where expertise fuels AI progress. Achieving this vision necessitates an ambitious, fast-paced, and deeply dedicated team. You will collaborate with researchers, operators, and AI firms that are at the forefront of transforming societal structures.Mercor is a thriving Series C company with a valuation of $10 billion. We operate five days a week in-person at our new headquarters in San Francisco.About the RoleAs a Site Reliability Engineer (SRE) at Mercor, you will take ownership of production reliability for our critical systems, working closely with our infrastructure leadership. You will play a pivotal role in establishing our SRE function and defining how Mercor manages large-scale, high-availability systems.Your ResponsibilitiesEnsure the reliability and safety of production for key shared services and customer-facing systems.Collaborate directly with infrastructure leadership to outline SRE priorities, reliability benchmarks, and the production safety roadmap.Enhance the structure of our production systems to ensure stability, resource efficiency, isolation, and observability.Advocate for and implement modern SRE methodologies (e.g., incident management, postmortems, SLIs/SLOs) across engineering teams.Work alongside engineering and applied AI teams to facilitate sustainable growth.Promote SRE best practices internally, supporting teams in a safe, scalable, and consistent production onboarding process.Who We SeekThe ideal candidate will have:Extensive experience in genuine SRE roles (not merely operations) across various positions or organizations.A deep understanding of SRE methodologies popularized by Google (e.g., error budgets, reliability vs. risk trade-offs, large-scale distributed systems).5+ years of SRE experience; ideally, 15+ years in total experience for this inaugural SRE position.A proven track record of managing systems at scale, with a strong grasp of the complexities involved.

Dec 27, 2025
Apply
company
Full-time|$214K/yr - $260K/yr|Hybrid|Hub - San Francisco

At Superhuman, we embrace a vibrant hybrid work model that offers our team members the ideal blend of focused individual work and collaborative in-person interactions, fostering trust, innovation, and a robust team culture.About SuperhumanSuperhuman, the AI productivity platform, is on a transformative mission to unlock the superhuman potential within everyone. With the integration of Grammarly's writing assistance and innovative tools like Coda’s collaborative workspaces and Go, our proactive AI assistant, we empower over 40 million individuals and 50,000 organizations globally. Founded in 2009, we strive to eliminate busywork and enhance productivity. Discover more at superhuman.com and explore our values here.The OpportunityTo meet our ambitious goals, we are seeking a Site Reliability Engineer (SRE) to join our infrastructure team. This pivotal role focuses on developing software solutions to maintain the reliability of our back-end systems while collaborating with engineering teams to strategize our future growth. You will also engage with our production engineering teams in Europe as we transition from a “you build it, you own it” approach.At Superhuman, our engineers and researchers enjoy the autonomy to innovate and drive breakthroughs, directly impacting our product roadmap. As we rapidly scale our interfaces, algorithms, and infrastructure, the complexity of our technical challenges is growing. Learn more about our technical endeavors on our technical blog.As an SRE, your responsibilities will include:Scaling our Kubernetes-based control plane that processes billions of events each day.Enhancing our automation mechanisms to efficiently respond to workload demands.Deploying machine learning systems across various departments.

Jun 18, 2025
Apply
companyAndromeda logo
Full-time|Remote|Global Remote / San Francisco, CA

Join Andromeda as a Senior Site Reliability Engineer specializing in AI Infrastructure. In this pivotal role, you will be responsible for ensuring the reliability, scalability, and performance of our cutting-edge AI systems. Collaborate with cross-functional teams to design and implement robust infrastructure solutions that support our innovative AI initiatives. Your expertise will play a crucial role in maintaining optimal service availability and improving system performance.

Apr 9, 2026
Apply
companyHeartFlow, Inc. logo
Full-time|$200.8K/yr - $250.9K/yr|On-site|San Francisco, California

About HeartFlow HeartFlow, Inc. is a medical technology company focused on improving the diagnosis and management of coronary artery disease. Our flagship product, the AI-powered HeartFlow FFRCT Analysis, provides a non-invasive, color-coded 3D view of a patient’s coronary arteries. Clinicians use our platform to identify blockages, assess blood flow, and analyze atherosclerosis, all in alignment with ACC/AHA Chest Pain Guidelines. HeartFlow’s technology supports care teams in the US, UK, Europe, Japan, and Canada, and has already impacted over 500,000 patients worldwide. As a publicly traded company (NASDAQ: HTFL), HeartFlow continues to expand its product line and modernize its platform to support the next generation of life-saving medical technologies. Role Overview: Staff/Lead Site Reliability Engineer (SRE) HeartFlow is searching for an experienced Site Reliability Engineer to join the cloud-native infrastructure team in San Francisco, California. This role works closely with Platform engineers and development teams to maintain and improve the reliability, scalability, observability, and performance of critical systems. What You Will Do Collaborate with Platform and development teams to ensure system reliability and performance Automate complex operational processes and reduce manual work Establish and promote standards for production excellence Support ongoing Platform Modernization initiatives Who We’re Looking For Extensive experience as a Site Reliability Engineer or in a similar role Strong background in cloud-native infrastructure Interest in automation, reliability, and scalable systems Comfort working with cross-functional engineering teams Location This position is based in San Francisco, California.

Apr 14, 2026
Apply
companyZyphra logo
Full-time|On-site|San Francisco

Zyphra is a cutting-edge artificial intelligence firm located in the heart of San Francisco, California.The Opportunity:As a Product Infrastructure Engineer specializing in Site Reliability, your primary focus will be on architecting and sustaining the frameworks that ensure Zyphra's infrastructure remains strong, observable, secure, and scalable. Your contributions will be pivotal in guaranteeing the dependability and reproducibility of machine learning workloads, managing deployment safety, and ensuring the long-term viability of our computational environments.Your Responsibilities:Enhancing and developing observability systems (monitoring, logging, alerting)Creating resilient build and deployment systems across both research and production settingsEstablishing secure release protocols with comprehensive audit trails and rollback capabilitiesCollaborating closely with ML engineers, DevOps, and infrastructure teams to optimize system reliability and performanceLeading incident response efforts, conducting root-cause analysis, and facilitating postmortems with a strong emphasis on learning and preventionThis position is perfect for individuals who are passionate about creating systems that empower other teams to be faster, safer, and more efficient.Qualifications:Proven experience in high-performance computing environments, such as machine learning clusters or GPU farmsStrong background in infrastructure as code tools (e.g., Ansible, Terraform)Familiarity with software release engineering tailored for ML/AI systems is advantageousExperience in designing reliable environments for experimental workloads and reproducible executionsUnderstanding of compliance and auditing standards related to deployment and system securityExperience with load testing, fault injection, and chaos engineering to strengthen systems under pressureA passion for developing tools that render infrastructure seamless and reliable for end usersPreferred Qualifications:Experience with infrastructure as code (e.g., Ansible, Terraform)Previous experience supporting ML/AI infrastructure, including GPU management and workload optimizationExposure to backend development for ML model serving (e.g., vLLM, Ray, SGLang)

Aug 22, 2025
Apply
companyMongoDB, Inc. logo
Full-time|$127K/yr - $249K/yr|Hybrid|United States

The TeamJoin our dynamic Platform Engineering team within Site Reliability Engineering (SRE), which is tasked with maintaining vital infrastructure and operational functions that empower our engineering organization. We manage multi-cloud Kubernetes infrastructures, deployment systems, and observability frameworks.The Fabric team specializes in ensuring secure communication between systems and the public internet. We focus on network architecture, service mesh, and edge load balancing, safeguarding customer data during transit. Our work is essential in building and sustaining a reliable, globally-connected multi-cloud network for MongoDB products.This position is available in our New York City headquarters, smaller offices in Austin, Palo Alto, and San Francisco, or as a fully remote role from anywhere in North America. Our hybrid work model accommodates both in-office and remote work.

Apr 8, 2026

Sign in to browse more jobs

Create account — see all 5,209 results

Tailoring 0 resumes

We'll move completed jobs to Ready to Apply automatically.