Elm Street’s IDX Broker team is seeking an experienced Senior DevOps Engineer to drive reliability, scalability, and modernization across our distributed systems. You will manage complex infrastructure in a hybrid Google Cloud and AWS environment, using Kubernetes/GKE, Docker, Terraform, and Helm to support mission-critical SaaS applications.
This role is ideal for someone who combines deep technical expertise in cloud infrastructure, networking, and security with a strong commitment to site reliability engineering. You will help optimize CI/CD pipelines, integrate AI-assisted workflows, strengthen operational performance, and modernize the architecture for future growth.
What You’ll Do
- Architect, implement, and maintain scalable deployment infrastructure and CI/CD pipelines that support continuous delivery.
- Drive platform modernization by automating workflows, integrating AI-assisted DevOps tools, and optimizing infrastructure as code with Terraform.
- Lead incident response, performance tuning, and capacity planning to ensure the availability and reliability of critical SaaS applications.
- Proactively diagnose and resolve complex issues across distributed systems, with a focus on service dependencies, observability, and log analysis.
- Partner with Software Engineering, QA, and Product teams to advance reliability, security, and the developer experience.
- Champion best practices in security, networking, and cloud architecture to maintain robust and compliant application environments.
- Organize and lead complex infrastructure projects from conception through completion, communicating progress, risks, and timelines to stakeholders.
- Coordinate cross-functional response efforts and guide contributors when their expertise is needed to resolve complex issues.
What You’ll Bring
- Proven experience managing production-grade infrastructure in Google Cloud Platform (GCP) or Amazon Web Services (AWS).
- Deep expertise in container orchestration using Kubernetes/GKE and Helm, along with strong experience using Docker.
- Demonstrated proficiency with infrastructure as code, particularly Terraform.
- Experience building and managing CI/CD pipelines using tools such as Argo CD and Bitbucket Pipelines.
- Strong fluency in at least one scripting or automation language, such as Python.
- Demonstrated experience applying SRE principles, including incident response, performance tuning, and capacity planning for large-scale, business-critical SaaS applications.
- Proven ability to troubleshoot distributed systems using observability, log analysis, and a strong understanding of service dependencies.
- Ability to prioritize and manage complex technical projects while clearly communicating status, risks, and timelines.
- Strong written and verbal communication skills, including the ability to explain technical concepts to non-technical audiences.
Preferred Experience
- Familiarity with PHP and the Laravel framework.
- Experience with relational and search databases, including MySQL, PostgreSQL, and Elasticsearch.
- A solid understanding of networking fundamentals, security best practices, and web application performance optimization.
Location
This is a fully remote position open to candidates located anywhere in the United States.
Compensation
The anticipated base salary range for this position is $150,000–$200,000, depending on experience and qualifications. This role is also eligible for an annual bonus and equity.