Design, develop, and operate distributed services that power Azure Storage Geo Replication capabilities Build solutions that improve data durability, regional resiliency, failover readiness, and disaster recovery experiences Collaborate with engineers, architects, and product teams to define and deliver new functionality. Participate in the design, implementation, testing, and deployment of cloud-scale services Analyze service telemetry and customer scenarios to identify reliability and performance improvements. Diagnose and resolve complex production issues involving replication, storage, and distributed systems Act as a Designated Responsible Individual (DRI) and contribute to operational excellence and service health Leverage AI-assisted engineering tools to improve productivity, code quality, troubleshooting effectiveness, and operational efficiency. Continuously improve service availability, observability, scalability, and performance Bachelor's Degree in Computer Science, Software Engineering, or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, or Java. Master's Degree in Computer Science, Software Engineering, or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, or Python OR Bachelor's Degree in Computer Science, Software Engineering, or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, or Python OR equivalent experience 6+ years of software development, including building products and services, scalable distributed systems, and cloud services Designing and building large-scale distributed systems Working with cloud infrastructure, storage systems, databases, or replication technologies Distributed systems concepts, including durability, consistency, availability, and fault tolerance Troubleshooting complex production issues in large-scale cloud environments Multithreaded and concurrent programming Data structures, algorithms, testing, and performance optimization Geo-distributed architectures, disaster recovery, and business continuity solutions Leveraging AI technologies for code generation, testing, debugging, documentation, and development productivity AI-assisted operational workflows, anomaly detection, telemetry analysis, and incident response automation Incorporating AI tools into software engineering workflows while maintaining high standards of quality and reliability
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.