As a member of our team, you will participate in developing innovative solutions across AI Platform. Collaboration with engineers and researchers to build and optimize training infrastructure and tools for LLMs, SLMs, multimodal, and code-specific models. Design, build and improve services with high scalability and reliability. Design and implement the services to serve the prod traffic and fulfill the security and privacy requirements. Participate in efforts to deliver and improve engineering systems and practices to ensure service quality in complex cloud environments. Contribute to the deployment and monitoring of services in production environments. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. 5+ years of software engineering experience, with significant ownership of production services, cloud platforms, distributed systems, or developer infrastructure. Strong experience building and operating containerized platforms using Kubernetes or similar orchestration systems. Strong coding skills in one or more systems or backend languages such as Python, Go, Rust, C++, C#, or Java. Experience designing reliable production APIs, backend services, or control-plane systems that manage compute, storage, networking, or runtime environments. Solid understanding of cloud infrastructure fundamentals, including identity, networking, storage, observability, capacity planning, security, and safe deployment practices. Experience diagnosing production issues using logs, metrics, traces, dashboards, and incident response processes. Demonstrated ability to lead technical design and retrospective, drive ambiguous projects to completion, mentor other engineers, and collaborate across teams. Experience building multi-tenant platforms where reliability, fairness, quota management, isolation, and security are important. Experience with sandboxed execution environments, remote development environments, hosted notebook/tool environments, evaluation infrastructure, or ephemeral compute platforms. Experience with container image build systems, registry authentication, image caching, package caching, artifact distribution, or startup-latency optimization. Experience with cloud networking concepts such as ingress, DNS, proxies, egress control, private endpoints, service routing, and traffic management. Experience with secure runtime design, including authentication, authorization, workload identity, secret handling, network isolation, and protecting shared infrastructure from untrusted workloads. Experience with AI infrastructure, agent execution, evaluation platforms, GPU workloads, Windows/Linux runtime environments, or VM/container hybrid systems. Experience improving service operability through structured logging, distributed tracing, dashboards, alerting, automated validation, and incident playbooks.
Want jobs like this matched to you?
SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.