AI Engineer - Data Platform

New YorkFullTimePosted Aug 6, 2026

About the Role

Join a well-funded, Series A AI startup building the next generation of autonomous site reliability engineering for the enterprise. Backed by top-tier investors and trusted by some of the largest companies in the world, this team is tackling one of the hardest problems in AI: autonomously detecting, diagnosing, and remediating complex production incidents in real time.

As an AI Engineer on the Data Platform team, you'll design, build, and maintain the backend systems that power an AI-driven observability platform. This hands-on role blends distributed systems engineering, low-level system design, performance optimization, observability, and AI integration — across both cloud and on-premises deployments.

What You'll Do

  • Architecture & Implementation: Contribute to the design and implementation of scalable, resilient infrastructure systems powering AI-driven root cause analysis and observability workflows, including on-premises deployment environments.

  • Low-Level System Design: Work on the foundational building blocks of the infrastructure, ensuring efficient resource utilization and high performance at scale.

  • Performance Optimization: Profile and tune backend systems to improve throughput, reduce latency, and eliminate bottlenecks across the stack.

  • Observability Systems: Build and maintain the internal observability stack — logs, metrics, and traces — used by AI agents to understand and act on production issues.

  • Hybrid Infrastructure: Support cloud and on-premises architecture to serve both SaaS and enterprise customer deployment models.

  • Cross-functional Collaboration: Work closely with engineers across the company to deliver resilient infrastructure that enables AI agents to diagnose and remediate production incidents in real time.

What We're Looking For

  • Experience: 2–5 years of hands-on backend or infrastructure engineering experience.

  • Distributed Systems: Strong understanding of distributed systems design principles and trade-offs.

  • Performance Engineering: Proven experience profiling and optimizing high-throughput, low-latency systems.

  • Observability: Familiarity with observability tooling and concepts (logs, metrics, traces); experience with platforms such as Datadog, Grafana, Splunk, or similar is a plus.

  • Cloud & On-Prem: Experience with hybrid or multi-environment infrastructure (cloud + on-premises).

  • AI/ML Integration: Interest in or experience building systems that support AI/ML workloads at scale.

  • Background: Prior experience at observability, incident management, or data infrastructure companies is highly valued.

Note: Visa sponsorship is not available for this role.

Location

This is a fully on-site role based in New York, NY. Remote work is not available for this position.

Want jobs like this matched to you?

SimpleCareer scores fresh postings against your résumé so you only see the matches that matter.

Get started free