Site Reliability Engineer
SpainSite Reliability Engineer (SRE)€58k–€97kPosted Jul 15, 2026
Site Reliability EngineerSpainEngineering – Platform /Remote /Remoteapply for this job
About Tinybird
At Tinybird, we help developers and data teams unlock the power of real-time data. This enables them to build data pipelines and innovative data products quickly. With Tinybird, you can seamlessly ingest multiple data sources at scale, query them using the SQL you already know, and publish results as low-latency, high-concurrency APIs for your applications. Developers can create fast APIs; what used to take hours or days now takes only minutes. Tinybird is the essential tool that data engineers and software developers have been waiting for, making it easier to drive innovation.
About the Platform team
The Platform team builds, operates, and continually improves the technical foundations that Tinybird relies on.
We help Tinybird run safely at scale, evolve predictably, and support product and customer growth with minimal operational friction. This involves working on the systems that ensure reliability, observability, performance, infrastructure, cost efficiency, CI/CD, development environments, and critical backend services.
Platform is both an infrastructure team and an operations team. We focus on making Tinybird's foundations reliable, observable, scalable, and easier to evolve.
What we are looking for
We seek an experienced Site Reliability Engineer who enjoys keeping large-scale distributed systems reliable and adaptable as they grow. You should understand how to make hardware and software work well together and be eager to grasp both our product and the real challenges our customers and internal teams face.
You might be a good fit if:
You have strong experience in designing, building, and running distributed cloud architectures and large-scale web-based production systems.
You have deep knowledge of Kubernetes, which is essential for this role. You should be comfortable designing and operating production-grade clusters, writing custom controllers or operators as needed, and tuning autoscaling mechanisms (KEDA, Karpenter, and similar) to respond to real-time workloads. You know how Kubernetes manages networking, storage, scheduling, and resources, and you can analyze performance and failure scenarios at scale.
You are skilled in AWS and GCP.
Coding skills are required. We are not looking for a software developer, but baseline. You should be able to explore our codebase, ClickHouse source code, or any other software we use to understand how things work. Our primary languages are Python and some C++.
You are comfortable operating close to production: debugging incidents, understanding system behavior, improving observability, and enhancing service reliability.
You think in systems and pay attention to edge cases, failure modes, and specific implementation details.
You care about performance, reliability, cost efficiency, and operational simplicity. You prioritize action, iteration, and delivery. You know many decisions can be reversed quickly, and that speed is important in business and technology.
You take ownership, follow through, and are willing to tackle issues that may be broken, because you can fix them if necessary.
You enjoy data and SQL, and you are curious about how real-time analytical systems work. We use our own product, so you'll need some SQL experience to query our own data. Experience with ClickHouse and/or launching database systems at scale would be a big plus.
Familiarity with Traefik, Varnish, Redis, Terraform, or Ansible is not mandatory, but it's helpful and will get you up to speed quickly. We don't expect anyone to know the full stack upon arrival.
You communicate clearly in writing. This is important because we work asynchronously, document decisions, write operational notes, and share context across teams.
You use AI tools such as Claude Code, Cursor, ChatGPT, and others to enhance efficiency and improve your workflows.
You are fluent in English and Spanish. English is the primary...