Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the full potential of modern AI. Astera Labs’ Intelligent Connectivity Platform integrates CXL®, Ethernet, NVLink, PCIe®, and UALink™ semiconductor-based technologies with the company’s COSMOS software suite to unify diverse components into cohesive, flexible systems that deliver end-to-end scale-up, and scale-out connectivity. The company’s custom connectivity solutions business complements its standards-based portfolio, enabling customers to deploy tailored architectures to meet their unique infrastructure requirements. Discover more at www.asteralabs.com.
Fabric Modeling and Analysis Engineer- Scale Up Fabric
Role Overview
Astera Labs is powering the connectivity behind rack-scale AI, and our Scorpio Scale-Up Fabric is central to how tomorrow’s GPU clusters scale. As a Fabric Modeling and Analysis Engineer, you will own the performance models that shape this fabric — quantifying bandwidth, latency, and throughput ceilings, and predicting how Scorpio hardware behaves under real AI/ML workloads long before silicon exists.
This is a high-impact role for an engineer who thrives where architecture, performance analysis, and software converge. Your models will shape the next generate products and enable informed architectural decisions ahead of tape-out, surface bottlenecks that only appear at scale, and translate directly into roadmap and IP decisions. As Astera Labs continues its hyper-growth, you’ll be the internal authority on fabric modeling.
Key Responsibilities
- Simulation Infrastructure & Model Correlation
- Design, implement, and maintain features in AI/ML system simulators and fabric models, extending support for collective communication, transport, topology, congestion, routing, buffering, scheduling, and traffic management
- Improve simulator fidelity, scalability, debuggability, and runtime performance across packet-level, flow-level, and analytical backends, and build reusable abstractions and APIs for end-to-end simulation flows
- Continuously calibrate and validate models against benchmark data from the Performance Engineering team, owning the correlation between model predictions and measured silicon across scale up fabric generations
- Analytical & Fabric Performance Modeling
- Design and implement transaction-level or cycle-approximate system models of the Scale up fabric, enabling rigorous evaluation of new architectural ideas early in the design cycle, well before tape-out
- Develop and own theoretical roofline and analytical models that establish bandwidth, latency, and throughput ceilings for the Scale-Up fabric, mapping AI/ML workload demands against fabric capabilities to identify compute-bound vs. fabric-bound regimes
- Model AI/ML collective communication patterns (AllReduce, AllGather, ReduceScatter) and next-generation features such as advanced congestion control, In-Network Computing, and novel resiliency at scale
- Workload Characterization, Scalability & Topology Analysis
- Model fabric performance as clusters scale from tens to thousands of GPUs, evaluating topological choices and configurations and their impact on latency, bandwidth, and congestion
- Proactively identify architectural bottlenecks that would otherwise only surface at scale and propose solutions in the in-house network model
- Innovation, Competitive Analysis & Roadmap Influence
- Build analytical models of competing fabric architectures and evaluate emerging interconnect technologies — such as UALink, ESUN, and optical fabrics — to assess their impact on future Scorpio generations
- Translate modeling insights into architectural innovation proposals and patentable IP, and propose future switch features to architecture and engineering teams
- Serve as the internal authority on fabric modeling, partnering with ASIC architects and product teams to drive roadmap and feature decisions, support early customer engagement, and produce modeling reports, white papers, and presentations for engineering, product, and executive audiences
Basic Qualifications
- Bachelor’s degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field
- 8+ years of relevant background in network modeling, system architecture, and scaleup fabric benchmarking and analysis
- Strong programming skills in Python and/or C++ for building and analyzing system models
- Hands-on experience with cycle-approximate or event-driven system modeling and simulation
- Deep experience designing, enhancing, and maintaining fabric models using established frameworks such as Astera-Sim, Garnet, NS-3, ATLAHS, BookSim, or in-house simulators
- Solid understanding of high-speed interconnect and switching fundamentals, including the physical layer, and protocols such as PCIe (Gen 6/7), ESUN, UEC or UALink
- Experience characterizing performance across bandwidth, latency, buffering, and congestion in network model or cluster
Preferred Qualifications
- MS or PhD in Electrical Engineering, Computer Engineering, or Computer Science
- Familiarity with AI/ML collective communication (AllReduce, AllGather, ReduceScatter) and large-scale LLM training/inference traffic patterns
- Background correlating models against emulation or silicon measurements, and analyzing competitive or emerging fabric architectures
- Strong communication skills with a track record of influencing architectural decisions and contributing to IP/patents
Salary range is $207,000 to $230,000 depending on experience, level, and business need. This role may be eligible for discretionary bonus, incentives and benefits.
We know that creativity and innovation happen more often when teams include diverse ideas, backgrounds, and experiences, and we actively encourage everyone with relevant experience to apply, including people of color, LGBTQ+ and non-binary people, veterans, parents, and individuals with disabilities.