AI Performance Modeling Engineer
About Quadric
Quadric is redefining edge AI with the industry's first General Purpose Neural Processing Unit (GPNPU), enabling developers to run both neural network inference and conventional C++ code on a single programmable architecture. Our technology powers intelligent edge devices across automotive, industrial, robotics, and embedded systems.
Founded by technologists from MIT and Carnegie Mellon, Quadric is a well-funded growth-stage semiconductor IP company with a growing licensing business. As we enter our next phase of growth, we're looking for our first true marketing leader to build and scale the function.
The Opportunity
Quadric has created an innovative General-Purpose Neural Processing Unit (GPNPU) architecture. Unlike standard accelerators, the Quadric GPNPU executes both neural network graph code and conventional C++ DSP/control code across edge and endpoint devices.
As an AI Performance Modeling Engineer, you will build analytical, cycle-level performance models of AI inference workloads on our next-generation architecture in Python before silicon exists. These models directly guide team decisions on hardware lane bindings, tensor placement, and architecture trade-offs. We welcome candidates across all experience levels—from early-career engineers to seasoned experts—with direct mentorship provided to help you master mapping complex workloads (like LLMs) onto our custom hardware.
What You'll Do
Performance Modeling & Architectural Analysis
- Build analytical, cycle-level Python models of AI inference workloads executing on next-generation GPNPU hardware.
- Derive from first principles which hardware lanes operations bind on (compute, on-chip/external memory bandwidth, interconnect) and model software pipelining overlaps.
- Model tensor placement, tiling across processing elements, local memory residency, and data movement across memory tiers.
- Model sharding and collective boundary communication across multi-die systems.
Workload Adaptation & Technical Writing
- Incorporate architectural details across vision networks and Large Language Models (LLMs), including operator mix, sparsity, routing, and quantization/low-precision numeric formats.
- Calibrate performance models against an instruction-set simulator and profiling traces to meet stated accuracy targets.
- Write and defend technical studies presenting empirical evidence that directly informs architecture and product decisions.
- Balance single-stream latency against scaled throughput performance.
What Success Looks Like
Within your first 6–12 months, you'll:
- Own a full workload's model end to end, calibrated against simulation and trusted by the engineering team.
- Build performance models that consistently predict workload behavior within 10–15% of actual measurements.
- Publish a written study whose defended conclusions directly shape an architecture or product decision.
- Review and extend performance models beyond your initial starting domain.
What We're Looking For
Required
- Python & Quantitative Modeling: Strong Python skills with experience writing, validating, and calibrating numerical or quantitative models in code.
- Computer Architecture Fundamentals: Solid grasp of memory hierarchies, bandwidth/latency trade-offs, pipelining, and execution bottlenecks (via industry experience, coursework, or research).
- Technical Writing: Comfort writing clear technical studies that state and defend evidence-based conclusions.
- Core Technical Depth (One of the following):
- Option A: Deep understanding of NN inference operators and tensor shapes (e.g., Transformers, attention mechanisms, MoE, prefill/decode split).
- Option B: Proven performance modeling experience in another quantitative/technical domain.
- Education: BS, MS, or Ph.D. in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.
Preferred
- Prior experience with GPUs, custom AI accelerators, CUDA, or Triton kernels.
- Familiarity with roofline analysis, back-of-the-envelope estimation, or architecture simulators (e.g., gem5, Timeloop, MAESTRO, Accel-Sim).
- Background in compiler internals (cost models, autotuners) or proficiency in C++.
- Published performance studies or technical write-ups.
What We Offer
The base salary range for this position is $150,000 to $200,000. This range reflects the full span of levels and geographies at which Quadric hires for this role. The actual base salary offered will depend on a number of factors, including the specific level of the role, years and depth of relevant experience, technical skills and competencies, the criticality of the role to the business, internal equity, and work location. In addition to base salary, this role is eligible for equity and a discretionary annual performance bonus as applicable to the role and level.
In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings; for roles in other locations, benefits vary and are shared during the hiring process. These include:
- Medical, dental, and vision insurance from day one - Premiums covered at 99% for Employees
- Company-paid life Insurance
- Voluntary supplemental life insurance
- STD + LTD insurance
- Commuter support including parking or Caltrain reimbursement. Our office is conveniently located within walking distance of the Caltrain station
- FSA + HSA
- Equity with the business
- Paid Parental Leave
- 401(k) Retirement Plan
- Flexible PTO
- Winter holiday shutdown
- Catered lunch each day in our office
- Downtown Burlingame office location, close to shops, cafes, and local amenities
- Collaborative, low-ego culture with significant ownership and impact
- A work culture focused on innovative disruption
Founded in 2016 and based in downtown Burlingame, California, Quadric is building the world’s first supercomputer designed for the real-time needs of edge devices. Quadric aims to empower developers in every industry with superpowers to create tomorrow’s technology, today. The company was co-founded by technologists from MIT and Carnegie Mellon, who were previously the technical co-founders of the Bitcoin computing company 21.
Quadric is proud to be an equal opportunity employer. We are committed to creating an inclusive environment where people from all backgrounds can do their best work. We consider all qualified applicants without regard to race, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, or any other protected characteristic under applicable law.
If this role resonates with you, we encourage you to apply even if your experience does not perfectly match every qualification. We value potential, curiosity, and a willingness to learn just as much as direct experience. Skills and growth come in many forms, and we would love to hear your story.
By submitting an application, you acknowledge that Quadric will collect and process your personal information as part of the hiring process. Please review our Privacy Policy to understand how we handle your data.