ARM Neon,ARM SVE,Performance optimization, library optimization-Senior Staff Engineer- CPU Hardware
Bangalore, IndiaPosted Jun 29, 2026
Skip to main contentOur CompanyOur BusinessSearch for JobsJoin Talent NetworkFAQsProfileEnglishSign InSingle PositionView All JobsARM Neon,ARM SVE,Performance optimization, library optimization-Senior Staff Engineer- CPU HardwareBangalore, India No longer accepting applications.Job ID3092371As a leading technology innovator, Qualcomm pushes the boundaries of what's possible to enable next-generation experiences and drives digital transformation to help create a smarter, connected future for all. As a Qualcomm Software Engineer, you will design, develop, create, modify, and validate embedded and cloud edge software, applications, and/or specialized utility programs that launch cutting-edge, world class products that meet and exceed customer needs. Qualcomm Software Engineers collaborate with systems, hardware, architecture, test engineers, and other teams to design system-level software solutions and obtain information on performance requirements and interfaces. Detailed JD:==========Job Description: CPU Software & Hardware Co-Design Engineer (ML Systems)Location: Bangalore (or relevant)Levels: Engineer / Senior Engineer / Staff / Principal EngineerRole OverviewWe are building a high-impact team at the intersection of CPU architecture, machine learning workloads, and system-level performance optimization. This role focuses on CPU software–hardware co-design for next-generation QMX architectures, including workload characterization, simulation, kernel optimization, and driving architectural insights for future CPU designs. The ideal candidate will work across the full stack—from ML models to low-level kernels to architectural feedback—enabling efficient execution of ML workloads on CPU platforms.Key Responsibilities1. ML Workload Identification & CharacterizationIdentify and prioritize critical ML use cases and models for CPU-centric execution (LLMs, vision, speech, recommender systems, etc.)Analyze workload characteristics including:Compute intensityMemory bandwidth and cache behaviorParallelism and dataflow patterns2. Simulation & Trace GenerationGenerate detailed execution traces for ML workloads using QEMU or equivalent simulatorsDevelop tooling to:Capture instruction-level execution behaviorExtract performance counters and bottlenecksEnable accurate modeling of workload behavior for architectural exploration3. Bottleneck Analysis & Performance OptimizationIdentify system bottlenecks across:CPU pipelinesMemory hierarchyInstruction utilizationOptimize critical hotspots through:Kernel-level tuningAlgorithmic improvementsData layout and memory optimizationsDrive measurable improvements in workload performance4. Software–Hardware Co-DesignCollaborate with CPU architecture and design teams to:Provide data-driven insights from real workloadsIdentify inefficiencies and propose architectural enhancementsInfluence next-generation CPU features in:Compute unitsVector/SIMD extensions (e.g., QMX)Memory subsystems5. ML Kernel & Library Development (QMX Focus)Design and implement highly optimized ML kernels and libraries for QMX architectureDevelop kernels for:GEMM, convolution, attention, activation functions, etc.Enable integration with:Open-source ML frameworks (e.g., PyTorch, ONNX, XNNPACK, MLAS)Apply advanced optimizations:SIMD/vectorizationCache-aware executionParallel execution strategies6. Benchmarking & Performance EngineeringOptimize CPU-centric ML benchmarks such as:Geekbench AIInternal benchmarking suitesEstablish performance baselines and track improvements across hardware generationsPerform competitive analysis and performance positioningRequired QualificationsStrong background in:Computer Architecture / Systems ProgrammingMachine Learning fundamentalsProficiency in:C/C++ (mandatory)Experience with:Performance profiling, benchmarking, and optimizationPreferred QualificationsExperience with:QEMU or equivalent simulatorsML kernel development (GEMM, convolution, attention)Knowledge of:CPU architecture (pipelines, caching,...