Deep Learning Research Intern — Multimodal BEV Perception
We are seeking a highly motivated Deep Learning Research Intern to join our Perception team, working on Bird's Eye View (BEV) fusion models using multimodal sensor inputs, including camera, LiDAR, and radar. You will support the design and evaluation of scalable perception algorithms that enable 3D scene understanding for autonomous driving, working closely with research engineers on real-world, production-relevant problems. We are seeking a highly motivated Deep Learning Research Intern to join our Perception team, working on Bird's Eye View (BEV) fusion models using multimodal sensor inputs, including camera, LiDAR, and radar. You will support the design and evaluation of scalable perception algorithms that enable 3D scene understanding for autonomous driving, working closely with research engineers on real-world, production-relevant problems.
Responsibilities:
-
Research, prototype, and evaluate BEV-based perception models that fuse camera, LiDAR, and radar inputs.
-
Design and run experiments benchmarking model performance on large-scale autonomous driving datasets using well-defined quantitative metrics.
-
Investigate novel multimodal fusion architectures and techniques to improve accuracy, robustness, and efficiency.
-
Collaborate closely with research scientists and perception engineers to translate findings into production-ready components.
-
Document experiments, results, and insights clearly for the broader research team.
-
Perform work in accordance with the company's Quality Management System (QMS) requirements where applicable.
Required skills
-
Currently pursuing a Ph.D. or Master's degree in AI, Computer Science, Electrical Engineering, Robotics, or a related field.
-
Strong programming skills in Python and experience building deep learning pipelines.
-
Hands-on experience with PyTorch, TensorFlow, or JAX.
-
Solid foundation in computer vision, deep learning, and 3D geometry.
-
Coursework or project experience with LiDAR-based 3D perception or BEV representation models.
-
Understanding of multimodal sensor fusion concepts
Preferred Skills:
-
Prior research or project experience in autonomous vehicle perception, robotics, or related fields.
-
Familiarity with camera, LiDAR, and radar modalities, including synchronization, calibration, and integration in perception pipelines.
-
Experience with distributed training, high-performance computing, or GPU acceleration.
-
Publications or open-source contributions in perception, 3D vision, or multimodal learning.