About the Role
We're a seed-stage wearable AI company building an AI co-pilot for skilled field technicians — delivered through industrial smart glasses — that helps workers in high-stakes industries like data centers and energy infrastructure operate at an expert level. Our stack spans edge inference, real-time voice/video, and agentic visual reasoning running on real hardware in demanding environments.
We're looking for a Founding AI Engineer with 1–5 years of experience who has shipped production multimodal and agentic AI systems to real users. This is a hands-on, end-to-end ownership role at the core of our product — not a research or prototyping position.
What You'll Do
Build and ship the production agentic Vision-Language Model (VLM) pipeline running on industrial smart glasses — multi-step, tool-using visual-reasoning loops against real customer workflows (SOPs, inspection, field service).
Own model orchestration and runtime optimization for edge inference, balancing model quality against latency with graceful degradation across variable connectivity conditions.
Design and build the evaluation harness and data flywheel from scratch — failure-mode capture, customer-data fine-tuning loops, and measurable quality improvements each release cycle.
Ship real-time voice and video AI interfaces tailored to different end-user profiles: video-heavy, conversational speech, and proactive alerting.
Build RAG pipelines for efficient creation and querying of enterprise knowledge bases from field operator data.
Drive multimodal model training for on-premise deployments: open-source model SFT, RL post-training, and quantization.
What We're Looking For
Required (Dealbreakers):
Demonstrable track record shipping production multimodal and computer vision systems in the VLM era, owning the model layer end-to-end — with hands-on expertise in visual-language and/or video-language VLMs/VLAs.
Bachelor's degree in Computer Science, Machine Learning, Engineering, or equivalent — graduated 2018 or later.
Willingness to work on-site 5 days/week in San Francisco, CA.
Also Required:
Experience with applied agentic AI or model orchestration in production settings.
Experience building production AI products at a startup or high-ownership AI team, or relevant big-company experience (AR/smart glasses, real-time video/streaming, on-device/edge ML) paired with a strong builder signal (e.g., early startup, side projects, open-source contributions).
Experience with rigorous evaluation methodologies — ground-truth evals, trajectory evals, tool-call accuracy, and regression testing for comparing models and orchestration stacks.
Strong foundation in CS, ML, or engineering, or a demonstrated equivalent through shipping history.
Nice to Have:
Experience with production AR or wearable AI (e.g., AR headsets, mixed reality platforms) or autonomous driving computer vision.
On-prem / self-hosted model deployment, including serving and optimizing open-weight models on customer hardware, or hands-on fine-tuning and deployment of vLLMs.
Exposure to industrial domains such as data centers, energy grids, aerospace, or manufacturing.
Experience with in-context grounding or RAG against a knowledge base, including tool and knowledge base wiring.
Master's degree with a vision or multimodal research component.
Tech Stack Includes: Python, PyTorch, vLLM, Triton, Ray Serve, ONNX, TensorRT, Hugging Face Transformers, LangChain, RAG, RLHF, SFT, Quantization (GPTQ, AWQ), Edge AI, Multimodal LLMs, Vision-Language Models, Docker.
Compensation & Benefits
Salary: $180,000 – $240,000 USD annually
Early-stage equity as a founding team member
Visa sponsorship: Not available
Location
On-site, 5 days/week in San Francisco, CA. This is not a remote or hybrid role.