Software Engineer, Artificial Intelligence/LLM (Multiple Seniority Levels)
San Carlos, CA · San CarlosSoftware Engineer$135k–$260kPosted Jul 15, 2026
Software Engineer, Artificial Intelligence/LLM (Multiple Seniority Levels) LocationSan Carlos - HybridEmployment TypeFull timeDepartmentEngineeringCompensation$135K – $260K • Offers Equity • The base salary range listed reflects the full scope of this position, which is open to multiple levels (L3–L6), and is informed and defined through professional-grade salary surveys and compensation data sources. The actual offer will be determined based on the candidate's experience, background, and the level at which they are hired.We use clear benchmarks to provide fair salaries and meaningful equity, ensuring compensation aligns with your skills and contributions. Beacon is financially healthy, well-funded, and ready to reward those who help achieve our ambitious goals. We use Pave to keep up to date with market pay rates and make proactive adjustments.About Beacon AIWe’re a fast-moving team of aviators, engineers, and operators building an AI platform to make flying safer, more efficient, and more capable. Backed by top investors, we’ve secured a dozen Department of Defense contracts and partnered with major airlines to deliver mission-critical systems. We operate without silos or heavy processes. Small, focused teams own what they build, ship quickly, and learn fast, pushing the boundaries of how humans and AI work together in aviation.You will ship LLM-powered product features end-to-end. That means designing retrieval and tool-calling flows, writing the services that run them, building evals and guardrails, and watching cost, latency, and quality in production. You’ll partner with the ML/infra teammates on embeddings, indexing, and model hosting, and with the product teammates on user experience and outcomes. We move fast, and we care about reliability in a safety-critical domain.We’re hiring across levels. Senior engineers own features and services. Staff engineers own systems, standards, and cross-team technical direction.What you’ll doBuild user-facing LLM featuresDesign and implement retrieval-augmented generation and tool-calling flows using frameworks like LangChain or equivalent primitives, where simpler is better.Deliver robust JSON and schema-bound outputs with validation, retries, and fallbacks.Add function calling to integrate with internal tools, search, routing, and data services.Own the service layerShip APIs and workers in Python or TypeScript with clear contracts, streaming, and backoff.Add caching, request shaping, prompt templates, and context packing to control latency and cost.Integrate with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints as needed.Retrieval and data prepCollaborate with infrastructure teammates to develop chunking, embeddings, and indexing capabilities for documents, time series, and multimedia.Choose and tune vector backends such as OpenSearch, pgvector, or Pinecone.Keep knowledge bases fresh with data syncs from S3, Aurora, DynamoDB, and external sources.Evaluation and qualityCreate offline evals and golden sets for prompts, retrievers, and tools.Stand up online metrics for task success, hallucination rate, retrieval precision/recall, p95 latency, and cost per request.Run A/B tests and prompt/version rollouts with guardrails and canaries.Safety, privacy, and complianceImplement content and policy checks, PII detection and redaction, access controls, and auditing.Design human-in-the-loop paths for sensitive actions.Handle aviation data with care and follow internal security standards.Operate what you buildAdd tracing, logs, and dashboards for model calls, token usage, errors, and saturation.Debug tricky failures across retrieval, prompts, tools, and providers.What will make you successfulShipped LLM apps: You’ve put LLM features in front of users and improved them with data.Strong builder: Comfortable writing production code, tests, and docs. You keep things simple and observable.RAG and tools depth: You understand embeddings, chunking, vector search tradeoffs, and function...