---
title: Staff AI Platform Engineer at Code Metal in Boston, MA
description: Code Metal is hiring for the Staff AI Platform Engineer role in Boston, MA. Pays
  $205K–$250K per year. See the full description and apply.
type: job
url: https://www.foundrole.com/jobs/staff-ai-platform-engineer-at-code-metal-01a10f1a-7d3b-7510-a985-e5a336e0cefd
date: 2026-10-06T03:16:23Z
og_description: Join Code Metal as Staff AI Platform Engineer in Boston, MA. Pays $205K–$250K
  per year. Full-time, hybrid role.
og_image: https://www.foundrole.com/og/wna3b8.png
breadcrumbs:
  - label: Home
    url: https://www.foundrole.com/
  - label: Search
    url: https://www.foundrole.com/jobs
---

| | |
|---|---|
| **Company** | Code Metal |
| **Location** | Boston, MA (Hybrid) |
| **Salary** | $205K/yr - $250K/yr |
| **Type** | Full Time |
| **Posted** | Oct 05, 2026 |

## Remote check

**Fully remote: passed our check.** The role itself is remote, with no recurring office days, travel under 10%, no on-site duties and no end date on the remote setup.

- "Flexible hybrid or remote work arrangement"

**Where you can work from:** Not stated in the posting. Ask the recruiter before you apply.
## Description

# About Code Metal

Code Metal is the leader in automated software engineering you can trust. As AI writes more of the world's code, the bottleneck in software has shifted from writing code to verifying it works, and AI cannot verify its own work with certainty. Code Metal takes a fundamentally different approach: constrain AI to what it does reliably, verify every step independently of the model using formal methods, and keep engineers in the loop on the decisions that matter. The result isn't code that probably works — it's code that is provably correct, with auditable proof. Customers including the U.S. Air Force, L3Harris, RTX, and Toshiba use Code Metal to modernize legacy code, optimize performance on real hardware, and move prototypes to production, fast. Founded in 2023 with offices in Boston and San Francisco, Code Metal is funded by Accel, Salesforce Ventures, B Capital, Smith Point Capital, J2 Ventures, Shield Capital, Overmatch, RTX, and others.  
Learn more at [codemetal.ai](http://codemetal.ai).

**The Role**

Code Metal's engineering teams are building AI-driven code transpilation and AI-enabled mission planning and wargaming. Both need the same foundations: models to serve, agents to run, context to manage, and results to measure. Our AI Platform team builds those foundations.

As a Staff AI Platform Engineer, you'll be the technical lead of this new four-person team. You'll architect and build the AI enablement stack our engineers depend on, from GPU inference serving and a model gateway up through agent harnesses, context engineering, observability, and AI experimentation management. It starts as an internal platform, but we're building it to product standard.

This is an engineering role first. Most of your time goes to designing, building, and operating production systems. You'll also need solid data science and AI research fundamentals: you'll work closely with our Applied AI Research team, and you'll sometimes run experiments yourself when a platform decision needs evidence.

**Core Responsibilities**

- Set the technical direction and architecture for Code Metal's AI platform and lead the team building it. Own the design docs and RFCs, help with build-vs-buy decisions, and mentor the team.

- Deploy, benchmark, and tune production inference for open-weight models on vLLM, SGLang, and TensorRT-LLM.

- Own the model gateway that teams use to reach self-hosted and commercial models, with consistent auth, routing, failover, quotas, and cost attribution.

- Design reusable agent harnesses and orchestration primitives that product teams can compose into reliable, verifiable workflows instead of rebuilding them for each product.

- Build context-engineering services for memory, retrieval, and data discovery, so agents get the right information within their context and cost budgets.

- Instrument the stack end to end with OpenTelemetry traces and service metrics, and build the experiment-tracking and artifact layer that lets engineers and researchers reproduce and compare results.

- Design for productization from day one (multi-tenancy, versioned APIs, security, and deployment in customer and air-gapped environments), and partner with Applied AI Research, product teams, and DevOps so the platform stays aligned with what they need.

**Required Qualifications**

- Production-grade Python and strong platform engineering fundamentals: API and service design, distributed systems, containers and Kubernetes, CI/CD, and testing.

- Shipped production agentic systems, with a clear sense of where they break and how to make them reliable.

- Experience with context engineering: retrieval-augmented generation, embeddings, vector or hybrid search, and memory for agents, ideally over code or large technical corpora.

- Experience instrumenting services (for example, with OpenTelemetry tracing and metrics) and operating AI services against SLOs.

- Solid data science and AI research fundamentals: how transformers and LLM inference work, experiment design, benchmarking, and model evaluation. Working familiarity with PyTorch and Hugging Face, and experience fine-tuning, evaluating, or serving language models.

- Staff-level technical leadership: owned architecture across multiple systems or teams, written design docs and RFCs, turned ambiguous needs from several internal customers into a roadmap, and mentored engineers.

**Preferred Qualifications**

- Highly desirable: Production experience running an LLM gateway or proxy such as SMG or Bifrost, or equivalent experience building an API gateway, including routing, auth, rate limiting, quotas, failover, and cost attribution.

- Hands-on experience deploying and tuning LLM inference engines such as vLLM, SGLang, or TensorRT-LLM on GPU infrastructure, with measurable gains in throughput, latency, or cost per token.

- Familiarity with inference optimization: speculative decoding, prefix caching, tensor/pipeline/expert parallelism, disaggregated prefill and decode, GPU profiling.

- Experience building evaluation harnesses for LLMs and agents, and experiment-tracking or artifact systems such as MLflow or Weights & Biases.

- Experience taking an internal platform to an external product: multi-tenancy, SDKs, versioned APIs, and documentation.

- Experience deploying AI systems on-prem or in air-gapped or classified environments, or in regulated domains such as defense or aerospace.

**Experience Level**

Typically 8+ years of software engineering experience, including 4+ years building and operating ML or LLM systems in production, with demonstrated Staff-level scope and impact. Equivalent depth of experience is valued over a rigid year count.

# Benefits

- Pay depends on experience, but we strive to be at the upper end of the salary range

- Health care plan with 100% premium coverage, including medical, dental, and vision

- 401k with 5% matching

- Paid Time Off (uncapped vacation, plus sick and public holidays)

- Flexible hybrid or remote work arrangement

- Relocation assistance for qualifying employees

*Wage Transparency - The salary range for this role is not a guarantee of compensation or salary, as the final offer amount may vary based on factors including, but not limited to, individual proficiency, skills, experience, and location.*

*We are an equal opportunity employer. US Citizenship may be required for certain project assignments involving security clearance.*

*Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.*
## Skills

- Distributed Systems
- Hybrid Search
- Vector Embeddings
- Security Clearance
- Experiment Tracking
- DevOps
- Technical Leadership
- Air-Gapped Networks
- Mlflow
- GPU Profiling
- Experimental Design
- Data Science
- Authentications
- Caching
- Software Deployment
- Python
- Cost Allocation
- Technical Direction
- Agentic Systems
- Machine Learning Model Fine-Tuning
- API Gateway
- Service Design
- Artificial Intelligence
- LLM Inference
- Applied AI
- Model Evaluation (Machine Learning)
- Distributed Tracing
- SGLang
- Kubernetes
- Retrieval-Augmented Generation (Rag)
- TensorRT-LLM
- Inference Optimization
- Climbing Equipment
- New Product Development
- CI/CD
- Benchmarking
- Hugging Face
- Team Building
- Failover
- Classified Environments
- Large Language Models (Llm)
- Software Development Kit (Sdk)
- Context Engineering
- RFC Writing (Request for Comments)
- Technical Documentation
- Platform Engineering
- Multi-Tenant Architecture
- Product Roadmaps
- Proxy Servers
- Speculative Decoding
- OpenTelemetry
- Electrical Machines
- GPU Infrastructure
- Rate Limiting
- vLLM
- Budgeting
- Data Discovery
## Benefits

- Health Insurance
- Dental Insurance
- Vision Insurance
- Paid Time Off (Pto)
- 401(k) Plans
- Relocation Assistance
- Company Events

## Related

- [Data & Analytics jobs](https://www.foundrole.com/careers/data-and-analytics?utm_source=ai_markdown)- [Machine Learning & AI jobs](https://www.foundrole.com/careers/data-and-analytics/machine-learning-and-ai?utm_source=ai_markdown)- [Artificial Intelligence Engineer jobs](https://www.foundrole.com/careers/data-and-analytics/machine-learning-and-ai/artificial-intelligence-engineer?utm_source=ai_markdown)- [Browse all companies](https://www.foundrole.com/companies?utm_source=ai_markdown)

## Explore this job market

- [Software Products](https://www.foundrole.com/sectors/technology/software-products?utm_source=ai_markdown) — 117493 jobs
- [Massachusetts](https://www.foundrole.com/locations/us/massachusetts?utm_source=ai_markdown) — 77021 jobs
- [Boston, MA](https://www.foundrole.com/locations/us/massachusetts/boston?utm_source=ai_markdown) — 17853 jobs

## Related pages

- [Software Products](https://www.foundrole.com/sectors/technology/software-products?utm_source=ai_markdown) — 117493 jobs
- [Massachusetts](https://www.foundrole.com/locations/us/massachusetts?utm_source=ai_markdown) — 77021 jobs
- [Boston, MA](https://www.foundrole.com/locations/us/massachusetts/boston?utm_source=ai_markdown) — 17853 jobs

## Browse

- [Companies](https://www.foundrole.com/companies?utm_source=ai_markdown)
- [Jobs by country](https://www.foundrole.com/locations?utm_source=ai_markdown)
- [Jobs in the United States by state and city](https://www.foundrole.com/locations/us?utm_source=ai_markdown)
- [H1B sponsors by location](https://www.foundrole.com/h1b-sponsors?utm_source=ai_markdown)
- [H1B salaries](https://www.foundrole.com/h1b-salaries?utm_source=ai_markdown)
- [Sectors and industries](https://www.foundrole.com/sectors?utm_source=ai_markdown)
- [Jobs by career field](https://www.foundrole.com/careers?utm_source=ai_markdown)
- [Latest jobs](https://www.foundrole.com/jobs/latest?utm_source=ai_markdown)