---
title: Senior HPC & GPU Infrastructure Engineer at Sciforium
description: Sciforium is hiring for the Senior HPC & GPU Infrastructure Engineer role in San
  Francisco, CA. See the full description and apply.
type: job
url: https://www.foundrole.com/jobs/senior-hpc-and-gpu-infrastructure-engineer-at-sciforium-01a10e8d-42a9-78da-9bee-b633b70bf8af
date: 2026-10-06T00:18:57Z
og_description: Join Sciforium as Senior HPC & GPU Infrastructure Engineer in San Francisco, CA.
  Pays $150K–$220K per year. Full-time, on-site role.
og_image: https://www.foundrole.com/og/a90znc.png
breadcrumbs:
  - label: Home
    url: https://www.foundrole.com/
  - label: Search
    url: https://www.foundrole.com/jobs
---

| | |
|---|---|
| **Company** | sciforium.com |
| **Location** | San Francisco, CA |
| **Salary** | $150K/yr - $220K/yr |
| **Type** | Full Time |
| **Posted** | Oct 05, 2026 |
## Description

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

## **About the role**

We are seeking a Senior HPC & GPU Infrastructure Engineer to take full ownership of the health, reliability, and performance of our GPU compute cluster. You will be the primary custodian of our high-density accelerator environment and the linchpin between hardware operations, distributed systems, and machine learning workflows. This role spans everything from hands-on Linux systems engineering and GPU driver bring-up to maintaining the ML software stack (CUDA/ROCm, PyTorch, JAX, vLLM). If you love squeezing every bit of performance out of hardware, enjoy debugging GPUs at scale, and want to build world-class AI infrastructure, this role is for you.

## **What you'll do**

1. System Health & Reliability (SRE)

- **On-Call Response:** Act as the primary responder for system outages, GPU failures, node crashes, and cluster-wide incidents. Minimize downtime by resolving issues rapidly.  

- **Cluster Monitoring:** Implement and maintain monitoring for GPU health, thermal behavior, PCIe/NVLink topology issues, memory errors, and overall system load.  

- **Vendor Liaison:** Coordinate with data center staff, hardware vendors, and on-site technicians for repairs, RMA processing, and physical maintenance of the cluster.  

2. Linux & Network Administration

- **OS Management:** Install, patch, and maintain Linux distributions (Ubuntu / CentOS / RHEL). Ensure consistent configuration, kernel tuning, and automation for large node fleets.  

- **Security & Access Controls:** Configure VPNs, iptables/firewalls, SSH hardening, and network routing to secure our computer infrastructure.  

- **Identity & Storage Management:** Manage LDAP/FreeIPA/AD for user identity, and administer distributed file systems such as NFS, GPFS, or Lustre.  

3. GPU & ML Stack Engineering

- **Deployment & Bring-Up:** Lead deployment of new GPU nodes, including BIOS configuration, NUMA tuning, GPU topology validation, and cluster integration.  

- **Driver & Kernel Management:** Build and optimize kernel modules, maintain GPU drivers and runtime stacks for both NVIDIA (CUDA) and AMD (ROCm).  

- **Software Stack Maintenance:** Maintain and optimize ML frameworks and libraries PyTorch, JAX, CUDA toolkit, cuDNN, ROCm, NCCL, and supporting runtime systems.  

- **Advanced Debugging:** Troubleshoot complex interactions involving GPUs, compilers, ML frameworks, and distributed training runtimes (e.g., vLLM compilation failures, CUDA memory leaks, ROCm kernel crashes).

## **Ideal candidate profile**

- 5+ years of experience in HPC, GPU cluster operations, Linux systems engineering, or similar roles.

- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.

- Strong expertise with NVIDIA (H100/B200) or AMD (MI325x/MI355x) GPUs, including driver and kernel-level debugging.

- Deep understanding of Linux internals, kernel modules, hardware bring-up, and systems performance tuning.

- Experience with network security, including VPNs, iptables/firewalld, SSH, and identity management (LDAP/FreeIPA/AD).

- Proficiency in Bash and Python for scripting, automation, and workflow tooling.

- Familiarity with ML software stacks: CUDA toolkit, cuDNN, NCCL, ROCm, JAX/PyTorch runtime behavior.

- Deep debugging experience with NVLink/NVSwitch fabrics and RDMA networking.

## **Nice-to-have**

- Experience with job schedulers such as Slurm, Kubernetes, or Run:AI.

- Exposure to vLLM, model serving optimizations, or inference systems.

- Hands-on experience with configuration management tools (Ansible, SaltStack, Terraform).

- Previous experience supporting ML research teams in a startup or research-heavy environment.

## **Benefits include**

- Medical, dental, and vision insurance

- 401k plan

- Daily lunch, snacks, and beverages

- Flexible time off

- Competitive salary and equity

## **Equal opportunity**

Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
## Skills

- High-Performance Computing
- IpTables
- IBM Spectrum Scale (Gpfs)
- LDAP
- Debugging
- Scripting
- cuDNN
- Slurm Workload Manager
- Hardening
- PCI Express PCIe
- Access Controls
- Ubuntu
- Systems Engineering
- Software Deployment
- CentOS
- Python
- Automation
- Distributed File Systems
- Lustre File System
- ROCm
- Linux Kernel Development
- Performance Tuning
- Artificial Intelligence
- Data Centers
- Terraform
- Kubernetes
- Virtual Private Networks (Vpn)
- Network Routing
- Hardware Bring-up
- Network File Systems
- Configuration Management
- Ansible
- Red Hat Enterprise Linux
- Distributed Training (Machine Learning)
- Compilers
- Bash
- Remote Direct Memory Access (Rdma)
- JAX
- GPU (Graphics Processing Unit)
- Network Administration
- NVIDIA CUDA
- Computer Engineering
- Maintenance
- Secure Shell SSH Software
- Computer Science
- NVLink
- Linux
- Node.js
- Electrical Engineering
- On-Call Coverage
- GPU Cluster Management
- Machine Learning
- Model Serving
- Kernel Tuning
- vLLM
- Salt (Tools for Software Configuration Management)
- BIOS/Firmware Development
- Network Security
- Site Reliability Engineering (Sre)
- Firewall Configuration
## Benefits

- Health Insurance
- Dental Insurance
- Vision Insurance
- Paid Time Off (Pto)
- 401(k) Plans

## Related

- [Technology & Software jobs](https://www.foundrole.com/careers/technology-and-software?utm_source=ai_markdown)- [Cloud & DevOps jobs](https://www.foundrole.com/careers/technology-and-software/cloud-and-devops?utm_source=ai_markdown)- [DevOps Engineer jobs](https://www.foundrole.com/careers/technology-and-software/cloud-and-devops/devops-engineer?utm_source=ai_markdown)- [Browse all companies](https://www.foundrole.com/companies?utm_source=ai_markdown)

## Explore this job market

- [California](https://www.foundrole.com/locations/us/california?utm_source=ai_markdown) — 264806 jobs
- [San Francisco, CA](https://www.foundrole.com/locations/us/california/san-francisco?utm_source=ai_markdown) — 26836 jobs
- [Technology, Information and Internet](https://www.foundrole.com/sectors/technology/technology-information-and-internet?utm_source=ai_markdown) — 14869 jobs

## Related pages

- [California](https://www.foundrole.com/locations/us/california?utm_source=ai_markdown) — 264806 jobs
- [San Francisco, CA](https://www.foundrole.com/locations/us/california/san-francisco?utm_source=ai_markdown) — 26836 jobs
- [Technology, Information and Internet](https://www.foundrole.com/sectors/technology/technology-information-and-internet?utm_source=ai_markdown) — 14869 jobs

## Browse

- [Companies](https://www.foundrole.com/companies?utm_source=ai_markdown)
- [Jobs by country](https://www.foundrole.com/locations?utm_source=ai_markdown)
- [Jobs in the United States by state and city](https://www.foundrole.com/locations/us?utm_source=ai_markdown)
- [H1B sponsors by location](https://www.foundrole.com/h1b-sponsors?utm_source=ai_markdown)
- [H1B salaries](https://www.foundrole.com/h1b-salaries?utm_source=ai_markdown)
- [Sectors and industries](https://www.foundrole.com/sectors?utm_source=ai_markdown)
- [Jobs by career field](https://www.foundrole.com/careers?utm_source=ai_markdown)
- [Latest jobs](https://www.foundrole.com/jobs/latest?utm_source=ai_markdown)