Member of Technical Staff - AI Platform Engineer

Patronus AI
Patronus AI

Software Engineering, IT, Data Science

San Francisco, CA, USA

USD 125k-200k / year + Equity

Posted on Sep 3, 2026

About Patronus AI

Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence.

We are the team behind some of the earliest and most influential research in AI evaluation like FinanceBench, Lynx, SimpleSafetyTests, CopyrightCatcher, Humanity’s Last Exam, and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe. We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more.

Responsibilities

The AI Platform Engineer sits between research and engineering, turning research workflows into services that other teams can self-serve. The work centres on building and operating the several internal platforms that the research and engineering teams rely on daily. These are production products with real users — backends, dashboards, CLIs and SDKs — built for colleagues inside the company rather than for an abstract audience.

The second half of the role is the foundation those platforms stand on: deploying models and keeping them served, both on managed inference providers and on internally operated GPUs, together with the cluster services that make GPU capacity self-serve rather than a matter of negotiation. This is a hands-on, delivery-oriented platform position rather than a research one: the emphasis falls on the systems and tooling that make other teams' work possible. Model training is part of the surrounding environment and remains accessible, but it is not the focus of the position.

In this role, you will:

  • Building and operating the internal platforms end to end — backends, storage, dashboards, and the CLI and SDK surfaces they are driven through — including multi-tenancy, sign-in and access control.
  • Owning workload orchestration on the GPU clusters and the services around them — submission and scheduling, provisioning, quotas and placement.
  • Deploying models and keeping them served, on managed inference providers and on internally operated GPUs — sizing each deployment for its hardware, and writing serving wrappers where no off-the-shelf engine fits.
  • Building the evaluation surface, so that a result stays comparable across months, colleagues and models.
  • Establishing and hardening CI/CD, containerized workflows and release safety.
  • Instrumenting the platform with logging, metrics and alerting, so that divergence between what a service promises and what it serves is caught by a test rather than by a customer.
  • Adding agent surfaces to the platform, with whatever an agent resolves written back as auditable configuration.
  • Partnering with research engineers to productionize their experiments, and carrying production issues through to resolution — including on-call for systems built in this role.

Qualifications

"The number one qualification to succeed in this machine learning course is gumption” - John Lafferty, CS Professor at Yale

Above all, we look for a proactive mindset, willingness to learn, unlimited energy, and relentless optimism. You are a great fit if you have a background in the following:

  • Several years of hands-on experience building and operating production applications.
  • Strong backend engineering in production-grade Python — API design, async services, relational databases and object storage.
  • Hands-on experience deploying and serving models in production, on managed inference providers and on self-operated GPUs, with at least one modern LLM serving stack.
  • Experience running workloads on shared GPU clusters through a scheduler such as Slurm, and working knowledge of cloud GPU infrastructure.
  • Experience with CI/CD, containerized workflows and Kubernetes.
  • Experience with MLOps and LLMOps tooling — model registries and hubs such as HuggingFace, experiment tracking and monitoring such as Weights & Biases, and deployment telemetry.
  • Experience with production observability and on-call operation — logging, metrics and alerting.
  • Experience with modern agentic frameworks.
  • A BS, MS, or PhD in Computer Science, Machine Learning, Software Engineering, or a related quantitative field.
  • Clear written and verbal communication, and experience collaborating cross-functionally with research, product and platform teams.
  • Strong integrity, sound judgment, and respect for others.

To support close collaboration, this role is based in our San Francisco headquarters and requires in-office attendance 5 days a week.
The expected base salary range for this role is $125,000 - $200,000 USD. In addition to base salary, we offer equity and benefits. Actual compensation will be determined based on experience, qualifications, skills, and location.

Benefits

  • Competitive salary and equity packages
  • 15 days of paid vacation per annum
  • Parental & sick leave
  • Health, dental, and vision insurance plans
  • 401(k) plan + matching
  • In-office lunch & dinner
  • Whoop band
  • Monthly meal stipend
  • Monthly health and wellness stipend
  • Equinox membership
  • Fun global offsites!

Patronus AI is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.

By clicking ‘Apply’, you agree to Greenhouse's Terms of Service and Privacy Policy.

By clicking 'Apply', you agree to Patronus AI, Inc. Privacy Policy.