Senior ML / AI Engineer
Remote R&D - Research and Development
Job Type
Full-time
Description

About The Role


We're building the AI/ML enablement backbone for Nonstop Health - a shared platform on which every AI initiative ships as a governed, production-grade service. The flagship work today is an agentic platform for healthcare claims (EOB parsing, claim substantiation against member policy + IRS/MERP rules) built on self-hosted LLMs/VLMs, a multi-agent framework, and a drag-and-drop Agent Studio - all under HIPAA / SOC 2 / ISO 27001 controls. You'll own hard parts of this platform and stand up new initiatives on top of the same foundation.


This is a builder role for someone equally comfortable reasoning about an agent's failure modes, tuning a retrieval pipeline, and making a service survive an RDS failover. You care that AI systems handling sensitive, money-moving decisions are correct, grounded, observable, and defensible - not just demo-able.


What You'll Do

  • Design and ship agentic systems - dynamic tool-calling agents (LangGraph + structured Pydantic outputs) that reason at runtime over a governed tool catalog, with per-step verification, grounding checks, LLM-as-judge, human-in-the-loop, and bounded, auditable loops.Build the RAG layer - document ingestion ? chunking ? embeddings ? hybrid retrieval ? reranking, with strict per-member/per-client access scoping and injection/poisoning defenses.
  • Operate self-hosted inference at scale - vLLM (chat/vision) and an embeddings/rerank service on EKS GPU nodes; optimize throughput (continuous batching, prefix caching, quantization, KV-cache tuning) and enforce fair-share concurrency across tenants.
  • Own the MLOps / governance plane - offline eval harnesses and golden sets, quality scoring, canary/gray-release with auto-rollback, cost/budget governance, and end-to-end observability (OpenTelemetry traces + metrics, immutable audit).
  • Make it production-safe - durable state, circuit breakers, retries, PHI-safe logging/redaction, RBAC, fail-closed defaults, and horizontal scalability on AWS (EKS, RDS/pgvector, SQS, S3/KMS, Cognito) via Terraform.
  • Extend the platform to new domains - take a new AI initiative from problem framing to a governed, evaluated, deployed service on the shared foundation, and mentor engineers on doing the same.
  • Raise the bar - clean, typed, tested code (the platform holds a ruff + mypy + full-test-green bar on every change); thoughtful design reviews; pragmatic trade-offs between autonomy and determinism for high-stakes tasks.
Requirements

What we're looking for (Required)

  • 5+ years building and shipping ML/AI systems in production (not just notebooks or POCs), including hands-on LLM application work in the last 1–2 years.
  • Strong Python (typed, tested, production-grade) and solid software-engineering fundamentals; comfortable across an async web service, a data layer, and infra.
  • Practical depth in LLM application patterns: prompting, structured/ function-calling outputs, RAG (embeddings, vector search, retrieval quality), agent/tool-use loops, and - critically - how to evaluate and de-risk them (grounding, hallucination control, eval sets, guardrails).
  • Cloud + MLOps experience on AWS (or equivalent): containers + Kubernetes, IaC (Terraform), CI/CD, observability, and cost/perf tuning of model-serving.
  • Track record of owning reliability: state durability, failure handling, scaling, and debugging production incidents.
  • Clear written communication and the judgment to make sensible calls under ambiguity.
  • Self-hosting/optimizing open-weight models (vLLM, TGI, or similar) on GPUs; embeddings/rerank serving (e.g., bge / Infinity).
  • LangGraph / LangChain or comparable agent frameworks; multi-agent orchestration.
  • Regulated-data experience - HIPAA / SOC 2 / ISO 27001, PHI/PII handling, RBAC, audit trails; healthcare/insurance/claims domain (X12 835, EOB, benefits) a strong plus.
  • pgvector / Postgres, MongoDB/DocumentDB, SQS, Cognito/OIDC.
  • OpenTelemetry, Prometheus/Grafana; frontend comfort (React/TypeScript) to extend an internal builder/console.
  • Amazon Bedrock or a multi-provider abstraction; on-prem/air-gapped deployment.

Tech You'll Work With


Python · FastAPI · LangGraph · Pydantic · vLLM · pgvector · MongoDB · Postgres · Redis · SQS · React/TypeScript · AWS (EKS + GPU, Cognito, RDS, S3/KMS, SQS) · Terraform · OpenTelemetry · Docker/Kubernetes


Why Join?


You'll work on genuinely hard, high-impact AI - systems that make sensitive decisions and therefore have to be right - with the mandate and the platform to do it well: real governance, real evals, real observability, and a team that treats "governed autonomy" as the goal, not an afterthought.


Compensation


The base salary range for this role is $150,000 - $210,000 depending on experience.


Great benefits aren't just for our clients! We cover 100% of medical, dental and vision benefits for employees and dependents. We also contribute to your retirement goals, with a 401k match up to 4%.


**Must be authorized to work for any employer in the United States. This job does not offer visa sponsorship.


Who We Are


Our main offering is a MERP that, when paired with a HDHP, offers first dollar coverage to employees at a lower premium cost. This means people have access to early care without the worry of a copay, coinsurance, or deductible. Ultimately, early care drives down overall cost and improves health and happiness. We think this is the way healthcare should work! Our primary sales channel is selling into brokers. We have a national presence with a growing book of business and expanded product offerings.


Our team at Nonstop is where our magic really comes to life. We are driven, curious, collaborative, passionate, mission oriented, and most of all - we are real people who love what we do and believe in the real impact we are driving for people all over the US.