Clera

ML Infrastructure Engineer

San Mateo, California · On-site · Posted 2 days ago

Opens jobs.ashbyhq.com

Get a version of your resume written for this job.

Salary
Not listed
Job type
Full-time
Work mode
On-site
Source
Ashby (employer's hiring system)

Skills mentioned

Python, Java, Rust, C++, TensorFlow, Machine Learning, AWS, Azure

About the role

About the Role

This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI company building a context and data governance layer that makes AI agents reliable in production. You will own the inference and model-serving infrastructure end to end, ensuring agents run fast and reliably at increasing concurrency. The work is squarely production-focused with real-world impact across regulated industries like insurance, banking, healthcare, and asset management.

What You'll Do

  • Design, build, and scale inference and model-serving infrastructure from the ground up through production deployment.

  • Optimize systems for latency, throughput, and reliability under high concurrency.

  • Collaborate closely with ML and infrastructure teams to ensure seamless integration and surface performance bottlenecks.

  • Drive solutions to infrastructure challenges across a fast-moving, cross-functional team.

What We're Looking For

  • 5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.

  • Hands-on experience designing and scaling inference-serving systems using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom solutions.

  • Strong distributed systems fundamentals, including containerization and orchestration with Docker and Kubernetes.

  • Proficiency with monitoring and observability tooling for production systems, such as Prometheus, Grafana, or distributed tracing frameworks.

  • Experience deploying and managing ML workloads on cloud platforms (AWS, GCP, or Azure).

  • Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.

  • Comfort collaborating across both ML and infrastructure disciplines in a fast-paced environment.

  • Nice to have: experience with knowledge graphs, semantic search, or graph databases; real-time or low-latency inference systems; agentic or multi-step AI pipelines; enterprise data integration or pipeline infrastructure.

Location

On-site in San Mateo, California, United States. Visa sponsorship is not available for this role.

Job ID ab-clera-88276d4c-0d4b-47d8-984c-0b315403d107 · Original posting ↗