PhillyTech.Co

Senior ML Ops Engineer | Hybrid + Equity | AI Powered Outage Intelligence SaaS Startup

King of Prussia, Pennsylvania · On-site · Posted today

Opens jobs.smartrecruiters.com

Get a version of your resume written for this job.

Salary
US$120,000–130,000 / year
Job type
Full-time
Work mode
On-site
Source
SmartRecruiters (employer's hiring system)

Skills mentioned

SaaS, Machine Learning, Python, PostgreSQL, MLOps, Data Engineering, Kubernetes

About the role

Job Description

As a Senior MLOps Engineer, you will have a high level of ownership over the platform the machine learning team builds on, spanning model training, experimentation, deployment, monitoring, and the data infrastructure supporting production ML.

You must have firsthand experience building an ML platform from the ground up or taking an existing platform through significant scaling and continuing to operate it as it matured.

Responsibilities include:

  • Own and extend the ML platform end to end, including training orchestration, experiment tracking, model registry, deployment, and production monitoring.
  • Build and operate data pipelines supporting model training and online inference.
  • Design processes for backfills, replays, and recovery when upstream data feeds fail.
  • Ensure training data accurately reflects what was known at the point in time it represents.
  • Build and manage reliable processes for moving trained models into production.
  • Ensure model training runs and results are reproducible.
  • Monitor deployed model performance over time, including models where outcomes are confirmed later.
  • Manage training and inference costs as data volume and the number of production models grow.
  • Contribute to technical architecture design and reviews.

Qualifications

  • 4+ years of professional software engineering or data engineering experience.
  • 2+ years building and operating machine learning systems in production.
  • Hands on experience building an ML platform from the ground up or significantly scaling an existing ML platform.
  • Experience continuing to own and operate ML infrastructure as it matured.
  • Experience supporting multiple production models with complex training workloads.
  • Experience building and maintaining production data pipelines at scale, including managing their operating costs.
  • Experience working with multiple external data sources that behave differently and may arrive inconsistently.
  • Working knowledge of machine learning modeling and evaluation, with enough depth to review and challenge the work of ML engineers.
  • Entrepreneurial mindset and interest in working within a fast moving startup environment.

Preferred Experience

  • Advanced proficiency with Python 3 in mature production environments + Kubernetes + PostgreSQL.
  • Experiment tracking, model registry, and model serving frameworks.
  • Workflow orchestration.
  • Infrastructure as code tooling.
  • Geospatial or time series systems.
  • Large real time data flows.
  • Experience working for a data intelligence company.
  • Bachelor's degree in Computer Science or a related field.

Additional Information

About SaaS Talent

SaaS Talent is more than just a recruiting company. We're your hiring, business development and growth partner with 20+ years of experience in SaaS and Hi-Tech that helps you scale and transform your business. We've worked with 100+ companies and helped them achieve their goals. From streamlining sales, marketing, and operations to hiring ideal talent and getting funding, if you're struggling to grow, we're an ideal choice.

Reach out to us at www.saas-talent.com to learn more about how we can help you.

SMS Communication Consent Disclaimer

By applying for this position, you agree to receive text message updates from SaaS Talent related to job opportunities. Standard message and data rates may apply, and messaging frequency varies. Text HELP for help and STOP to cancel. Learn more about our opt-in SMS Communication consent policy here: https://www.saas-talent.com/opt-in-sms-communication-consent

Company Description

This is a 3-day in-office hybrid role in King of Prussia, PA

ARE YOU READY TO BUILD AND SCALE THE MACHINE LEARNING INFRASTRUCTURE POWERING REAL TIME INTELLIGENCE FOR SOME OF THE WORLD’S MOST CRITICAL INFRASTRUCTURE?

Our client is a fast growing B2B SaaS outage intelligence company helping major enterprises reduce downtime through advanced automation and real time intelligence. Their technology helps organizations understand outages faster, automate operational workflows, reduce unnecessary costs, and accelerate repair times.

Their platform supports critical infrastructure operations where reliability matters. As the company expands its customer base and develops new products, they are investing further in the machine learning systems behind their intelligence platform.

This is not a role where you inherit a finished ML platform and simply maintain it. They are looking for someone who has previously helped build an ML stack from the ground up, understands what production ML infrastructure looks like as it scales, and wants meaningful ownership over the systems supporting machine learning in production.

Why Join

This is an opportunity to join a growing technology company where you will have significant ownership of the ML platform and work directly with the engineers building the models that power the product.

  • Own and influence the ML platform from model training through production deployment and monitoring.
  • Build systems that customers rely on in real time.
  • Work directly with ML engineers on model development, evaluation, and productionization.
  • Help shape technical architecture and the future of the company's ML and operational infrastructure.
  • Solve complex engineering problems involving machine learning, large scale data pipelines, external data sources, and real time systems.
  • Join an entrepreneurial environment where engineers are expected to take ownership and influence how things are built.
  • Opportunity to earn equity in a growing technology company.

Benefits

  • Hybrid work model, onsite in King of Prussia 3 days per week
  • Equity in a fast-scaling SaaS company
  • Fully paid medical, dental, and vision options
  • Life and AD&D insurance
  • PTO

Job ID sr-phillytechco-744000152245344 · Original posting ↗