HEB
Senior Manager-MLOps (Austin)
Austin, Texas · On-site · Posted today
Opens careers.heb.com
Get a version of your resume written for this job.
- Salary
- Not listed
- Job type
- Full-time
- Work mode
- On-site
- Source
- iCIMS (employer's hiring system)
Skills mentioned
MLOps, GCP, CI/CD, BigQuery, Python, SQL, Databricks, Spark
About the role
Responsibilities We are looking for an execution-driven Senior Engineering Manager – MLOps Platform to build and lead the engineering team powering H-E-B’s enterprise Machine Learning platform on Google Cloud Platform (GCP). You'll work closely with stakeholders from product and design, and other engineering leaders, to provide high-quality, repeatable technology delivery for the digital engineering organization. Responsible for managing a team(s) that may include multiple related departments. Partner with senior leaders to define engineering strategy, roadmap priorities, operational standards, and organizational objectives aligned with business goals. Responsible for resource allocation, prioritization, and financial responsibility. Responsible for hiring, firing, and performance / pay reviews. Key Responsibilities & Essential Functions: In this role, you will treat MLOps as a product, providing an end-to-end, self-service platform that empowers Data Scientists and ML Engineers to move models seamlessly from experimentation to production. You will bridge modern data platform foundations with scalable ML infrastructure, driving the vision for model training, CI/CD automation, scalable inference, LLMOps, and continuous monitoring. People & Technical Leadership: Build, mentor, and lead a high-performing team of MLOps and platform engineers, including hiring, coaching, performance management, career development, and organizational planning. End-to-End MLOps Platform: Architect and deliver a cohesive ML lifecycle platform— spanning feature stores, model registry, experimentation tracking, CI/CD for ML, and automated pipelines. Establish engineering standards, provide technical guidance on complex challenges, and drive improvements to processes, tools, and platform capabilities. Developer Experience for AI/ML: Partner with Data Scientists and ML Engineers to improve developer experience through self-service tooling, SDKs, templates, and automated workflows. Scalable Serving & LLMOps: Establish high-throughput batch and real-time inference infrastructure on GCP, integrating LLMOps capabilities such as RAG pipelines, fine-tuning infrastructure, and vector search. Data Platform Integration & Governance: Leverage enterprise data platforms (BigQuery, data lakes) to ensure efficient data ingestion for training/serving, robust RBAC, model governance, drift detection, and cost optimization. The responsibilities and essential functions outlined above describe the general nature and level of work assigned to this position. This is not an exhaustive list of all duties, responsibilities, and skills required. Duties and responsibilities may be modified at any time based on business needs. Employees may be required to perform other job-related tasks as requested by their supervisor, subject to reasonable accommodations. Qualifications & Key Requirements: Work Experience: 7+ years of experience in software, data, or platform engineering 2+ years of experience managing engineering teams delivering production platform infrastructure. Experience successfully delivering timely, high-quality software. Knowledge/Skills/Abilities: MLOps & AI Tooling: Hands-on experience with production ML frameworks and orchestrators (Vertex AI, Kubeflow, MLflow, Ray, Feast/Feature Stores, Triton, or vLLM). GCP & Cloud Infrastructure: Strong expertise in GCP (Vertex AI, BigQuery, GKE, Cloud Composer/Airflow, Pub/Sub, Cloud Storage) and Infrastructure as Code (Terraform). Engineering Practices: Solid background in container orchestration (Docker, Kubernetes), modern CI/CD automation, and distributed systems in Python. Data Platform Acumen: Practical understanding of distributed data processing (Spark, SQL) and how feature pipelines integrate into core data warehouses and data lakes. Preferred Qualifications Experience migrating legacy ML workloads or scaling GenAI/LLM pipelines in an enterprise retail, e-commerce, or supply chain environment. Familiarity with hybrid/multi-cloud data ecosystems (AWS, Databricks). Track record of building a strong "Platform-as-a-Product" culture with high internal customer adoption. Education: Bachelor's / Master's degree in a relevant field of education and / or work experience leading successful projects and coaching / mentoring others in functional area. Physical Demands & Working Conditions: Function in a fast-paced multi-priority environment Travel by car or plane with overnight stays Regularly lift up to 20 lbs Work extended hours and / or rotating schedules The work environment characteristics described here are representative of those a Partner encounters while performing the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. Last revised: 11/01/2024 JDENGINEERING
We are looking for an execution-driven Senior Engineering Manager – MLOps Platform to build and lead the engineering team powering H-E-B’s enterprise Machine Learning platform on Google Cloud Platform (GCP). You'll work closely with stakeholders from product and design, and other engineering leaders, to provide high-quality, repeatable technology delivery for the digital engineering organization. Responsible for managing a team(s) that may include multiple related departments. Partner with senior leaders to define engineering strategy, roadmap priorities, operational standards, and organizational objectives aligned with business goals. Responsible for resource allocation, prioritization, and financial responsibility. Responsible for hiring, firing, and performance / pay reviews. Key Responsibilities & Essential Functions: In this role, you will treat MLOps as a product, providing an end-to-end, self-service platform that empowers Data Scientists and ML Engineers to move models seamlessly from experimentation to production. You will bridge modern data platform foundations with scalable ML infrastructure, driving the vision for model training, CI/CD automation, scalable inference, LLMOps, and continuous monitoring. People & Technical Leadership: Build, mentor, and lead a high-performing team of MLOps and platform engineers, including hiring, coaching, performance management, career development, and organizational planning. End-to-End MLOps Platform: Architect and deliver a cohesive ML lifecycle platform— spanning feature stores, model registry, experimentation tracking, CI/CD for ML, and automated pipelines. Establish engineering standards, provide technical guidance on complex challenges, and drive improvements to processes, tools, and platform capabilities. Developer Experience for AI/ML: Partner with Data Scientists and ML Engineers to improve developer experience through self-service tooling, SDKs, templates, and automated workflows. Scalable Serving & LLMOps: Establish high-throughput batch and real-time inference infrastructure on GCP, integrating LLMOps capabilities such as RAG pipelines, fine-tuning infrastructure, and vector search. Data Platform Integration & Governance: Leverage enterprise data platforms (BigQuery, data lakes) to ensure efficient data ingestion for training/serving, robust RBAC, model governance, drift detection, and cost optimization. The responsibilities and essential functions outlined above describe the general nature and level of work assigned to this position. This is not an exhaustive list of all duties, responsibilities, and skills required. Duties and responsibilities may be modified at any time based on business needs. Employees may be required to perform other job-related tasks as requested by their supervisor, subject to reasonable accommodations. Qualifications & Key Requirements: Work Experience: 7+ years of experience in software, data, or platform engineering 2+ years of experience managing engineering teams delivering production platform infrastructure. Experience successfully delivering timely, high-quality software. Knowledge/Skills/Abilities: MLOps & AI Tooling: Hands-on experience with production ML frameworks and orchestrators (Vertex AI, Kubeflow, MLflow, Ray, Feast/Feature Stores, Triton, or vLLM). GCP & Cloud Infrastructure: Strong expertise in GCP (Vertex AI, BigQuery, GKE, Cloud Composer/Airflow, Pub/Sub, Cloud Storage) and Infrastructure as Code (Terraform). Engineering Practices: Solid background in container orchestration (Docker, Kubernetes), modern CI/CD automation, and distributed systems in Python. Data Platform Acumen: Practical understanding of distributed data processing (Spark, SQL) and how feature pipelines integrate into core data warehouses and data lakes. Preferred Qualifications Experience migrating legacy ML workloads or scaling GenAI/LLM pipelines in an enterprise retail, e-commerce, or supply chain environment. Familiarity with hybrid/multi-cloud data ecosystems (AWS, Databricks). Track record of building a strong "Platform-as-a-Product" culture with high internal customer adoption. Education: Bachelor's / Master's degree in a relevant field of education and / or work experience leading successful projects and coaching / mentoring others in functional area. Physical Demands & Working Conditions: Function in a fast-paced multi-priority environment Travel by car or plane with overnight stays Regularly lift up to 20 lbs Work extended hours and / or rotating schedules The work environment characteristics described here are representative of those a Partner encounters while performing the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. Last revised: 11/01/2024 JDENGINEERING
Job ID ic-careers-heb-com-240590 · Original posting ↗
Similar jobs
- Engineering ManagerNewarlohotelsheadofficeWashington, District of Columbia · On-site
- Engineering ManagerNewarlodcWashington, District of Columbia · On-site
- Technical Project Manager (AI)New3EBethesda, Maryland · Hybrid · US$90,000–105,000 / year
IT Project ManagerNewEnerfabBlue Ash, Ohio
Staff Product Manager, Growth Automation & EnablementNewHeadwayRemote, United States · Remote · US$265,000–331,500 / year