prometheus
MTS - Pre-Training Data & Acquisition Engineer
San Francisco, California · Hybrid · Posted today
Opens jobs.ashbyhq.com
Get a version of your resume written for this job.
- Salary
- Not listed
- Job type
- Full-time
- Work mode
- Hybrid
- Source
- Ashby (employer's hiring system)
Skills mentioned
Python, Rust, C++, Prometheus, Distributed Systems
About the role
The Role
We’re hiring a Pre-Training Data & Acquisition Engineer to build the data systems powering Prometheus’s foundation models for the physical world. You’ll work closely with research and infrastructure teams to acquire, process, and deliver large-scale training datasets across engineering, scientific, and multimodal domains. This role spans distributed crawling, source integration, data processing, and production operation, with end-to-end ownership from raw content to training-ready datasets.
What You’ll Do
Identify and integrate valuable data sources across engineering and scientific domains.
Build distributed crawlers, API integrations, and bulk ingestion systems with effective scheduling, rate limiting, retries, and incremental updates.
Develop pipelines for parsing, extraction, normalization, deduplication, quality filtering, and tokenization across heterogeneous formats.
Optimize throughput and cost across networking, compute, storage, and databases as acquisition and processing workloads scale.
Build monitoring and tooling to track source coverage, ingestion failures, processing throughput, and usable data yield.
Work closely with pre-training researchers to translate data requirements into reliable pipelines and deliver datasets ready for large-scale training.
Own dataset reproducibility, versioning, provenance, and recovery from acquisition through delivery.
What We’re Looking For
Experience building and operating large-scale distributed systems, web crawlers, or data processing pipelines.
Strong programming ability in Python and Rust, Go, C++, or a comparable systems language.
A practical understanding of web infrastructure, including HTTP, DNS, concurrency, caching, and common failure modes.
Hands-on experience with databases, object storage, and distributed batch or streaming processing.
Experience designing fault-tolerant systems that handle partial failures, resume interrupted work, and prevent unintended duplication or data loss.
Ability to profile and debug performance across CPU, memory, disk, and network usage.
Strong technical judgment when integrating unfamiliar sources, APIs, and file formats.
Bias toward fast iteration and end-to-end ownership, from initial implementation through reliable production operation.
Experience with search indexing, document extraction, or foundation-model data pipelines is a plus.
Why Join Us
Work with world-class researchers on frontier AI systems for the physical world.
Build the acquisition and processing systems that supply engineering, scientific, and multimodal data to large-scale model training.
Competitive compensation and flexible work arrangements.
High-impact, mission-driven environment.
Job ID ab-prometheus-4ab810ab-df74-4bae-94e9-8186a47012c7 · Original posting ↗
Similar jobs
Software Engineering PMTSNewSalesforceSan Francisco, California · On-site
Software Engineering MTS- Platform Security2dSalesforceBellevue, Washington · On-site
- MTS - Engineering (Security)2dCollinear AiSan Francisco, California · Hybrid
Software Engineering MTS3dSalesforceBellevue, Washington · On-site
Data Engineer (SMTS / LMTS) - MDM3dSalesforceSan Francisco, California · On-site