Innodata Inc.

Research Scientist, Video & Multimodal

Remote, United States · Remote · Posted today

Opens job-boards.greenhouse.io

Get a version of your resume written for this job.

Salary
Not listed
Job type
Not specified
Work mode
Remote
Source
Greenhouse (employer's hiring system)

Skills mentioned

PyTorch, Generative AI, Data Engineering

About the role

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role: 

Video is where multimodal models are weakest and hardest to grade. Temporal reasoning, long-form understanding, grounding events in time, and holding audio, video, and text together do not fall out of image benchmarks — and the evaluations for them are still immature. Closing that gap is gated as much by how we design data and evaluation as by architecture. Innodata builds that data and those evaluations for the customers and frontier labs advancing video and multimodal models, and we are hiring a Research Scientist to own the science behind it. 

You will partner directly with the customers and frontier labs building video understanding, video-language, and video-generation models, as interested in the data behind them as in the models themselves. Video spans two model families judged in completely different ways: models that understand video — answering questions, localizing events, grounding language in time — where the question is whether the answer is correct; and models that generate it, where fidelity, temporal coherence, and physical plausibility matter and no automatic metric is settled. You own the evaluation science for both, and knowing when model-based scoring can stand in for a human versus when it can't. Your conclusions shape what our partners measure and collect next. 

What You’ll Own:

You will define how Innodata designs, structures, and evaluates video data for video and multimodal models, and you will validate those choices experimentally. Concretely, you will: 

  • Translate the requirements of video and multimodal models — video understanding, temporal and event localization, action recognition, long-form video, video-language models, video generation, cross-modal reasoning, and multimodal retrieval and grounding — into concrete data specifications: modalities, annotation schemas, sampling, and evaluation criteria. 
  • Build evaluation methodology for video understanding — temporal grounding accuracy, long-context and long-horizon reasoning, and dynamic multi-turn, cross-modal, and retrieval-and-grounding evaluation — clear about when model-based scoring is trustworthy and when a human is needed. 
  • Build evaluation methodology for video generation — fidelity, temporal coherence, and physical plausibility, including generative video used as a world model — the regime where automatic metrics are weakest and human judgment matters most. 
  • Decide how existing and incoming video should be structured, enriched, and sampled to extract the most model value from it, including from messy, domain-specific footage. 
  • Run experiments that prove data decisions matter: fine-tune and evaluate models on Innodata data, with ablations tying specific data choices to measurable improvement. 
  • Design adversarial and stumping evaluations that surface where video and multimodal systems fail, and turn those failures into better data. 
  • Publish. Turn what you learn into benchmarks, methodology, and papers that advance the field and earn the trust of the customers and frontier labs we partner with. 
  • Work with annotation teams, subject-matter experts, and the synthetic-data pipeline to turn specifications into operational collection and labeling plans. 

 You’ll Thrive in This Role If You Have:

  • Roughly 5+ years of hands-on industry experience in video understanding or multimodal ML. We weight practical experience over formal credentials; a PhD with a compelling, current research agenda can offset the lower end. 
  • A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree (MS or PhD) in a relevant field is preferred. 
  • Trained and evaluated video or multimodal models yourself, with strong PyTorch fundamentals. 
  • Fluency in the formats and tooling video work runs on: ffmpeg and decord pipelines, temporal and COCO-style annotation, WebDataset, Parquet and Arrow, and HuggingFace datasets. 
  • Experience fine-tuning large video or vision-language models with the modern toolchain (HuggingFace transformers, PEFT, efficient inference), and with long-form video, streaming, temporal segmentation, or synthetic video generation. 
  • A way of thinking in datasets and benchmarks: you have built evaluation sets, calibrated difficulty, and argued about what makes video data good for a given objective. 
  • A track record the field recognizes: first-author publications or strong open-source contributions at venues such as CVPR, ICCV, ECCV, NeurIPS, or ICLR. 
  • The ability to work directly with the research scientists at the customers and frontier labs we partner with, and to explain data and modeling decisions clearly to both expert and non-expert audiences, backed by a rigorous, reproducible approach to experiments and documentation. 
  • Bonus: interest or hands-on experience in responsible-AI evaluation and red-teaming — safety and robustness testing for video and multimodal systems. 

The expected salary range for this position is $160,000 - $185,000 p/year, based on experience, skills, and qualifications.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams. 

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.

Job ID gh-innodatainc-4415074009 · Original posting ↗