AI Research Scientist

Embedding Vc
Milpitas, CA
On-site

Who this role is best for

Best suited to candidates with expertise in vision-language-action models and multi-modal data integration working in robotics for industrial environments, requiring in-office collaboration in Milpitas, CA.

Best fit for

  • Candidates with advanced degrees in ML/Robotics and experience in world model architectures
    — “PhD degree in Machine Learning, Robotics, or related field, or Master's degree with 4+ years of relevant experience
  • Individuals skilled in action-conditioned prediction and long-horizon planning for embodied AI
    — “Develop and train world models for action-conditioned prediction, long-horizon planning, and environment simulation
  • Professionals with a track record of integrating vision, language, and spatial data for robotics
    — “Integrate multi-modal data sources (vision, language, speech, etc.) to enable natural human-robot communication

Things to consider

  • Full-time in-office presence is required, with no remote flexibility mentioned
    — “Requires 5 days/week in-office collaboration with the teams
  • Deployment expertise is critical, not just research capabilities
    — “Optimize and deploy models as production-grade solutions on RoboForce robotic platforms

How to stand out

  • Highlight production deployment experience with TensorRT or CUDA in your technical projects
    — “Expertise in neural network deployment (e.g., TensorRT) and GPU programming with CUDA
  • Quantify contributions to multimodal model development in your resume's achievements section
    — “Integrate multi-modal data sources (vision, language, speech, etc.)
  • Emphasize research in spatial reasoning and environment simulation for foundation models
    — “Develop foundation models with spatial reasoning capabilities
  • Showcase experience with synthetic data generation for robot learning if applicable
    — “Experience with video generation or prediction models... synthetic data generation for robot learning
Pace · Fast PacedCollaboration · HighAutonomy · HighDecision Impact · CompanyLevel · Senior

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • design and deploy vision-language-action models
  • develop and train world models
  • research approaches to improve world model fidelity
  • integrate multi-modal data sources
  • optimize and deploy models as production-grade solutions
Typical background
AI researchroboticsmachine learning

Skills & requirements

Required

Vision-language-action ModelsWorld ModelsAction-conditioned PredictionLong-horizon PlanningEnvironment SimulationMultimodal ModelsModern ML ArchitecturesNeural Network DeploymentGPU Programming

Preferred

Video GenerationPrediction ModelsDiffusion-based Video ModelsAutoregressive Video TransformersNeural Network DeploymentGPU Programming

Stack & domain

Machine LearningRoboticsPythonDeep Learning FrameworksLarge Foundation ModelsWorld Model ArchitecturesAction-conditioned Generative ModelingMultimodal ModelsModern ML ArchitecturesNeural Network DeploymentGPU ProgrammingLeadershipCommunicationProblem-solvingAdaptabilityOwnershipAttention To DetailProactivitySystematic ApproachCritical ThinkingAI

About the role

Original posting from Embedding Vc via Ashby

Why RoboForce

RoboForce is an AI robotics company developing Physical AI–powered Robo-Labor for dull, dirty, and dangerous work. The company's robots are engineered for demanding industrial environments, with a focus on real-world deployment and scalability.

We are looking for a Senior / Staff AI Research Scientist, Foundation Models to advance robotic embodied intelligence. In this role, you will develop algorithms that enable robots to understand their environment, interpret and execute tasks, and communicate seamlessly with humans — with a particular focus on building and training world models that allow robots to predict, plan, and generalize across complex physical tasks.

Responsibilities

  • Design and deploy vision-language(-action) models (VLM/VLA) for contextual understanding and generalized robot action policies.
  • Develop and train world models for action-conditioned prediction, long-horizon planning, and environment simulation — enabling robots to reason about the consequences of their actions before execution.
  • Research approaches to improve world model fidelity using multi-modal inputs including vision, language, proprioception, and spatial representations.
  • Develop foundation models with spatial reasoning capabilities to achieve high-precision robotic actions.
  • Integrate multi-modal data sources (vision, language, speech, etc.) to enable natural human-robot communication.
  • Optimize and deploy models as production-grade solutions on RoboForce robotic platforms.

Requirements

  • PhD degree in Machine Learning, Robotics, or related field, or Master's degree with 4+ years of relevant experience.
  • Proficiency in Python and deep learning frameworks (e.g., PyTorch, JAX).
  • Expertise in large foundation models (VLM, VLA, etc.).
  • Strong understanding of world model architectures and action-conditioned generative modeling for robot learning.
  • Decent understanding of multimodal models, modern ML architectures (transformers, diffusion models, etc.).
  • Requires 5 days/week in-office collaboration with the teams.

Bonus Qualifications

  • Experience with video generation or prediction models (e.g., diffusion-based video models, autoregressive video transformers) and their application to world modeling or synthetic data generation for robot learning.
  • Strong publication record at top conferences (NeurIPS, ICML, CVPR, ICCV, CoRL, ICRA, or equivalent).
  • Expertise in neural network deployment (e.g., TensorRT) and GPU programming with CUDA.
  • Proven ability to design scalable experimentation and data pipelines.

Benefits

  • Competitive stock options/equity programs.
  • Health, dental, and vision insurance, 401(k) plan.
  • Visa sponsorship and green card support for qualified candidates.
  • Lunches and dinners, a fully stocked kitchen, and regular team-building events.

Source: Embedding Vc careers (Ashby)

Similar roles