AI Intern – Vision-Language-Action (VLA) & Data

Rivr
Zürich, Switzerland
On-siteCareer-pivot friendly

Who this role is best for

Geared toward candidates with a technical background in robotics and deep learning, comfortable with processing multi-modal sensor data and working in a collaborative, in-person environment in Zürich.

Best fit for

  • Candidates with a BSc in Computer Science or Robotics and experience with computer vision projects
    — “At least BSc in Computer Science, Robotics, Machine Learning, or a related field.
  • Individuals with prior exposure to VLA models or sensor fusion techniques
    — “Knowledge or experience with transformers, Vision-Language Models and/or VLA’s.
  • Those who have worked with large-scale datasets and have a strong interest in real-world robotics applications
    — “Previous experience working with large-scale image or video datasets.

Things to consider

  • In-person presence is required despite the internship nature of the role
    — “require in-person presence in our office locations.
  • Geographic eligibility is restricted to Schengen Area citizens, excluding many international candidates
    — “only accept applications from citizens of Schengen Area countries.

How to stand out

  • Highlight experience with PyTorch and specific computer vision tasks in your resume and interviews
    — “Proficiency in Python and experience with deep learning frameworks (preferably PyTorch).
  • Demonstrate familiarity with data visualization tools and debugging model performance
    — “Assist in developing software tools to visualize data and debug model performance.
  • Emphasize any prior work with robotic systems or motion/action prediction in your application materials
    — “Experience working with robotic systems.
Pace · Fast PacedCollaboration · HighAutonomy · MediumDecision Impact · Individual

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • successful processing and analysis of multi-modal sensor data
  • development of robust VLA models
Typical background
computer scienceroboticsmachine learning

Skills & requirements

Required

PythonDeep LearningComputer VisionData AnalysisProblem-solvingCollaboration

Preferred

TransformersVision-language ModelsVLA3D GeometrySensor FusionRobotic SystemsMotion/action Prediction

Stack & domain

Multi-modal Sensor DataVLA ModelsSoftware ToolsData VisualizationDebug Model PerformanceData StrategiesRobotic SystemsSelf-supervised LearningGenerative AIPythonDeep Learning FrameworksComputer VisionData Manipulation LibrariesNumPyPandasOpencv3D GeometryCamera ProjectionsSensor FusionMotion/action PredictionProblem-solvingCollaborationAIRoboticsVisionLanguageAction

About the role

Original posting from Rivr via Lever

RIVR, part of Amazon is a robotics company pioneering Physical AI through real-world doorstep delivery. Founded in 2024 as an ETH Zurich spin-off, RIVR, part of Amazon developed wheeled-legged robots designed to operate in complex, unstructured environments such as stairs, gates, doors, and uneven urban terrain. We believe that achieving general physical intelligence requires solving real customer problems in the real world, where robots can learn from rich operational data at scale.

Following our acquisition by Amazon in March 2026, we are continuing this mission with greater reach and speed. By combining custom robot hardware, onboard autonomy, and cloud-based coordination, RIVR, part of Amazon is building the next generation of safe, reliable autonomous robots for last-mile delivery.

Important Notice: For this position, we can unfortunately only accept applications from citizens of Schengen Area countries. This restriction does not apply to ETHZ and EPFL students who are required to complete compulsory internships as part of their studies.

What you’ll be doing:

-

Support the team in processing, curating, and analyzing multi-modal sensor data for training VLA models.

-

Get hands-on experience with state-of-the-art VLA models.

-

Assist in developing software tools to visualize data and debug model performance.

-

Work closely with senior engineers to implement data strategies that improve the robustness of our robotic systems.

-

Help integrate software components to evaluate algorithms in simulation and on hardware.

-

Engage in continuous learning and gain exposure to recent developments in VLA, self-supervised learning, and generative AI.

What you must have:

-

At least BSc in Computer Science, Robotics, Machine Learning, or a related field.

-

Proficiency in Python and experience with deep learning frameworks (preferably PyTorch).

-

Hands-on experience (either through coursework, previous internships, or other projects) with deep learning for computer vision.

-

Familiarity with data manipulation libraries (e.g., NumPy, Pandas, OpenCV).

-

Strong problem-solving skills and an eagerness to work with complex, real-world data.

-

Eagerness to learn and contribute in a collaborative team environment.

Get some bonus points:

-

An MSc or ongoing PhD in a related field.

-

Previous experience working with large-scale image or video datasets.

-

Knowledge or experience with transformers, Vision-Language Models and/or VLA’s.

-

Experience with 3D geometry, camera projections, or sensor fusion.

-

Experience working with robotic systems

-

Experience in motion/action prediction in robotics context.

RIVR, part of Amazon is committed to building a diverse and inclusive team that values every perspective. If you’re passionate about driving innovation in robotics and creating meaningful impact, we encourage you to apply and bring your unique self to our team.

We believe the best work is done when collaborating and therefore require in-person presence in our office locations.

Source: Rivr careers (Lever)

Similar roles