Machine Learning Intern

Bland
San Francisco, CA
On-siteCareer-pivot friendly

Who this role is best for

Best suited to early-career researchers with strong ML foundations working in audio/speech AI, with a focus on production-ready systems and collaboration within a research team in San Francisco.

Best fit for

  • Early-career researchers with ML expertise and a clear interest in audio/speech AI
    — “Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience.
  • Candidates who can independently drive research projects and deliver impactful results
    — “We scope internships around a single meaningful question that can be answered in the time you have.
  • Individuals with experience in self-supervised or generative modeling and a deep understanding of audio quality
    — “Experience with self-supervised, generative, or multimodal modeling.

Things to consider

  • This role is based in San Francisco and likely requires in-person presence
    — “Beautiful office in Levi's Plaza, SF with rooftop views
  • The internship is full-time and may involve collaboration with engineers on production systems
    — “Where the result warrants it, work with engineers to move it toward production.

How to stand out

  • Highlight projects where you independently scoped and completed a research task from start to finish
    — “Own a research question end to end
  • Emphasize experience with real-world audio data and handling production-level challenges
    — “Train and evaluate models on large-scale, real-world telephony audio
  • Showcase your ability to design ablations and defend your methodology in presentations
    — “Design ablations that isolate what actually caused an improvement.
  • Demonstrate fluency in PyTorch and experience working in large codebases
    — “Fluent in PyTorch and comfortable in a real codebase.
  • Include any prior work or contributions in speech or language AI, even if not explicitly published
    — “Prior publications or open source contributions in speech or language AI are a strong signal
Pace · Fast PacedCollaboration · HighAutonomy · HighDecision Impact · Individual

Derived from job-description analysis by Serendipath's career intelligence engine.

What success looks like

  • Own a research question end to end
  • Work on real systems
  • Present findings to the research team
Typical background
MS or PhD in ML, CS, EEExperience with self-supervised, generative, or multimodal modeling

Skills & requirements

Required

Machine LearningSpeech-to-textLarge Language ModelsNeural Audio CodecsText-to-speech

Preferred

Audio Quality AssessmentReal-time Inference

Stack & domain

PyTorchAudioSpeech

About the role

Original posting from Bland via Ashby

THE ROLE: MACHINE LEARNING RESEARCH INTERN, AUDIO

As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or text-to-speech. You will work alongside our research team on the same problems they are working on, not on a side track built to keep interns busy.

We scope internships around a single meaningful question that can be answered in the time you have. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling millions of calls.

WHAT YOU WILL DO

Own a research question end to end

  • Take one well-scoped problem from literature review through implementation, experimentation, and results.
  • Design ablations that isolate what actually caused an improvement.
  • Present your findings to the research team and defend the methodology.

Work on real systems

  • Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard.
  • Use our distributed GPU infrastructure rather than toy-scale setups.
  • Where the result warrants it, work with engineers to move it toward production.

Choose your depth

Depending on your background and interests, your project may focus on:

  • Expressive and controllable text-to-speech, including prosody and emotion modeling
  • Neural audio codecs and discrete or continuous speech representations
  • ASR robustness for telephony, accents, and code switching
  • Real-time and streaming inference under latency constraints
  • Full-duplex conversation and turn-taking dynamics

WHAT MAKES YOU A GREAT FIT

Research foundations

  • Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience.
  • Comfortable reading a paper and reimplementing it without hand-holding.
  • Experience with self-supervised, generative, or multimodal modeling.

Audio or speech grounding

  • Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning.
  • Strong intuition for audio quality and what makes synthetic speech sound wrong.
  • Prior publications or open source contributions in speech or language AI are a strong signal, though not required.

Engineering ability

  • Fluent in PyTorch and comfortable in a real codebase.
  • Able to run your own experiments on GPU clusters without waiting to be unblocked.

HOW YOU SHOW UP

  • You identify the single experiment that validates an idea in days, not months.
  • You measure everything and let data drive decisions.
  • You are honest about negative results, because they are how we narrow the search.
  • You are obsessed with making voice agents sound truly human.
  • You use AI tools aggressively to amplify your own impact.

BENEFITS

  • Competitive intern compensation
  • Mentorship from researchers working on frontier voice AI
  • Every tool you need to succeed
  • Beautiful office in Levi's Plaza, SF with rooftop views
  • A real shot at a return offer

Source: Bland careers (Ashby)

Similar roles