Tesla is building robust embodied intelligence through autonomous driving and humanoid robots. Core to reaching this goal is developing intelligent agents that see, move, and speak in the real world. In this role, you’ll architect and deploy models that allow our agents to listen, understand, and respond in real time powered by our custom AI chip on our embodied products. Your work will be central to making Optimus or our cars as seamless as talking to a friend. Most importantly, you will see your work repeatedly shipped to and utilized by thousands of Humanoid Robots in real world applications and millions of cars.
Define long-term architectural vision and key milestones for speech-powered, expressive character systems
Design and develop low-latency, streaming-friendly neural models that integrate audio, language, and non-verbal cues
Implement multimodal models for robust speech-to-speech, emotion recognition, and expressive response generation
Drive the character-building lifecycle from research prototypes to polished, production-grade experience
Test your models E2E on robotic platforms
Strong software engineering practices and is very comfortable with Python and Numpy programming, debugging/profiling, and version control
Experience with speech-related machine learning tasks: ASR, emotion detection, speaker diarization, or multimodal input processing
Experience training and fine-tuning large-scale speech models, LLMs, or VLMs
Familiarity with real-time audio processing and latency-constrained systems
