Open position
Staff Software Engineer (Data Collection Lead)
We are seeking a Staff Software Engineer to build the data function for AMP. The learning workstreams on the roadmap, from perception to manipulation policies to fleet learning, depend on data from real deployments that does not exist yet in usable form.
You will seed the Data team and own the pipeline end to end: what gets captured on the robot, how it moves off the robot, how it is stored, labeled, and versioned, and how it reaches training.
You will work with the AI Lead on requirements, the on-robot team on capture, and the platform team on the data path, and hire the team behind it.
What you’ll do
- Own the AMP data strategy and build the pipeline from robot capture through to training-ready datasets.
- Define what gets captured on the robot, at what rate and fidelity, within the available compute, storage, and bandwidth.
- Build the ingest path from edge to cloud, including buffering, prioritization, and handling intermittent connectivity.
- Build dataset storage, indexing, versioning, and lineage so training runs are reproducible.
- Own annotation: tooling, quality control, and vendor management where labeling is outsourced.
- Define the automated curation and mining approach that finds the useful data instead of storing everything.
- Establish data governance for customer data: consent, retention, isolation, and access control.
- Build and mentor the Data team.
Job requirements
Experience
- 8 to 12 or more years of professional engineering experience, with 4 or more years building data infrastructure for machine learning.
- Demonstrated ownership of a data pipeline that fed production models, including the annotation and quality side.
- Experience with data captured from physical systems or sensors rather than only web or transactional sources.
- Experience leading a team or workstream, including hiring.
Education
- Bachelor's degree or higher in Computer Science, Robotics, Electrical Engineering, or a closely related technical field. Advanced degree preferred.
- Equivalent advanced industry experience with demonstrated data infrastructure leadership may be considered.
Core technical background
- Strong data engineering skills: pipeline orchestration, columnar and object storage, and large-scale processing.
- Experience with multimodal sensor data at scale: images, point clouds, trajectories, and time series.
- Experience with dataset versioning and experiment reproducibility.
- Experience with annotation tooling and measurable label quality control.
- Proficiency in Python, plus enough C++ to work on the capture side.
- Working knowledge of cloud data infrastructure, with Azure preferred.
Systems and architectural mindset
- Ability to work backward from what a model needs to what the robot should record.
- Comfortable with the storage and bandwidth trade-offs of high-rate sensor capture.
- Puts metrics on data quality.
Technical leadership and collaboration
- Able to set data practices across teams that do not report to you.
- Communicates clearly with AI, robotics, platform, and legal stakeholders.
- Able to work effectively with remote teams across multiple time zones.
Nice to have
- Experience with robot learning datasets or teleoperation data collection.
- Experience with fleet data collection from deployed devices.
- Experience with data privacy and residency requirements in industrial or EU contexts.
- Experience with active learning or automated data curation.
- Prior open-source contributions to data or ML tooling.