Senior Machine Learning Engineer, Speech & LLM Training Data
Own Your Impact. At Propio, we don't believe careers happen to people. We believe people create them.
Here, you're trusted to make decisions, challenge assumptions, drive innovation, and shape outcomes. Your success is not limited by hierarchy or tenure. It's fueled by your ambition, your curiosity, and your willingness to own your impact. If you're looking for a role where you can simply maintain the status quo, this probably isn't it, but if you're looking for a place where your ideas matter, your growth is accelerated, and your work creates meaningful impact across the world, we'd love to talk.
Why Propio?
Every day, communication changes lives. A patient receives care they otherwise couldn't access. A family gains critical information. A business connects with a customer. A community becomes more inclusive. These moments happen because barriers are removed. And behind those moments are Propio team members who show up every day to solve problems, innovate, and build the future. This isn't just work. This is world impact.
As a Senior Machine Learning Engineer, Speech & LLM Training Data, you'll have the opportunity to make a meaningful contribution to the continued growth and transformation of Propio. You will transform large volumes of multilingual conversational audio into high-quality training and evaluation datasets. This hands-on role owns audio processing, dataset curation, annotation and QA workflows, model training, and evaluation for our multilingual speech, translation, and conversational AI systems.
You'll Be Empowered To
- Take ownership of important initiatives and outcomes.
- Drive meaningful business results.
- Influence decisions and contribute new ideas.
- Partner with talented, high-performing team members.
- Challenge yourself through continuous learning and growth.
- Help shape the future of a rapidly growing organization.
What You'll Own
- Define the data roadmap for multilingual speech, translation, multimodal LLMs, and conversational AI.
- Build audio-processing pipelines covering resampling, channel handling, VAD, diarization, language identification, transcription, alignment, and quality filtering.
- Build dataset pipelines for cleaning, deduplication, PII/PHI redaction, quality scoring, sampling, balancing, versioning, and lineage.
- Design annotation guidelines, QA rubrics, golden datasets, and reviewer workflows.
- Build evaluation datasets, analyze model failures, and translate performance gaps into targeted data improvements.
- Run training, fine-tuning, post-training, and evaluation experiments, including SFT, preference data, DPO/RLHF-style workflows, and synthetic data generation.
- Productionize secure, traceable, and reproducible data and ML workflows on AWS.
What Makes Someone Successful Here
The most successful people at Propio aren't necessarily the ones with the longest resumes. They're the people who:
- Take ownership instead of waiting for direction.
- Embrace challenges as opportunities to grow.
- Continuously seek better ways of working.
- Turn ideas into action.
- Hold themselves and others accountable to high standards.
- Are driven by making a measurable impact.
What You'll Bring
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science, Machine Learning, Data Science, Electrical Engineering, Computational Linguistics, or a related field, or equivalent practical experience.
- 5+ years of experience in ML engineering, speech/audio ML, ML data engineering, NLP, or LLM training-data workflows.
- Strong hands-on experience with Python, SQL, Linux, Git, and Docker.
- Experience training or evaluating models using PyTorch, Hugging Face, or comparable ML frameworks.
- Experience with FFmpeg and audio-processing libraries such as TorchCodec, torchaudio, librosa, or equivalent tools.
- Experience with speech-processing tasks such as VAD, diarization, ASR, forced alignment, language identification, and audio-quality analysis.
- Experience with Databricks/Spark, Parquet/Arrow, and large-scale dataset pipelines.
- Working knowledge of AWS S3, SageMaker, Glue, Step Functions, IAM, and KMS.
- Experience with an annotation platform such as Labelbox, Label Studio, Scale AI, Prodigy, Argilla, or custom internal tooling.
- Experience with experiment tracking and data versioning tools such as MLflow, Weights & Biases, DVC, Delta Lake, or LakeFS.
- Experience with multilingual speech, translation, annotation workflows, and evaluation datasets.
Preferred Qualifications:
- Experience with multilingual telephony, healthcare, interpretation, or call-center audio.
- Experience with tools such as Silero VAD, pyannote, WhisperX, NeMo, Kaldi, or equivalent speech technologies.
- Experience with distributed processing or training using Ray, PySpark, or similar frameworks.
- Experience with HIPAA, PHI/PII redaction, and secure data governance.
- Experience with low-resource languages, accents, dialects, and code-switching.
- Experience with synthetic data, active learning, weak supervision, or LLM-as-judge evaluation.
Even if your experience doesn't perfectly match every qualification, we encourage you to apply. We're looking for potential, drive, and a commitment to growth as much as experience.
What You'll Gain
Own Your Growth: We invest in people who invest in themselves. You'll have opportunities to learn, develop, and expand your capabilities while building a meaningful career.
Own Your Impact: You'll see the connection between your work and our success. We believe great people deserve the opportunity to make a real difference.
Own Your Innovation: The best ideas can come from anywhere. We encourage curiosity, creativity, and challenging the way things have always been done.
Own Your Success: Whether you're building expertise, pursuing leadership opportunities, or expanding your career path, we'll give you room to grow and the support to get there.
At Propio, your work doesn't just move a company forward. It helps connect people, communities, and opportunities across the world. Ready to Build Something Bigger? Apply today and discover what happens when you own your success.
#LI-JS1
Notice of AI Use in Job Application Review
As part of our commitment in creating a fair, efficient, and consistent hiring process we may use artificial intelligence (AI) to help our recruiting teams organize, summarize, and analyze information provided by candidates, including resumes, application responses, and other materials submitted during the application process.AI may be used to identify patterns, highlight relevant skills, and experience, and assist in comparing a candidate’s qualifications with the requirement of a specific role. These tools are to improve efficiency and consistency while supporting more informed hiring decisions, which will ultimately be made by the hiring team.