Nile Bits

Senior Applied ML Engineer (Speech Audio)

Nile Bits القاهرة, C, EG
Full-time Posted 2 months ago

Responsibilities

  • check_circle Benchmark and evaluate TTS and ASR models using Arabic-specific test sets, measuring metrics such as Word Error Rate (WER), naturalness, and dialect coverage.
  • check_circle Fine-tune generative models for voice cloning, zero-shot speaker adaptation, and speech synthesis.
  • check_circle Build and maintain Arabic-focused data pipelines, including: Audio collection and preprocessing Diacritization (Tashkil) Data cleaning and augmentation
  • check_circle Audio collection and preprocessing
  • check_circle Diacritization (Tashkil)
  • check_circle Data cleaning and augmentation
  • check_circle Optimize model inference for production environments using: Quantization KV-cache tuning Streaming inference techniques
  • check_circle Quantization
  • check_circle KV-cache tuning
  • check_circle Streaming inference techniques
  • check_circle Integrate and evaluate complete speech-to-speech conversational pipelines.
  • check_circle Conduct experiments based on recent research papers and convert findings into production-ready solutions.
  • check_circle Collaborate with engineering and product teams to deploy robust and scalable speech systems.

Basic qualifications

  • 5+ years of experience in Machine Learning, Applied AI, or AI Research.
  • Strong programming skills in Python.
  • Extensive hands-on experience with PyTorch and the Hugging Face ecosystem.
  • Proven experience training and fine-tuning neural models for: Text-to-Speech (TTS) Automatic Speech Recognition (ASR) Audio codecs
  • Text-to-Speech (TTS)
  • Automatic Speech Recognition (ASR)
  • Audio codecs
  • Deep understanding of modern speech architectures such as: Whisper Conformer HiFi-GAN Diffusion-based models
  • Whisper
  • Conformer
  • HiFi-GAN
  • Diffusion-based models
  • Experience with audio processing techniques including: Voice Activity Detection (VAD) Speaker Diarization Neural Vocoders
  • Voice Activity Detection (VAD)
  • Speaker Diarization
  • Neural Vocoders
  • Demonstrated ability to implement and adapt research papers into practical production experiments.
  • Strong understanding of Arabic language challenges, including: Diacritization (Tashkil) Dialectal variations Code-switching
  • Diacritization (Tashkil)
  • Dialectal variations
  • Code-switching
  • Experience with inference optimization techniques such as: Quantization Streaming inference NVIDIA TensorRT
  • Quantization
  • Streaming inference
  • NVIDIA TensorRT

Preferred qualifications

  • Experience developing custom NVIDIA CUDA kernels for high-performance model inference.
  • Familiarity with speculative decoding and other advanced acceleration techniques.
  • Experience deploying models at scale in cloud or GPU-based production environments.
  • Contributions to open-source speech or machine learning projects.

Benefits

  • check_circle All employees benefits for free (our famous games room, daily breakfast, fruits, coffee and other hot drinks, soft drinks and juices, company days out and parties…).
  • check_circle Flexible and comfortable schedule.
  • check_circle Social insurance.
  • check_circle Paid annual and national vacation.
  • check_circle Working remotely.
  • check_circle Competitive salaries.
  • check_circle Monetary rewards and incentives.
  • check_circle Career possibilities with growing team.
  • check_circle Open-door management policy.
  • check_circle Full Medical insurance.
  • check_circle Accommodation and transportation allowance.
  • check_circle Friendly environment that values innovation and efficiency.
  • check_circle Exciting opportunities for career growth and talent development.
  • check_circle Feedback encouragement.
  • check_circle Recognition and reward programs.
  • check_circle Friendly environment.
  • check_circle Fun committees.
  • check_circle Fun, smart and creative people.
  • check_circle Social benefits.
  • check_circle Natural Text-to-Speech (TTS)
  • check_circle Real-Time Automatic Speech Recognition (ASR)
  • check_circle End-to-End Speech-to-Speech Conversational Systems

Tags & Focus Areas

Fulltime Remote Machine Learning Ai

About Nile Bits