Aloha Consulting Group

AI Engineer (Voice & Speech)

Aloha Consulting Group  •  Hanoi, VN (Onsite)  •  1 hour ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description


ACG_3756_JOB

Our client is a leading software company in Vietnam who is looking for a qualified candidate to join their firm.

JOB OVERVIEW

  • The AI Engineer will work at the intersection of AI Research, Product Engineering, and Production, focusing on the development of advanced Speech AI and Voice AI capabilities.

  • The role involves researching and developing speech technologies, building end-to-end AI pipelines, deploying models into production environments, and continuously optimizing AI systems for accuracy, scalability, latency, and real-world performance.

KEY RESPONSIBILITIES


Speech AI Research & Development


  • Explore and prototype new approaches to solve complex speech and audio AI problems.

  • Develop, fine-tune, and adapt speech models for practical real-world applications.

  • Evaluate different modeling approaches and conduct systematic experiments to improve model performance.

  • Research and reproduce state-of-the-art Speech AI methods and translate promising research findings into production-ready solutions.

  • Design rigorous experiments, benchmarks, and error analyses to understand model limitations and identify opportunities for improvement.

  • Explore alternative modeling approaches when existing solutions are insufficient.


Speech AI System Development


  • Develop and optimize intelligent voice technologies, including Automatic Speech Recognition, Text-to-Speech, audio processing, speech enhancement, speech understanding, and voice interaction systems.

  • Improve speech recognition accuracy, robustness, and performance across different environments and use cases.

  • Improve speech synthesis quality, naturalness, and reliability.

  • Design AI solutions that can operate effectively under real-world audio conditions.


End-to-End Audio AI Pipelines


  • Design, develop, and optimize production-ready audio AI pipelines that connect AI models with real-world applications.

  • Build audio preprocessing and feature extraction pipelines.

  • Develop noise suppression and speech enhancement capabilities.

  • Build and optimize real-time audio streaming and processing systems.

  • Develop model serving and inference pipelines.

  • Integrate AI models into cloud, mobile, and other deployment environments.

  • Design and implement end-to-end AI workflows from data preparation and model development through deployment and monitoring.


AI-Powered Product Development


  • Collaborate closely with Product, Mobile, Backend, and other Engineering teams to transform AI capabilities into practical product experiences.

  • Contribute to voice recording and transcription capabilities.

  • Develop AI-powered voice assistants and voice-driven workflows.

  • Build natural conversational experiences and speech-enabled product features.

  • Integrate AI capabilities into mobile applications and business workflows.


AI Model Production & Scalability


  • Take AI models from research prototypes through deployment into reliable production systems.

  • Optimize inference latency and computational resource consumption.

  • Optimize AI models for different deployment environments and hardware constraints.

  • Build scalable, stable, and reliable AI services.

  • Monitor model performance in production environments and use real-world feedback to continuously improve models and systems.

  • Identify performance degradation, model limitations, and opportunities for optimization after deployment.


Data & Experimentation


  • Build and improve high-quality datasets for model training and evaluation.

  • Contribute to data collection, synthetic data generation, data augmentation, and quality validation processes.

  • Develop appropriate evaluation datasets and benchmarks for Speech AI applications.

  • Continuously improve models and datasets based on experimental results and production feedback.


Requirements


  • At least 3 years of experience developing AI/ML systems or AI-powered products.

  • Strong Python programming skills.

  • Hands-on experience with deep learning frameworks such as PyTorch, TensorFlow, or equivalent technologies.

  • Proven experience deploying AI models into production environments.

  • Solid understanding of speech processing, audio machine learning, and deep learning techniques.

  • Hands-on experience with at least one area of Speech AI, including Automatic Speech Recognition, Text-to-Speech, Audio AI, or Voice AI applications.

  • Ability to design and implement end-to-end AI and machine learning workflows.

  • Experience conducting model experimentation, evaluation, benchmarking, and performance optimization.

  • Ability to translate research concepts into practical and scalable engineering solutions.

  • Strong analytical and problem-solving skills.

  • Product-oriented mindset with the ability to balance model performance, engineering constraints, and user experience.

  • Ability to collaborate effectively with cross-functional engineering and product teams.

Nice to have

  • Experience working with modern Speech AI foundation models, frameworks, or technologies such as Whisper, wav2vec 2.0, NVIDIA NeMo, XTTS, VITS, SpeechBrain, Kaldi, or equivalent solutions.

  • Experience deploying AI models on mobile or edge devices.

  • Experience with real-time audio streaming and processing systems.

  • Knowledge of speaker recognition and speaker diarization.

  • Experience with voice cloning technologies.

  • Strong knowledge of digital audio and signal processing.

  • Experience integrating Large Language Models into AI applications.

  • Experience developing conversational AI or intelligent voice interaction systems.


Benefits


  • Competitive compensation and performance-based incentives based on capabilities and contribution.

  • Full statutory benefits in accordance with applicable labor regulations and company policies.

  • Opportunities to work on challenging AI problems and apply state-of-the-art research to real-world applications.

  • Ownership of the complete AI development lifecycle, from research and experimentation to model development, deployment, monitoring, and continuous optimization.

  • Opportunities to collaborate with experienced AI Engineers, Software Engineers, and Product professionals.

  • Strong exposure to production-scale AI engineering and the development of practical Speech AI and Voice AI products.

  • Continuous learning and professional development opportunities in artificial intelligence, machine learning, speech technologies, and production AI systems.

  • Opportunities to create measurable product impact by developing AI capabilities that directly improve user experiences.


Contact: Giang Van or Thao Phan or Thuong Le


Due to the immense number of applications, only shortlisted candidates will be contacted.
Aloha Consulting Group

About Aloha Consulting Group

Our vision is "To be the leading consulting firm in South East Asia, focused on leveraging human elements and technology to lead our partners towards higher growth and performance together with joy"

Industry
HR & Recruiting
Company Size
11-50 employees
Headquarters
Hanoi & HCMC, VN
Year Founded
2021
Social Media