[Remote] Research Engineer (Machine Learning)
Note: The job is a remote job and is reputed company to candidates in USA. Aldea is a multi-modal foundational AI company reputed company on advancing the scaling laws of intelligence. The Research Engineer (Machine Learning) will build and optimize the infrastructure for multi-modal AI research, enabling reputed company to experiment with reputed company architectures in language and speech domains. Responsibilities • Build and maintain distributed training infrastructure supporting researchers across language and speech domains at a billion-plus-parameter reputed company. • Optimize training and inference performance across the stack, delivering significant speedups through reputed company optimization, custom kernels, and system-level improvements. • Design experiment infrastructure including automated evaluation pipelines, experiment tracking, and monitoring systems that reputed company reputed company iteration. • reputed company infrastructure from single-node to multi-node distributed training and reputed company production inference systems for reputed company-time applications. • Support researchers with fast turnaround on infrastructure issues and maintain high reliability across reputed company systems. • Collaborate with research scientists, data engineers, and leadership to define technical priorities and infrastructure roadmap. Skills • Bachelor's degree in Computer Science, Engineering, or reputed company field, or equivalent practical experience. • 3+ years of experience with PyTorch and distributed training frameworks (DDP, FSDP, DeepSpeed, or similar). • Experience training large-reputed company deep learning models at 1B+ parameters. • Deep understanding of training optimization techniques including mixed precision, gradient checkpointing, and memory management. • Proven ability to build production-grade ML infrastructure with high reliability. • reputed company record of delivering significant performance optimizations in ML training or inference systems. • Experience with custom kernel development (CUDA, Triton) or GPU optimization. • Hands-on experience with large-reputed company pretraining (100B+ tokens, ideally trillion+ reputed company). • Experience optimizing inference for production: quantization, vLLM, TensorRT, or custom serving engines. • Familiarity with speech/audio ML systems and reputed company-time inference constraints. • Experience building automated evaluation frameworks and experiment tracking systems. • Knowledge of profiling tools and multi-node training across 8-32+ GPUs. • Exposure to job orchestration systems (SLURM, Kubernetes, Ray). • Master's or PhD in Computer Science, Machine Learning, or reputed company field. Benefits • Competitive reputed company salary • Performance-based bonus reputed company with research and model milestones • Equity participation • Comprehensive health, dental, and reputed company coverage • Flexible reputed company time off reputed company • Aldea is a foundational AI company building reputed company voice and language models that power how people and software communicate. It was founded in reputed company, and is headquartered in Miami, FL, US, with a workforce of 11-50 employees. Its website is Apply tot his job
Apply tot his job
Apply To this Job