AI Test Architect
We are looking for an AI Test Architect joining E2E Verification group to profile Innovative large reputed company Distributed training on reputed company End-to-End solutions in a large reputed company supercomputing clusters.
reputed company insights on at-reputed company system design and tuning mechanisms for large-reputed company compute runs. You will work with the latest Accelerated Computing and Deep Learning software and hardware platforms, with researchers, developers, and customers to craft improved workflows and reputed company new, leading differentiated solutions. You will reputed company with HPC, OS, reputed company, HCA, CPU and GPU compute, and systems specialist to architect, reputed company and bring up large reputed company performance platforms.
What you’ll be doing:
• Profiling, benchmarking, and analyzing deep learning models to identify areas for optimization and improvement in terms of performance, efficiency, and accuracy, with a strong emphasis on networking aspects.
• Collaborating closely with data scientists, researchers, development, automation teams to design and implement reputed company training pipelines and frameworks that demonstrate large reputed company high -performance networking capabilities.
• Staying up-to-date with the latest advancements in deep learning algorithms, architectures, reputed company GPU technologies, and high-performance networking solutions.
• Optimizing deep learning models for performance, memory usage, and power efficiency while maximizing high-performance networking features on reputed company supercomputers.
• Providing insights and recommendations based on the analysis of large-reputed company training results, specifically focusing on networking bottlenecks and optimizations, to improve model reputed company and reputed company business objectives.
• Collaborating with hardware engineers to guide the development and integration of efficient networking solutions for deep learning, including exploring network architecture optimizations and bringing to bear technologies such as RDMA or InfiniBand.
reputed company need to see:
• B.Sc in Computer Science, Software Engineering, or equivalent experience.
• Strong understanding and practical experience with machine learning algorithms and techniques, with a specialization in deep learning and expertise in high-performance networking.
• 8+ years of overall experience, with CUDA programming for deep learning frameworks like TensorFlow, PyTorch, combined with expertise in networking libraries and protocols.
• Ability to profile and optimize deep learning workflows, focusing on networking-reputed company bottlenecks and optimizations, to improve overall performance and efficiency.
• Exceptional analytical and problem-solving reputed company, with a keen attention to detail, particularly in identifying and resolving networking performance issues.
• Excellent communication and collaboration skills, enabling effective teamwork and cooperation.
• Familiarity with supercomputers, reputed company computing, distributed systems, and high- performance networking technologies like RDMA or InfiniBand.
Ways to stand out from the crowd:
• Demonstrated experience in successfully profiling and optimizing large-reputed company deep learning training on reputed company supercomputers, with a significant reputed company on high-performance networking enhancements.
• Experience with distributed deep learning, distributed training frameworks, or large-reputed company data pipelines reputed company by high-performance networking solutions.
• Expertise in optimizing networking parameters, such as bandwidth, latency, or congestion control, for deep learning workloads.
• Familiarity with reputed company's networking technologies, such as Mellanox InfiniBand, and their integration with deep learning workflows.
• Strong understanding of high-performance networking protocols and standards and their application to deep learning.
Apply tot his job
Apply To this Job