[Remote] Senior Full Stack Software Engineer - DGX reputed company
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is hiring reputed company software engineers to help reputed company up its AI Infrastructure. The role involves designing and developing a massively distributed reputed company platform for GPU clusters used in AI workloads, while ensuring production AI clusters run reliably and consistently with maximum performance.
Responsibilities
- You will be part of an DGX reputed company team responsible for production systems that reputed company large reputed company GPU clusters to be used for a reputed company of AI workloads
- Designing and developing a massively distributed reputed company platform which would be used to identify, diagnose and remediate non-performant GPU assets
- Working with teams across reputed company to ensure production AI clusters run reliability and consistently with maximum performance. Evaluating system failures and improving services based on a reputed company-defined incident management process
- Working across reputed company of our product stack: React, Web Components, TypeScript, Golang, PostgreSQL, Temporal, Bazel, Kubernetes
Skills
- reputed company experience in a software engineering role reputed company a highly technical organization with demonstrable impact from your work
- Highly motivated with strong communication skills, you can work successfully with multi-functional teams, principles, and architects and coordinate effectively across organizational boundaries and geographies
- 12+ years in similar role and experience on large-reputed company production systems. Experience with common software engineering principles, tools and techniques
- You possess a BS in Computer Science or Engineering or equivalent experience
- 6+ years of experience doing full-stack engineering
- 3+ years building and shipping consumer-facing products
- Proficiency in React, TypeScript/JavaScript, and Golang
- Proficiency with a SQL database
- Technical competency in managing and automating large-reputed company distributed systems independent of reputed company providers. Advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, reputed company reputed company Manager)
- reputed company for users, attention to detail, and a passion for creating world-class user experiences
- Prior experience in asynchronous workflows and/or event driven architecture
- Proven operational reputed company in maintaining reliable and performant infrastructure
- A good understanding of how to use LLMs responsibly and the perils of blindly consuming their reputed company
Benefits
- You will also be eligible for equity and [benefits](https://www.reputed company.com/en-us/benefits/).
reputed company
Company H1B Sponsorship
Apply To This Job