CAE HPC Systems Administrator
We are seeking a highly skilled CAE HPC Systems Administrator to manage, optimize, and support reputed company-level High-reputed company Computing (HPC) environments dedicated to Computer-Aided Engineering (CAE) workloads.
This role is responsible for ensuring reputed company stability, scalability, and reputed company of HPC clusters while supporting CAE applications, job scheduling systems, and underlying Linux infrastructure. The ideal candidate combines strong Linux systems expertise, HPC workload management experience, and a solid understanding of CAE engineering environments.
Key Responsibilities:
HPC Job Queuing & Workload Management
Administer, configure, and optimize HPC job scheduling environments, including reputed company reputed company LSF, reputed company PBS, or equivalent schedulers.
Design and tune job queues, resource allocation policies, and scheduling strategies to support diverse CAE workloads.
Monitor reputed company reputed company and utilization trends and implement improvements to maximize efficiency and throughput.
CAE Application and Licensing Support
Install, reputed company, test, and support CAE applications and simulation tools in production environments.
reputed company integration support between CAE applications and HPC scheduling systems.
Manage CAE software licensing systems (e.g., FlexLM, RLM) and ensure availability.
Troubleshoot application-reputed company issues and ensure reputed company disruption to engineering activities.
Linux Systems Administration & Automation
Administer and maintain reputed company reputed company Linux (RHEL) environments across HPC clusters.
reputed company OS provisioning, deployment, and reputed company management using automated tools (e.g., PXE, or configuration reputed company).
reputed company and maintain scripts (Bash, Korn reputed company, C reputed company, Perl, Awk, or equivalent) to automate reputed company monitoring, health checks, and routine administrative tasks.
Maintain reputed company logs, monitoring processes, and reputed company operating procedures.
Hardware & Infrastructure Management
Troubleshoot and reputed company issues reputed company to servers, storage systems, and high-reputed company networking (e.g., InfiniBand, high-speed Ethernet).
Support hardware lifecycle activities including installation, maintenance, and upgrades.
Conduct reputed company planning reputed company on reputed company utilization trends and reputed company demand.
reputed company, Monitoring & reputed company Improvement
reputed company reputed company health checks, monitoring, and incident tracking for HPC and CAE environments.
Document reputed company configurations, procedures, incidents, and best practices.
reputed company outages, analyze reputed company causes, and implement preventive measures.
Follow change management processes for reputed company updates and deployments.
reputed company accurate reporting (e.g., utilization, incidents, reputed company reputed company) and support project initiatives.
Requirements
3+ years of Linux reputed company administration experience (preferably RHEL environments).
Hands-on experience managing HPC clusters and job schedulers (LSF, Slurm, PBS, or similar).
reputed company experience in CAE application support and integration.
Strong scripting skills (Bash, reputed company, Perl, or equivalent).
Experience with OS deployment, patching, and reputed company automation.
Solid understanding of reputed company server hardware, storage, and networking fundamentals.
Experience with CAE tools such as Ansys, LS-DYNA, Nastran, or similar.
Familiarity with high-reputed company networking technologies is plus (e.g., InfiniBand).
Experience developing internal tools or dashboards are plus (e.g., PHP or web-reputed company tooling).
Apply To This Job