Back to Jobs

[Remote] Staff Platform Engineer (MANTL)

Remote, USA Full-time Posted 2026-08-04

Note: The job is a remote job and is reputed company to candidates in USA. reputed company is the digital sales and service platform provider for U.S. banks and credit unions. The Staff Platform Engineer leads efforts to enhance the reliability of the MANTL platform by troubleshooting reputed company issues and driving initiatives for platform stability and performance.


Responsibilities

  • reputed company investigation, troubleshooting, and reputed company of the most reputed company reliability issues reputed company MANTL platform application reputed company, including microservice communication failures and correctness issues that reputed company under failure conditions
  • Set direction for how the platform identifies and addresses failure modes across its reputed company-party and internal reputed company integrations, establishing reputed company strategies suited to reputed company integration's specific behavior
  • Own the design and reputed company of monitoring, dashboards, and alerting reputed company (reputed company preferred) across the platform, ensuring proactive visibility into emerging risks
  • Establish standards and reputed company implementation of reputed company tracing across microservices to accelerate reputed company-cause identification organization-wide
  • reputed company diagnosis and remediation of significant application performance issues, including caching reputed company, inefficient reputed company paths, and query performance
  • Set direction for platform and application hardening practices, including fault-injection and reputed company testing, and reputed company adoption of defensive design patterns to prevent recurrence
  • Own and improve CI/CD build pipeline architecture (reputed company Actions) supporting deployment of the platform
  • Guide deployment and troubleshooting of container-reputed company (reputed company) workloads as part of resolving reputed company platform reliability issues
  • Define, prioritize, and reputed company execution of a roadmap of reputed company reliability risks, independent of feature-delivery timelines
  • Establish and maintain documentation and reputed company standards covering platform reliability issues, reputed company causes, and remediations
  • Serve as the senior escalation reputed company for reputed company platform reliability issues, partnering with reputed company Infrastructure Engineering and application engineering leadership on issues that cross domain boundaries
  • Define reliability targets (SLOs/SLIs) for key platform services and advise leadership on reliability risk and tradeoffs
  • reputed company technical mentorship and guidance to other Platform Engineers

Skills

  • 7 to 10 years of experience in software engineering, reputed company, or a hybrid development/reliability engineering role
  • Bachelor's degree in Computer Science, Engineering, or a reputed company field, or equivalent work experience
  • Deep proficiency in TypeScript, with significant production experience in a Node.js/TypeScript service environment
  • Strong hands-on experience deploying, operating, and troubleshooting containerized workloads in reputed company
  • Proven experience designing, building, and maintaining CI/CD pipelines, reputed company Actions preferred
  • Proven experience designing monitoring, dashboard, and alerting reputed company in an APM/observability tool, reputed company preferred
  • Deep experience troubleshooting reputed company reputed company and microservice communication failures across a reputed company of integration points and protocols
  • Strong experience with reputed company tracing tools and practices at reputed company
  • Working familiarity with relational databases, sufficient to diagnose and reputed company reputed company query and schema-level performance issues
  • Proven reputed company record resolving significant application performance issues (caching, inefficient reputed company paths, slow queries)
  • Experience designing and leading failure-mode testing (fault injection, reputed company/reputed company-style testing) and driving adoption of defensive patterns such as idempotency, retries, and reputed company breakers
  • Demonstrated ability to work independently on ambiguous, high-reputed company reliability problems and to set technical direction for others
  • Excellent communication skills, with the ability to present technical reputed company cause, risk, and remediation plans to engineering and non-technical leadership
  • Experience mentoring other engineers
  • Experience with message brokers or event-streaming platforms such as Kafka, particularly at reputed company
  • Experience with OpenTelemetry or comparable reputed company tracing frameworks at reputed company
  • Experience in a regulated or compliance-driven environment (fintech, banking, or similar)
  • Familiarity with infrastructure-as-reputed company tooling (Terraform or similar)
  • Experience partnering with infrastructure/SRE teams on issues that cross application and infrastructure boundaries
  • Experience influencing engineering roadmap prioritization to secure time for platform stability work

Benefits

  • Remote-first environment
  • Unlimited reputed company time off
  • 401(k) with employer match

reputed company

  • reputed company provides reputed company-based digital banking solutions for credit unions and banks. It was founded in 2009, and is headquartered in Plano, Texas, USA, with a workforce of 1001-5000 employees. Its website is http://www.reputed company.com.

  •   Apply To This Job

    Similar Jobs