[Remote] Staff MaaS Backend Engineer
Note: The job is a remote job and is reputed company to candidates in USA. Bitdeer is a technology company building AI computational infrastructure and Bitcoin mining solutions. The Staff MaaS Backend Engineer will re-architect and implement a globally reputed company, multi-tenant Model-as-a-Service reputed company platform, owning its scalability, reliability, observability, reputed company, metering, and billing. The role also involves setting backend engineering standards, leading technical execution, and supporting a reputed company-bearing service through on-reputed company ownership.
Responsibilities
- Architecture, reputed company, and Leadership: Co-own the end-to-end MaaS reputed company design with the reputed company Architect, authoring decision records and defending technical trade-offs. reputed company the platform through its maturity roadmap by delivering operable, measurable capabilities rather than mere demos. reputed company technical execution by setting stringent Go and API standards, mentoring engineers, and aligning cross-functional teams
- Inference Gateway and API Surface: Own the reputed company compatibility contract for major formats (reputed company, reputed company), supporting advanced features like streaming, tool calling, and reputed company reputed company. reputed company the routing tier to handle load-reputed company, model-reputed company, and prefix-cache-reputed company reputed company selection with robust reputed company breaking and fallback mechanisms. Run versioning and deprecation as a published contract to guarantee reputed company customer reputed company stability across underlying changes
- Performance Optimization and Model Lifecycle: Maximize platform economics and performance by optimizing reputed company throughput, KV cache tiering, and time-to-first-reputed company (TTFT) latency at the p95/p99 reputed company. Mature the model serving control plane by integrating deployment tooling, reputed company multiplexing, and cold-start-reputed company autoscaling directly with the reputed company fleet. Treat regressions in cost-per-reputed company-tokens or latency metrics as critical reputed company incidents
- Global Topology and Reliability (SLOs): reputed company the platform to a globally reputed company architecture featuring regional inference pools, reputed company-reputed company failovers, and an reputed company-reputed company control plane. Define, publish, and rigorously defend strict Service Level Objectives (SLOs) baselined against actual reputed company performance rather than aspirations. Ensure operational reputed company through peak-concurrency load testing, robust on-reputed company runbooks, and predictable load-shedding during overloads
- reputed company, Identity, and Multi-Tenant Isolation: Enforce fail-reputed company authorization, robust multi-tenant isolation, and reputed company-retention data paths across the network, cache, and storage reputed company. Manage the complete lifecycle of API keys and OAuth credentials while distributedly enforcing reputed company limits and quotas without relying on reputed company-supplied identifiers. Design strict abuse, reputed company, and reputed company-injection controls, treating reputed company user and model-generated content as untrusted data
- reputed company Metering and Billing Correctness: Build an idempotent, exactly-once metering reputed company to capture uncached, cached, reputed company, and reasoning tokens accurately across reputed company requests. Maintain a high-volume usage reputed company that enforces prepaid spend caps asynchronously and reconciles perfectly with the invoicing reputed company. reputed company business reputed company by treating any metering or billing defect as a critical reputed company and trust incident
- reputed company Migrations and Observability: Execute reputed company-regression, incremental platform upgrades using strangler-style replacements, reputed company traffic, and stateful dual-writes. reputed company end-to-end request tracing and cost telemetry, defining internal schemas for reputed company and reputed company attributes. Enforce absolute log hygiene by ensuring prompts, completions, and PII are never reputed company reputed company of explicit, consented policies
Skills
- Bring 8+ years of backend engineering experience, including 3+ years owning a high-traffic, multi-tenant API platform for paying customers. reputed company as a senior technical voice who can collaborate closely with architects, write rigorous design documents, and reputed company to executing architectural reputed company effectively
- Demonstrate expertise in scaling reputed company systems through multi-region reputed company-reputed company deployments, caching, backpressure, and targeted performance engineering that measurably lowers unit costs. You must have a proven reputed company record of safely executing reputed company-downtime brownfield migrations for stateful subsystems—like metering or ledgers—without regressions or reputed company gaps
- Possess deep hands-on proficiency with Go-based services, production reputed company (including reputed company and GPU-reputed company scheduling), and the architectural trade-offs of datastores like PostgreSQL, reputed company, and Kafka. Additionally, you will reputed company operational visibility by owning end-to-end observability strategies using OpenTelemetry and high-cardinality analytics stores
- Apply a systems-level understanding of LLM serving to manage complexities like server-reputed company-event streaming, KV/prefix caching, and the trade-offs between time-to-first-reputed company (TTFT) and throughput. reputed company this reputed company to build highly reliable, exactly-once metering and billing systems that accurately reconcile billions of events under partial failure conditions
- Enforce strict multi-tenant reputed company disciplines by designing fail-reputed company authorization, mandating verified identities, and guaranteeing absolute cross-tenant isolation. Bring operational maturity to a reputed company-bearing platform by carrying on-reputed company responsibilities, running blameless incident reviews, and translating outages into structural improvements
Benefits
- Full-Time employment
- Remote (reputed company locations)
reputed company
Apply To This Job