Careers at Lexin
Senior Platform / Site Reliability Engineer
Engineering
Australia
Reporting to: | Technical Lead (CTO) – Interim Manager in place until Technical Lead/CTO is in place |
Employment Type: | Full-time — $160,000–$180,000 + super |
Responsible for: | No direct reports (senior individual contributor) |
Department: | Product |
Location: | Remote (Australia only) |
Role Summary:
This is Lexin’s first dedicated Platform / Site Reliability Engineer role — a hands-on senior individual-contributor position with real ownership from day one, not a management seat. Lexin builds SaaS solutions for asset intensive industries, focusing on the indirect material supply chain, connecting into SAP workflows. Platform reliability and uptime, fast recovery and data integrity are core to how the platform is built. This role owns production, safe and consistent deployment, delivers quick recovery times, and helps the business scale as wegrow. This is not a role where the candidate waits for a team to be built around them.
Key Responsibilities / Your role in detail:
– Own the AWS production environment (AU + US), managed as code with Terraform.
– Evolve Observability - build on existing monitoring and alerting systems and processes so issues are identified early and for resolution actions to be executed.
– Own deployment integrity - health checks and rollbacks that keep the weekly release cadence fast and reliable.
– Own Postgres/RDS reliability and performance at scale - backups, restores, query performance, and scaling as data and customer base grow.
– Own the disaster-recovery strategy - set the standard for how Lexin recovers, and validatethrough regular testing.
– Own capacity and scale planning - ensure onboarding new customers never degrades the experience for existing ones.
– Partner with developers on the reliability of the application itself, not just the infrastructure under it.
– Own infrastructure and server patching - keep systems secure, up to date, and reliable through controlled patching.
KPI’s:
– Disaster-recovery plan documented and validated via a live recovery test within the first 3 months, with a clear recovery time objective (RTO) the business can rely on.
– Deploy health checks and rollbacks so failed releases are caught and reversed automatically, securing the integrity and safety of our release schedule.
– Postgres/RDS backup and restore process tested and proven, with query performance monitored and scaling headroom identified ahead of need.
– New customer onboarding (Rest of the World + North America) happens without degrading performance or reliability for existing customers.
– Patching cadence controlled and documented, with infrastructure and servers kept current and secure on a predictable schedule.
Skills, qualifications & Experience:
– 5+ years in Senior platform or infrastructure roles, including 3+ years owning AWS production SaaS.
– Hands-on exposure with Terraform, CI/CD, and Observability today - not just managingothers who do it.
– Deep knowledge of Postgres/RDS in production – must be as comfortable investigating a slow query as writing a Terraform module.
– Has been the first or only platform owner previously and is comfortable building from incomplete documentation.
– Improves systems incrementally without blocking the roadmap - strengthening and optimizing what's already available instead of pushing for unnecessary large-scale change.
– Nice to have: exposure to SOC 2, ISO 27001, or similar controls.
– Nice to have: experience working alongside external managed-service or support partners.
– Nice to have: background in SaaS serving enterprise or industrial customers.