Site Reliability Engineer - Appian

LUXOFT · Sydney NSW 2000 · Full time
Posted 1d ago

Site Reliability Engineer - Appian

KEY POINTS WE FOUND
  • Support the Australian banking sector as a Site Reliability Engineer.
  • Enhance service reliability and operational resilience through monitoring and automation.
  • Drive capacity management and performance optimisation for cloud-native technologies.

Project description

We are seeking a Site Reliability Engineer with more than 10+ years of experience to support our client in the Australian banking sector. The role centers on improving service reliability, availability, and operational resilience through comprehensive monitoring, resilience testing, disaster recovery planning, and management of SLIs and SLOs.

Responsibilities

Improve service reliability, availability, and operational resilience through robust monitoring, resilience testing, disaster recovery planning, and management of SLIs, SLOs, and error budgets.

Design and enhance observability capabilities, including monitoring, alerting, dashboards, synthetic monitoring, automated health checks, and visibility across customer journeys and cloud environments.

Develop automation, self-healing capabilities, operational tooling, Infrastructure as Code and CI/CD practices to reduce manual effort and improve platform reliability.

Participate in major incident response, service restoration, root cause analysis, and continuous improvement initiatives to minimize customer impact and improve recovery times.

Design and implement AI-driven operational solutions, utilize LLMs for incident triage and remediation, and drive adoption of GitHub Copilot and AI-assisted engineering practices.

Drive capacity management, performance optimization and reliability improvements to ensure services meet current and future demand.

Skills

Must have

Demonstrated experience in Site Reliability Engineering (SRE), Production Engineering, DevOps, or Platform Engineering, with a strong understanding of distributed systems, cloud-native technologies, Kubernetes/OpenShift, and modern application architectures.

Appian – Ability to use Appian futures and use scripting & tools to implementation solutions to reads logs, analyses performance impact/issues and drives fixes to resolution.

Having experience & knowledge on Appian environment setups including On Prem & Cloud.

Observability

Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement. Familiarity with tools such as Elastic, Splunk, Grafana, Prometheus, OpenTelemetry, or similar enterprise monitoring platforms, and drive preventative improvements that enhance application performance and availability.

Knowledge of enterprise application, middleware, and content management platforms, including FileNet, ICN, WAS, container platforms, databases and integration technologies.

Reliability and Scalability

Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand.

Familiarity with Infrastructure as Code (IaC), DevSecOps practices and Zero Trust security principles.

Proficiency in DevOps toolkits such as Ansible and Chef to support and drive CI/CD and reliability-automation initiatives.

Programming and scripting languages including Python, Java, Go, PowerShell, or equivalent technologies to support automation and platform engineering initiatives.

Troubleshooting

Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations.

Nice to have

Python

Other

Languages

English: C1 Advanced

Seniority

Senior

Stay Safe While Job Hunting

We vet all employer accounts and do our best to keep job ads safe, but scams can still occur. Be cautious when sharing personal information — never provide financial details or make payments during the application process. For extra security, use the Apply button on our site when proceeding.

Skills

0 of 29 matched
Ai-driven operational solutions and llmsAnalytical thinkingAppianAutomation and self-healing capabilitiesCapacity management and performance optimizationCi/cd tools (e.g., jenkins, github actions)Cloud-native technologiesCollaboration and communication skillsDatabases and integration technologiesDev opsDevsecops practicesDisaster recovery planningDistributed systemsEnglish language proficiency (c1 advanced)Enterprise application platforms (filenet, icn, was)High availability and fault tolerance designIncident response and root cause analysisInfrastructure as code (iac)KubernetesLog analysisMonitoring platforms (elastic, splunk, grafana, prometheus, opentelemetry)OpenshiftPlatform engineeringProduction engineeringProgramming languages (python, java, go, powershell)Scripting & tools for appianSite reliability engineeringTroubleshooting and diagnostic skillsZero trust security principles

LUXOFT

More details

Expiring date
Site Reliability Engineer - Appian | LUXOFT | Glow Up