Level 3 Observability Engineer
Woodside EnergyBengaluru
- Experience
- 6 – 9 yrs
- Compensation
- Not disclosed
- Department
- Digital Operations
Skills
About Woodside Energy
Proudly founded in Australia, we are a global energy company providing reliable and affordable energy to help people lead better lives. Woodside Global Solutions in Bengaluru brings together talent, digital expertise and operational excellence to support Woodside's global operations.
Woodside Global Solutions in Bengaluru is being built as a hub of excellence, to drive innovation and global collaboration. We are looking for talented professionals who are eager to make a global impact, helping to shape the future of Woodside together.
About the role
The Observability Engineer is responsible for operationalising Woodside’s enterprise observability capability across applications, infrastructure, cloud, digital platforms and operationally critical services.
The role will support the DSC Observability Centre of Excellence by onboarding applications into Dynatrace and related observability platforms, implementing telemetry standards, building dashboards, defining service health views, tuning alerts, improving log, metric and trace quality, and enabling actionable insights for Digital Operations.
The role partners with Digital IT Teams and requires strong hands-on technical capability, a practical understanding of observability engineering, operational discipline, automation mindset, stakeholder collaboration and a focus on safe, reliable and resilient digital operations through proactive monitoring, observability and operational analytics.
Duties & Responsibilities
- Lead the implementation and operationalisation of Woodside’s enterprise observability capability across applications, infrastructure, cloud, digital platforms and business-critical services.
- Partner with application owners, platform teams and service owners to onboard services into enterprise observability platforms, including service health, ownership, dependency mapping, escalation paths and operational runbooks.
- Configure, maintain and optimise Dynatrace and related observability capabilities including OneAgent, management zones, dashboards, alerting, synthetic monitoring, distributed tracing, service mapping, SLOs and service health views.
- Define, implement and continuously improve telemetry standards across metrics, logs, traces, events, topology, service maps, synthetic monitoring and business transaction monitoring.
- Support Open Telemetry and instrumentation patterns to ensure telemetry data is accurate, consistent, appropriately tagged and aligned to enterprise standards.
- Develop operational dashboards, service health views, reliability reports and leadership insights covering application health, infrastructure health, service performance, SLO adherence, incident trends and alert quality.
- Monitor and improve key operational outcomes including service availability, observability coverage, service performance, alert quality, Mean Time to Detect, Mean Time to Resolve and reduction of monitoring-related operational risk.
- Improve event quality and operational effectiveness through alert tuning, anomaly detection, event correlation, suppression of duplicate or low-value alerts and ServiceNow ITSM and ITOM or equivalent integration.
- Analyse incidents, major incidents and recurring problems to identify monitoring gaps, telemetry defects, reliability risks, service degradation patterns and continual improvement opportunities.
- Develop automation, observability-as-code capabilities and reusable onboarding patterns using appropriate scripting, API, CI/CD and source control practices to improve scalability, consistency and operational efficiency.
- Support AIOps use cases including event correlation, anomaly detection, ticket enrichment, predictive monitoring, operational analytics and human-in-the-loop remediation workflows.
- Provide technical leadership for the Observability Centre of Excellence by establishing engineering standards, reviewing observability designs and guiding engineers from suppliers, vendors in the delivery of enterprise observability capabilities.
- Coach, mentor and uplift observability engineering capability through knowledge sharing, technical governance, reusable standards, onboarding checklists, dashboard templates, runbooks and operational documentation.
- Partner with Enterprise Architecture, Cyber, Cloud, Infrastructure, Network, ServiceNow, Application and Managed Service Provider teams to embed observability requirements into enterprise technology services, project delivery and solution design.
- Ensure observability capabilities support cyber security monitoring, telemetry governance and forensic investigation requirements through the implementation of secure access controls, data protection standards and appropriate platform segregation.
- Own observability platform consumption and financial governance by monitoring licence utilisation, optimising telemetry ingestion and retention practices, and providing recommendations for cost-effective observability coverage, capacity planning and sustainable platform growth.
Skills & Experience
Experience
- 6 to 9 years of experience in IT Operations, Observability, Application Support, Infrastructure Operations, Cloud Operations, Site Reliability Engineering, Platform Engineering, Event Management or enterprise monitoring environments.
- Hands-on experience with enterprise observability or monitoring platforms such as Dynatrace, Grafana, Prometheus, New Relic, AppDynamics, Azure Monitor, AWS CloudWatch, Elastic, Splunk or similar tools.
- Practical experience onboarding applications and infrastructure services into observability platforms, including instrumentation, dashboarding, alerting, dependency mapping and runbook capture.
- Strong hands-on experience or working knowledge of Dynatrace capabilities including OneAgent, dashboards, management zones, alerting profiles, synthetic monitoring, service flow, distributed tracing, problem detection, tagging, SLOs and Davis AI.
- Good understanding of observability concepts including metrics, logs, traces, events, topology, service maps, synthetic monitoring, service health, full-stack monitoring and business transaction monitoring.
- Experience with Open Telemetry concepts, instrumentation patterns, distributed tracing and telemetry data quality.
- Experience creating dashboards, operational scorecards, health views and reports for application, infrastructure and service performance.
- Experience writing queries and analytics using Dynatrace Query Language, Grafana query patterns, PromQL, log analytics or equivalent observability query languages.
- Experience supporting cloud and hybrid environments such as Azure, AWS, VMware, Windows, Linux, databases, middleware, APIs, containers, kuberenetes and enterprise applications.
- Experience working with ITSM tools such as ServiceNow, including incident, problem, change, request and event management processes.
- Understanding of ServiceNow ITOM Event Management concepts including event ingestion, enrichment, deduplication, correlation, resolver routing and alert-to-incident workflows.
- Experience analysing operational data to identify trends, recurring issues, monitoring gaps, noisy alerts, service degradation and improvement opportunities.
- Experience supporting incident response, major incident reviews, problem investigations and continual service improvement activities.
- Experience with scripting and automation using Python, PowerShell, shell scripting, REST APIs, YAML, JSON, GitHub, Azure DevOps, Jenkins or equivalent tooling.
- Exposure to observability-as-code, infrastructure-as-code, CI/CD pipelines, source control and reusable configuration patterns is preferred.
- Understanding of reliability engineering practices including SLIs, SLOs, error budgets, operational readiness, service health and continuous improvement.
- Good knowledge of ITIL processes, especially Incident Management, Problem Management, Change Management, Event Management and Continual Service Improvement.
- Exposure to AIOps, anomaly detection, event correlation, predictive monitoring, AI-assisted troubleshooting or automation-led IT operations.
- Understanding of network performance monitoring concepts including latency, packet loss, DNS, WAN, SD-WAN, firewall dependencies, synthetic testing and service path visibility is preferred.
- Awareness of IT, OT and telecommunications dependencies in operationally critical environments is preferred.
Required Qualifications
- Bachelor’s degree in engineering, Computer Science, Information Technology or a related discipline, or equivalent practical experience in IT Operations, Observability, SRE, Platform Engineering or enterprise monitoring.
- Dynatrace Associate, Dynatrace Professional or equivalent observability platform certification.
- Grafana, Prometheus, Elastic, Splunk, New Relic, AppDynamics or equivalent monitoring / observability certification.
- Demonstrated hands-on experience with enterprise observability, monitoring, telemetry, dashboarding, alerting and operational analytics.
- Practical understanding of ITIL processes and operational support models.
- Ability to work with technical teams and application owners to onboard services, define monitoring requirements, improve alert quality and support operational outcomes.
Preferred Qualifications
- Microsoft Azure, AWS, Kubernetes, Linux or cloud operations certification.
- ServiceNow ITSM or ServiceNow ITOM Event Management certification or equivalent practical experience.
- Experience with Open Telemetry, distributed tracing, log analytics, AIOps, predictive monitoring or automation-led operations.
- Experience using Generative AI or AI-assisted operations tools to support troubleshooting, incident summarisation, knowledge discovery, event correlation or operational automation.
- Exposure to Dynatrace Davis AI, ServiceNow Now Assist, Microsoft Copilot, Azure AI or similar AI-enabled operations capabilities.
- Experience with Ansible, Terraform, GitHub, Azure DevOps, Jenkins or equivalent DevOps / automation tooling.
- Knowledge of CMDB, service mapping, application dependency mapping and event-to-CI relationships.
- ITIL Foundation certification.
- Experience in Oil & Gas, mining, utilities, manufacturing or similar operationally critical industries with awareness of IT, OT and telecommunications dependencies.
Recognition & Reward
Ongoing development
Commitment to your ongoing development, including on-the-job opportunities and formal programs.
Inclusive parental leave
Inclusive parental leave entitlements for both parents.
Values-led culture
A culture led by our Values, where people are trusted to bring their whole self to work.
Flexible work options
Flexible work options to suit how and where you do your best work.
Generous leave
Generous annual leave, sick leave and casual leave.
Cultural and religious leave
Cultural and religious leave, with flexible public holiday opportunities.
Competitive package
A competitive remuneration package featuring performance-based incentives, with uncapped Employer Provident Fund.
Woodside is committed to fostering an inclusive and diverse workforce culture, which is supported by our Values. Inclusion centres on all employees creating a climate of trust and belonging, where people feel comfortable to bring their whole self to work. We also offer supportive pathways for all employees to grow and develop leadership
Ready to apply?
It takes a few minutes. You will get a reference number as soon as your application is received.
Apply for this role