Back to jobs

Observability & Monitoring Engineer

كي كارد

محافظة بغداد بغداد العراق, IQFull-timeعامSeptember 14, 2026Source: bebee.com

Job Details

Job Summary: We are seeking a highly skilled Observability Engineer to design, implement, and manage comprehensive monitoring and observability solutions across infrastructure, applications, and digital services. The role will be responsible for ensuring system availability, performance, reliability, and proactive incident detection through the effective use of metrics, logs, traces, dashboards, and alerting. The position will work closely with Development, Infrastructure, Security, Network, and Operations teams to investigate incidents, perform root-cause analysis, improve monitoring coverage, and identify potential service issues before they impact business operations. The role will also contribute to automation, service-level monitoring through SLIs, SLOs, and SLAs, and the integration of observability platforms with ITSM tools. The ideal candidate will continuously improve system reliability, alert quality, operational efficiency, and service visibility, while maintaining proper documentation and supporting production incident-response activities. Statement of Duties: Design and manage observability and monitoring solutions across infrastructure and applications. Implement and maintain metrics, logs, traces, dashboards, and alerting. Develop and maintain monitoring dashboards using tools such as Grafana, Prometheus, Dynatrace, and Datadog. Configure centralized log collection, analysis, and correlation. Create and optimize alerts, thresholds, health checks, and anomaly detection. Monitor application performance, infrastructure capacity, CPU, memory, storage, network, and database performance. Monitor APIs, microservices, containers, and distributed applications. Perform incident investigation, troubleshooting, and root-cause analysis (RCA). Work closely with Development, Infrastructure, Security, Network, and Operations teams to resolve incidents. Implement proactive monitoring to identify potential failures before they affect services. Develop scripts and automation for health checks, monitoring, reporting, and operational activities. Define and maintain SLIs, SLOs, and SLAs for critical services. Participate in production monitoring and incident-response activities. Integrate monitoring and alerting platforms with ITSM tools such as ServiceNow and Jira. Maintain monitoring documentation, operational procedures, and incident reports. Continuously improve monitoring coverage, alert quality, system reliability, and operational efficiency. Skills & Educational Background: Bachelor's degree in Computer Engineering, Science or related field. Strong knowledge of monitoring and observability concepts. Experience with Grafana, Prometheus, Dynatrace, Datadog, or similar platforms. Experience with centralized logging and log analysis. Good knowledge of Linux and Windows systems. Understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, TLS, load balancing, and firewalls. Experience monitoring APIs, microservices, containers, and distribu

Similar jobs