Back to jobs
Observability & Monitoring Engineer
كي كارد
محافظة بغداد بغداد العراق, IQFull-timeعامSeptember 14, 2026Source: bebee.com
Job Details
Job Summary:
We are seeking a highly skilled Observability Engineer to design, implement, and manage comprehensive monitoring and observability solutions across infrastructure, applications, and digital services. The role will be responsible for ensuring system availability, performance, reliability, and proactive incident detection through the effective use of metrics, logs, traces, dashboards, and alerting.
The position will work closely with Development, Infrastructure, Security, Network, and Operations teams to investigate incidents, perform root-cause analysis, improve monitoring coverage, and identify potential service issues before they impact business operations. The role will also contribute to automation, service-level monitoring through SLIs, SLOs, and SLAs, and the integration of observability platforms with ITSM tools.
The ideal candidate will continuously improve system reliability, alert quality, operational efficiency, and service visibility, while maintaining proper documentation and supporting production incident-response activities.
Statement of Duties:
Design and manage observability and monitoring solutions across infrastructure and applications.
Implement and maintain metrics, logs, traces, dashboards, and alerting.
Develop and maintain monitoring dashboards using tools such as Grafana, Prometheus, Dynatrace, and Datadog.
Configure centralized log collection, analysis, and correlation.
Create and optimize alerts, thresholds, health checks, and anomaly detection.
Monitor application performance, infrastructure capacity, CPU, memory, storage, network, and database performance.
Monitor APIs, microservices, containers, and distributed applications.
Perform incident investigation, troubleshooting, and root-cause analysis (RCA).
Work closely with Development, Infrastructure, Security, Network, and Operations teams to resolve incidents.
Implement proactive monitoring to identify potential failures before they affect services.
Develop scripts and automation for health checks, monitoring, reporting, and operational activities.
Define and maintain SLIs, SLOs, and SLAs for critical services.
Participate in production monitoring and incident-response activities.
Integrate monitoring and alerting platforms with ITSM tools such as ServiceNow and Jira.
Maintain monitoring documentation, operational procedures, and incident reports.
Continuously improve monitoring coverage, alert quality, system reliability, and operational efficiency.
Skills & Educational Background:
Bachelor's degree in Computer Engineering, Science or related field.
Strong knowledge of monitoring and observability concepts.
Experience with Grafana, Prometheus, Dynatrace, Datadog, or similar platforms.
Experience with centralized logging and log analysis.
Good knowledge of Linux and Windows systems.
Understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, TLS, load balancing, and firewalls.
Experience monitoring APIs, microservices, containers, and distribu
Similar jobs
Salesperson (مندوب مبيعات)
Sahel Albaylsan · محافظة بغداد بغداد العراق, IQ
Sales Supervisor - Najaf Area
Jamjoom Pharma · محافظة بغداد بغداد العراق, IQ
Odoo Sales Specialist
Digital Creativity Technologies · محافظة بغداد بغداد العراق, IQ
Property & Logistics Specialist
RELYANT Global, LLC · محافظة بغداد بغداد العراق, IQ