Manifest Solutions is currently seeking a Senior Enterprise Monitoring and Event Management Engineer.
- Design, develop, and maintain enterprise event management and correlation solutions.
- Build, maintain, and optimize Netcool probes, gateways, automation policies, rules, and event processing workflows.
- Develop alarm correlation, suppression, enrichment, normalization, deduplication, and root cause analytics.
- Design enterprise event reduction and noise suppression strategies.
- Identify opportunities to automate operational monitoring processes.
Platform and Administrative Support
- Administer and support IBM Netcool Operations Insight (NOI) production and non-production environments.
- Maintain platform health, availability, performance, and resiliency.
- Perform platform upgrades, patching, capacity planning, and lifecycle management.
- Troubleshoot platform issues involving event processing, integrations, databases, and infrastructure components.
- Provide Level 3 support for enterprise monitoring and event management solutions.
Monitoring Architecture
- Design monitoring solutions across servers, databases, network devices, cloud platforms, applications, middleware, and infrastructure services.
- Develop monitoring standards, onboarding processes, and alerting strategies for business-critical systems.
- Partner with application teams to create meaningful actionable alerts and reduce alert fatigue.
- Improve service visibility and operational readiness across enterprise environments.
ServiceNow Integration
- Design and support integrations between monitoring platforms and ServiceNow.
- Implement automated incident creation, event enrichment, routing, escalation, and ticket lifecycle management.
- Support Event Management and AIOps initiatives within ServiceNow.
- Integrate monitoring tools with CMDB and service mapping solutions.
Integration Engineering
- Develop and support integrations using REST, SOAP, SNMP, Syslog, webhooks, messaging technologies, and custom APIs.
- Integrate monitoring platforms with infrastructure, application, database, and cloud environments.
- Support enterprise onboarding activities for new monitoring sources and technologies.
Automation & Continuous Improvement
- Develop event-driven automation and remediation workflows.
- Create self-healing and closed-loop operational automation capabilities.
- Implement monitoring automation using scripts, workflows, runbooks, and orchestration platforms.
- Drive monitoring modernization initiatives and platform improvements.
Production Operations Support
- Participate in on-call rotation and critical incident response activities.
- Support enterprise monitoring platforms that provide critical operational visibility to IT and business stakeholders.
- Analyze monitoring failures and implement corrective actions.
- Serve as a technical escalation point for monitoring and integration-related incidents.
Governance, Risk & Compliance
- Support monitoring solutions that are subject to SOX, audit, and compliance requirements.
- Maintain documentation, runbooks, operational procedures, and technical standards.
- Participate in audit activities and evidence collection when required.
- Ensure monitoring solutions align with cybersecurity and operational standards.
Leadership & Soft Skills
- Strong customer service and stakeholder management skills
- Exceptional troubleshooting and analytical abilities
- Ability to communicate effectively with technical and non-technical audiences
- Ability to prioritize multiple projects and operational demands
- Self-directed and highly accountable
- Strong collaboration skills across infrastructure, security, cloud, and application teams
- Ability to lead technical initiatives and mentor junior engineers
- Strategic mindset with the ability to align monitoring capabilities with business outcomes
Required
- 5+ years supporting enterprise monitoring or event management platforms
- IBM Netcool NOI / OMNIbus experience (or closely related enterprise event management platform)
- Linux administration and troubleshooting
- Scripting (Shell, Python, Perl)
- SNMP, Syslog, REST integrations
- ServiceNow integration experience
- Event correlation, suppression, and automation
- Enterprise production support/on-call experience
Preferred
- Dynatrace
- Splunk
- SolarWinds
- Utility industry experience
- SOX-regulated environment
- VMware
- Oracle
- Ansible
- ServiceNow ITOM