Monitoring and Incident Response Specialist
Washington, D.C.
Description

Remote Work: Hybrid 


Location: Washington, D.C. 


Position Description
The Monitoring and Incident Response Specialist is responsible for providing real-time monitoring, incident response, and operational support for an enterprise network environment. This role supports the Monitoring and Incident Response Team (MIRT), which operates in a 24x7x365 operational environment to ensure network availability, performance, and security. The specialist monitors network infrastructure and services, investigates alerts and incidents, performs initial troubleshooting and root cause analysis, and escalates issues to appropriate engineering teams.

Requirements

 Network and Service Monitoring

  • Continuously   monitor network infrastructure, applications, and services to ensure   system availability and performance.
  • Monitor  alerts generated by enterprise monitoring platforms and respond to  operational events.
  • Track  network performance metrics and identify anomalies or potential service  disruptions.
  • Monitor  enterprise infrastructure including routers, switches, firewalls, load  balancers, and WAN circuits.

Incident Response and Troubleshooting

  • Investigate  alerts related to network outages, service degradation, and security  events.
  • Perform  initial triage and root cause analysis of incidents affecting network or  application services.
  • Troubleshoot  connectivity issues and coordinate resolution with network engineering, security, and application teams.
  • Escalate critical incidents to appropriate support teams based on severity and  impact.

Network Infrastructure Support

  • Diagnose  issues related to enterprise networking equipment including routers,  switches, firewalls, and load balancers.
  • Assist  with configuration updates and operational changes under established  change management processes.
  • Utilize packet capture and network diagnostic tools to troubleshoot network  anomalies.

Incident Documentation and Reporting

  • Document  incidents, troubleshooting actions, and resolution steps within the IT  service management (ITSM) system.
  • Maintain  detailed incident logs and operational reports for network and  infrastructure events.
  • Provide  updates to stakeholders regarding incident status, impact, and resolution  timelines.

Operational Monitoring and Alert Management

  • Monitor  enterprise systems for health metrics including:
  • Network  availability
  • CPU  utilization
  • Memory  usage
  • Interface  performance
  • System  alerts and alarms

Investigate monitoring alerts and perform operational response procedures.


Requirements

  • Public Trust
  • Minimum  10 years of experience supporting network operations, IT infrastructure  monitoring, or incident response, with 8 years of experience providing IT technical support, performing network/service monitoring, and using   ticketing system
  • Experience  working in enterprise IT environments supporting network or infrastructure  operations.
  • Bachelor’s Degree