Business System Support Analyst
Fully Remote • Remote
Job Type
Full-time
Description

Job Location:

Remote 


Are you EPIC?
Do you have the ability to demonstrate, understand and apply HFD’s core purpose and values in all that you do? At HFD, our mission is to make healthcare more affordable by giving everyone a better way to pay. In order to accomplish this mission, we must ensure that our team is aligned with our E.P.I.C. values:

  • Excellence: Always exceeding expectations!
  • Passionate: Executing with boldness!
  • Innovative: Pioneering a better way!
  • Collaborative: Together we win!

The EPIC Business System Support Analyst we are looking for:
Reporting to the [SRE Manager - TBD], the Analyst provides dedicated operational support to the Site Reliability Engineering (SRE) team by owning ticketing, triage, documentation, and runbook management. This role acts as the first point of contact for incoming requests, business process issues, and defects, ensuring they are properly logged, categorized, and routed. By absorbing this coordination and documentation workload, the Analyst frees SRE engineers to focus on core reliability engineering - availability, performance, automation, and incident response. Secondarily, the role supports the General Engineering roadmap, helping drive execution of key initiatives across site reliability improvements, company efficiencies, and technical and foundational engineering improvements. 


As a Business System Support Analyst, you will:

  • Monitor incoming tickets/alerts across queues, triage by severity and urgency, and route to the correct SRE engineer or on-call rotation
  • Serve as the first point of contact for internal requests to the SRE team; ensure requests are complete, correctly prioritized, and tracked to resolution. 
  • Maintain accurate, current documentation for systems, processes, and recurring issues; ensure knowledge is captured after incidents and change events. 
  • Create, update, and organize runbooks in partnership with SRE engineers; retire outdated runbooks and standardize formatting. 
  • Track ticket volume, triage SLAs, and documentation coverage; produce regular status summaries for SRE leadership. 
  • Perform hands-on technical analysis of incidents and recurring alerts - correlate logs, metrics, and dashboards to isolate root cause and surface trends to SRE engineers. 
  • Write and run SQL queries against production and reporting databases to validate data, investigate anomalies, and answer operational questions without engineering escalation. 
  • Build and maintain recurring reports and dashboards (ticket trends, reliability KPIs, roadmap progress) from queried data. 
  • Use AI and LLM-based tools to accelerate issue analysis - summarize incident timelines and log output, cluster related tickets, surface emerging failure patterns, and draft first-pass runbook and post-incident content for SRE review. 
  • Identify recurring issues or gaps in documentation/runbooks and propose improvements to reduce repeat work. 
  • Liaise between SRE, support, and other engineering teams to keep ticket status and priorities visible. 
  • Coordinate execution of General Engineering roadmap initiatives spanning site reliability improvements, company efficiencies, and technical/foundational engineering improvements. 
  • Support planning and prioritization of the annual General Engineering roadmap in partnership with SRE leadership aligned with company growth and technical strategy
  • Success metrics and KPIs for roadmap initiatives (uptime improvement %, deployment frequency, mean-time-to-recovery, etc.)
  • Share roadmap status and priorities with engineering teams, leadership, and cross-functional stakeholders
  • Flag needed roadmap adjustments based on business needs, technical debt, and emerging infrastructure risks
Requirements
  • 5+ years in a technical support, IT operations, service desk, or analyst role. 
  • Experience with a ticketing platform (e.g., Jira Service Management, ServiceNow, Zendesk). 
  • Working proficiency in SQL - able to independently write joins, aggregations, and filtered queries against relational databases (e.g., PostgreSQL, MySQL, SQL Server). 
  • Demonstrated technical analysis skills: interpreting logs, metrics, and monitoring data to identify patterns and root causes. 
  • Experience with reporting or BI/observability tooling (e.g., Looker, Power BI, Tableau, Grafana, Datadog) a plus. 
  • Hands-on use of AI/LLM tools (e.g., Claude, ChatGPT, Copilot) to analyze issues and trends - summarizing logs and incident data, identifying patterns, and drafting documentation, with sound judgment about verifying output and handling sensitive data. 
  • Strong written communication skills; comfortable authoring and maintaining technical documentation. 
  • Basic understanding of incident management concepts (severity levels, escalation, on-call processes). 
  • Familiarity with runbook/knowledge-base tooling (e.g., Confluence, Notion, internal wikis). 
  • Ability to pay close attention to detail and be self-motivated 
  • Ability to multitask and excel in a fast-paced environment

Benefits: 

  • Medical, Dental, Vision Insurance 
  • 401k with 4% company match. 
  • Time off: Unlimited PTO, 6 days of paid sick time, plus 6 paid holidays and 1 floating holiday (from the HFD approved list).
  • EPIC company culture 
Salary Description
$96,000-$118,000, dependent on experience