At CyberMaxx, we believe it is our duty to defend against those committed to wide-scale societal disruption through cyberattacks.
We help our customers reduce risk by tightly integrating MDR with offensive security, threat hunting, security research, and digital forensics and incident response (DFIR) to continually adapt to new and evolving threats. Our modern MDR (Managed Detection & Response) approach is tailored to the unique characteristics and risk factors of each customer, enabling us to take full ownership of the response process and, optionally, manage key security controls. By thinking like an adversary and defending like a guardian, we help our customers stay a step ahead of threat actors.
At CyberMaxx, we value humility, transparency, intellectual curiosity, and a customer first approach.
The Director of Platform Engineering owns the architecture, reliability, and scale of the Elastic-based detection and log management platform that underpins CyberMaxx's MDR service.
This is a working architect role, not a pure management role. The Director is expected to be personally accountable for the Elastic estate — able to reason about cluster topology, read an ingest pipeline, and make the call on an upgrade path — while also leading and developing the engineering team that operates it.
The role balances two competing demands. The majority of the department's throughput is operational: keeping client clusters healthy, managing access and credentials, and tuning ingestion and detection quality. The remainder is forward-looking architecture and scale work. The successful candidate will protect capacity for the second without letting the first degrade.
The department's flagship near-term initiative is the consolidation of the Elastic estate onto a single operating model. A migration plan already exists and is owned by the team's senior engineers. The Director's job is to land it.
What You Will Do:
- Platform Architecture: Own the end-to-end architecture of the Elastic estate, including deployment model, cluster topology, tiering, and multi-tenant isolation across the client base. Holds final sign-off authority on platform architecture decisions.
- Platform Consolidation: Lead the migration of the remaining ECE-based estate onto ECK. Execute against the existing engineering plan, close technical gaps, and drive the program to completion without service disruption.
- Hands-On Technical Leadership: Remain directly engaged in the platform — design reviews, upgrade planning, escalated troubleshooting, and prototype work. This role retains working knowledge of the environment rather than managing it at a distance.
- Capacity and Scale: Maintain and evolve the capacity planning model for the Elastic estate. Forecast ingest growth, size clusters and hardware, and drive the procurement cycle for the capacity required to meet contracted service levels.
- Change Control: Chair the weekly Platform Engineering change control forum and represent the department at the enterprise Change Advisory Board. Ensure platform changes are reviewed, staged, and reversible.
- Ingestion and Data Quality: Own the design and health of ingest pipelines, log source onboarding, and data quality standards. Partner with Detection Engineering and Security Operations on detection tuning and alerting quality.
- Upgrade and Lifecycle Management: Plan and execute Elastic version upgrades across all deployment estates with minimal client-facing disruption. Maintain platform currency and manage end-of-life risk.
- Incident Escalation: Act as the named Platform Engineering escalation contact under the Major Incident Management process. Lead root cause analysis for platform-originated incidents and drive remediation to closure.
- Team Leadership: Lead, coach, and develop a team of platform engineers. Own hiring, performance management, career development, and workload allocation. Run the department's operating cadence.
- Cross-Functional Leadership: Chair the Engineering Leadership Roundtable, the standing forum for alignment across software engineering, infrastructure, detection engineering, security operations, and product. Translate platform constraints into terms product and commercial stakeholders can act on.
- Vendor and Licensing Management: Manage the Elastic relationship, including license entitlement, renewal, and true-up. Evaluate platform tooling and negotiate technical terms alongside procurement.
- Integration: Lead platform-side integration of acquired environments — cluster migration, ingest re-pointing, client mapping, and license consolidation.
- Documentation and Continuity: Establish and enforce documentation standards for platform architecture, runbooks, and recovery procedures such that no single person is a continuity risk. Reducing existing key-person dependency is an explicit early deliverable.
- Compliance: Serve as responsible officer for platform-relevant control documentation and support SOC 2 and client audit evidence requests.
What We Are Looking For:
Experience
- Minimum 10 years in infrastructure, platform, or systems engineering, with at least 4 years leading engineering teams.
- Demonstrated ownership of a large-scale, multi-tenant data platform in production.
- Experience operating in a managed services, MSSP, or MDR environment where platform availability is contractually bound.
- Track record of inheriting and stabilizing an environment built by others, including reverse-engineering undocumented systems.
- Track record of delivering a platform migration or consolidation program end to end.
Technical Skills
- Deep, hands-on expertise with the Elastic Stack at scale — Elasticsearch, Kibana, Logstash, Beats and Agent — including multi-tenant cluster design and operation. This is non-negotiable for the role.
- Elastic Cloud on Kubernetes (ECK) depth is strongly preferred. The estate is consolidating onto ECK, and prior experience building or migrating a multi-tenant platform on ECK is the single most valuable technical qualification for this role.
- Strong grounding in log ingestion architecture, data normalization (ECS), index lifecycle management, and retention design.
- Working knowledge of SIEM and detection engineering concepts; familiarity with detection-as-code practices.
- Linux systems administration, containerization and orchestration (Kubernetes), and infrastructure-as-code.
- Identity and access integration: SAML, LDAP, SSO, and role-based access design.
- Capacity modeling and cost-to-serve analysis for data-intensive platforms.
Leadership Skills
- Able to hold architectural authority while remaining genuinely open to being wrong.
- Able to lead execution of a plan authored by others, and to build the team's credibility alongside their own.
- Comfortable running a department where most of the work is unplanned, and able to defend strategic capacity anyway.
- Clear written and verbal communication with both engineers and executives.
- Effective in a distributed, multi-time-zone organization.
Preferred Qualifications
- Elastic certification (Elastic Certified Engineer or Elastic Certified Architect).
- Experience with 24x7 operational environments and formal change management frameworks such as ITIL.
- Prior experience presenting platform posture to enterprise clients or in audit contexts.
Some Of What We Offer
- Flexible Paid Time Off
- 401k with a company match
- Medical, Dental and Vision Coverage
- Voluntary Short Term and Long-Term Disability
- Employee Assistance Program with Mental Health Supplement
- Voluntary Basic, Accidental, and other ancillary life insurance
- Health Savings Account Contribution (with selection of a HDHP)
- 10 annual, paid holidays
CyberMaxx will consider all qualified applicants without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, disability, veteran or military status, age, genetic information, or other characteristics protected by federal, state, or local applicable law.