À distance, n'importe où aux États-Unis
Director of Technical Operations
Description du poste
Deque is seeking a Director of Technical Operations. You will own the infrastructure that Deque’s products run on, and you will still have your hands on it.
Deque’s products help organizations make the digital world accessible to people with disabilities. Those products run in our AWS environment and inside our customers’ networks, and many of them are load-bearing for our customers’ own compliance obligations. When our infrastructure is slow or unavailable, their work stops.
You’ll lead a small, capable systems operations team — small enough that you’ll be in the terminal yourself, not just in the roadmap doc. We want someone with the judgment and scope of a director who still wants to do the work: debug the failing deploy, write the Terraform, run the postmortem. If you’ve been looking for the step where you get real ownership without giving up hands-on engineering, this is that step.
Responsabilités
Product infrastructure and reliability
- Our AWS production environment: the infrastructure behind every Deque SaaS product, defined as code and deployed through pipelines your team owns
- Availability, performance, and capacity against our published customer commitments — and the observability, alerting, and SLO practice that makes those commitments measurable rather than aspirational
- Incident command for major events, and a blameless postmortem process that actually closes its action items
- Backup, disaster recovery, and business continuity, including tested restore procedures and defined recovery objectives
- AWS cost: visibility into where the money goes across production and development, and a standing pipeline of savings work
- Manage the team to achieve 99.99% uptime with the exception of planned maintenance and upgrades
- Managing changes, including releases with gates and maintaining all changes as documented artifacts
Internal IT and business systems
- Administration and integration of the systems Deque runs on — Jira, GitHub, ZenHub, Artifactory, Keycloak, Zendesk, and others
- Vendor relationships for those systems: evaluation, renewals, contract negotiation, and consolidation where it makes sense
- Identity and access management across our tooling, in partnership with Security
- End-user IT operations – managing employee laptops, onboarding/offboarding for the in scope IT and business systems
On-premises customer deployments
- Installation and upgrade of Deque products in customer environments, and the tooling and documentation that make those repeatable rather than bespoke
Security, compliance, and risk
- The operations half of maintaining our ISO-27001 certification: controls, evidence, audit response
- Meet the customer commitments for SLAs for remediating vulnerabilities
- Operational support for initiatives led by our Director of Security
- Data handling practices consistent with GDPR and comparable privacy regimes
Team and practice
- Hiring, developing, and retaining the systems operations team in multiple locations, currently including the US and India; setting the technical bar
- Scaling the team’s capabilities ahead of customer growth – this is the strategic core of the job, not an afterthought. Our tooling needs to be materially stronger a year from now than it is today.
- Documented runbooks, change management, and operational standards
- Accountability for the accessibility of the internal tools and admin interfaces you select and configure. Our accessibility experts can help with the evaluations, if you lack experience, our Deque University procurement courses will help you get up to speed.
- A practical plan for using AI in operations: where it genuinely helps, with governance around AI safety, data handling and change safety
What success looks like in your first year
- You understand our environment well enough to have found the things that worry you, and you’ve told us what they are
- Off-hours coverage and escalation are less frequent and more efficient than when you arrived
- Our ISO-27001 audit is uneventful
- The team has capacity it didn’t have before, through automation
- The IT costs are well understood, optimized and constantly monitored
- Met all SLA’s and SLO’s within acceptable bounds
- Formulated and executed plans to improve highest risk areas in your control
Hours and on-call
- We have follow-the-sun coverage on weekdays. Weekend on-call is limited and centered on planned upgrades and genuine emergencies; you’ll share in it alongside your team. We’d rather tell you this now than have you discover it in week three.
Conditions requises
- 7+ years in infrastructure, SRE, or IT operations with steadily growing responsibility, including at least two years leading engineers
- Deep hands-on, recent (last 18-24 months) AWS experience – not just oversight of someone else’s AWS
- Infrastructure as code in production (Terraform, CloudFormation, CDK, or similar)
- Containers and orchestration (EKS, ECS, Docker, Kubernetes)
- CI/CD pipelines you’ve built or substantially rebuilt, not merely inherited
- Observability and incident response: instrumentation, alerting, on-call practice, postmortems
- Experience in an organization with an active customer base and a support function — you’ve felt the pressure of a customer-facing outage
- Working knowledge of GDPR obligations as they affect infrastructure and data handling
- Clear communication with non-technical audiences, including reporting and dashboards for executives
Additional Skills
-
FedRAMP or StateRAMP experience — the most valuable thing on this list of advantages
-
SOC 2 Type II, VPAT / Section 508 / EN 301 549, or US state privacy regimes (CCPA/CPRA)
-
Experience deploying and supporting software in customer-controlled environments
-
Background in accessibility, assistive technology, or disability inclusion
-
Google Cloud or Azure hands on experience