Cvent

Senior Site Reliability Engineer (Edge Security)

Cvent  •  $120k - $150k/yr  •  United States (Remote)  •  5 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

Category: Site Reliability Engineering

Req ID: 10933

Remote Eligible: Hybrid

Locations: Tysons Corner, Virginia; Fredericton, Canada; Dallas, Texas

Overview:

Our Culture and Impact

Cvent is a leading meetings, events, and hospitality technology provider with more than 6,000+ employees and 34,000+ customers worldwide, including 60% of the Fortune 500. Founded in 1999, Cvent delivers a comprehensive event marketing and management platform for marketers and event professionals and offers software solutions to hotels, special event venues and destinations to help them grow their group/MICE and corporate travel business. Our technology brings millions of people together at events around the world. In short, we’re transforming the meetings and events industry through innovative technology that powers the human connection.

Cvent's strength lies in its people, fostering a culture where everyone is encouraged to think like entrepreneurs, taking risks and making decisions confidently. We value diverse perspectives and celebrate differences, working together with colleagues and clients to build strong connections.

AI at Cvent: Leading the Future

Are you ready to shape the future of work at the intersection of human expertise and AI innovation? At Cvent, we’re committed to continuous learning and adaptation—AI isn’t just a tool for us, it’s part of our DNA. We’re looking for candidates who are eager to evolve alongside technology. If you love to experiment boldly, share your discoveries, and help define best practices for AI-augmented work, you’ll thrive here. Our team values professionals who thoughtfully integrate AI into their daily work, delivering exceptional results while relying on the human judgment and creativity that drive real innovation.

Throughout our interview process, you’ll have the chance to demonstrate how you use AI to learn, iterate, and amplify your impact. If you’re excited to be part of a team that’s leading the way in AI-powered collaboration, we’d love to meet you.

Site Reliability is about combining development and operations knowledge and skills to help make the organization better. Whether you have a development background and are interested in learning more about operations and security, or have an operations or security background and are interested in developing internal tools and automation, Cvent SRE can benefit from your skillsets. Ultimately, we are looking for passionate people who love learning, love technology and always want to make things better.

The SRE Security team owns Cvent's edge security and routing layer: AWS WAF, Shield Advanced, CloudFront and F5 LTM. That layer protects every core production domain and thousands of customer-facing endpoints, and it is the team that stands between attacks and the applications enterprises depend on.

As a Senior SRE on the SRE Security team, you will be a core platform owner. You will own the infrastructure, carry the on-call rotation, handle the requests that gate deploys and customer onboarding, and ship the work that keeps the production edge secure and reliable at enterprise scale. We are looking for someone with the drive, ownership and ability to take on challenging problems, both technical and process related, in a dynamic, collaborative and highly distributed, multi-disciplinary team environment. You will work closely with product development teams, Information Security, Cloud Infrastructure and other SRE teams.

We use SRE principles such as blameless postmortems, error budgets and a focus on automation to ensure we're constantly improving our knowledge and maintaining a good quality of life. Overall, we're passionate about continuous improvement, learning and participating in dynamic day to day work where success is rewarded with recognition and upward mobility.

In This Role, You Will:

  • Own the edge security and routing layer: WAF rule management, CDK deployments, Shield Advanced policy configuration via AWS Firewall Manager (FMS), and CloudFront distribution management across multiple AWS accounts.
  • Carry the primary on-call rotation for WAF/Shield and the production edge; triage incidents by reading WAF logs and Splunk queries, act as incident commander when needed, write the postmortem and close the loop with stakeholders.
  • Handle the daily walk-up queue: requests for help troubleshooting blocked requests, spikes in traffic, bot scraping, or just guidance on the best design for a new application.
  • Manage the full edge protection estate, including WAF rule lifecycle and WCU cost governance, TLS certificate operations, F5 LTM virtual server management, and CloudFront dependency ownership.
  • Ensure the scalability, performance, and resilience of edge security systems and processes; identify recurring problems and anti-patterns and turn them into guardrails and automation.
  • Develop build, test and deployment automation for edge infrastructure using AWS CDK and PR-based workflows; champion Cvent standards and best practices.
  • Apply AI tooling to accelerate incident triage, cost analysis, and knowledge retrieval, and contribute to expanding the team's AI-assisted workflows.
  • Build the team's knowledge infrastructure: runbooks, WAF rule catalogues, Shield posture guidelines, and platform documentation.
  • Work with product development teams, Information Security and other SRE teams to ensure a holistic understanding of edge security concerns and their effective and efficient resolution.

Here's What You Need:

  • 5+ years of hands-on experience in Site Reliability Engineering, DevOps or cloud operations, with a demonstrated track record of owning reliability, security, and operational excellence in production environments.
  • Hands-on experience with AWS WAF, Shield Advanced, or CloudFront in production at real scale: operational ownership rather than proof-of-concept familiarity. You have authored and managed rules, resolved incidents, and understood the blast radius before deploying.
  • TLS and certificate lifecycle management: certificate renewals, CSR generation, domain routing, and multi-domain remediation for production systems.
  • Infrastructure as Code with AWS CDK (or CloudFormation). You write the infrastructure, review the diffs, and understand what a change does before it lands.
  • Splunk query fluency: WAF log analysis is a core on-call skill. You write the searches and build the dashboards rather than only reading them.
  • Incident management experience on security-adjacent events: able to act as IC, coordinate across product and security stakeholders in real time, write clear incident summaries, and drive RCA.
  • Change management discipline: ability to communicate changes proactively to stakeholders, document rollout strategies, and manage phased production deployments with rollback plans.
  • Fluent in at least one scripting language such as Python, TypeScript, or Bash, enough to automate your own triage and tooling rather than only run scripts others wrote.
  • On-call experience at real interrupt volume. You know what a sustainable rotation feels like and have managed a walk-up queue without losing project weeks.
  • SRE fundamentals: SLIs, SLOs, error budgets and toil accounting, used as working tools rather than vocabulary.
  • Excellent communication skills and a track record of driving alignment across multi-disciplinary teams.
  • Experience with SDLC methodologies and PR-based deployment workflows.

AI & Automation Literacy (Must Have):

Practical understanding and hands-on exposure to AI fundamentals as applied to SRE and operational workflows:

  • Prompt Engineering: ability to design effective prompts for LLMs to assist with incident analysis, RCA generation, runbook creation, and on-call triage.
  • Retrieval-Augmented Generation (RAG): basic understanding of RAG patterns; ability to leverage or contribute to RAG-based internal tools that surface relevant runbooks, past incidents, and knowledge base articles during operational events.
  • AI-assisted Workflow & Process Automation: experience using or building AI-powered automations in operational contexts, such as automated incident summarization, alert enrichment, WAF log analysis, change risk assessment, or post-mortem drafting using LLM integrations (for example via MCP tools, Claude Code, Slack bots, or custom pipelines).

Good to Have Skills:

  • AWS WAF WCU modeling: understanding of WebACL cost structure, rule complexity, and the WCU thresholds that translate directly to monthly spend.
  • AWS Shield Advanced across multi-account environments via Firewall Manager: policy management, account onboarding, transitioning endpoints from count to block mode, and Shield event analysis.
  • Experience with bot mitigation strategies, including AWS Bot Control, token-based traffic classification, and evaluation of third-party vendors such as Datadome.
  • F5 LTM configuration and management.
  • Cloudflare, Fastly, or Datadome: operational experience with complementary edge security, CDN, or bot-protection platforms.
  • CloudFront at scale: distribution management, cache behaviors, custom origins, Origin Access Control, and dependency mapping; experience with API Gateway and ALBs as part of a layered security posture.
  • LLM-based agents or Claude Code as a daily engineering accelerant: built or operated agents for triage, analysis, or automation in an operational context.
  • Experience with APM, monitoring and logging tools (Datadog, PagerDuty, Splunk) and with CI/CD tooling such as GitHub Actions or Jenkins.
  • Disaster recovery planning and execution: multi-region failover, DR runbooks, and RTO/RPO management.
  • Good understanding of containerization concepts (Docker, ECS, EKS, Kubernetes) and of basic networking concepts (DNS, TLS, HTTP, load balancing).
  • Familiarity with risk assessment and management concepts and practices.

Hybrid: 2 days in office

This job posting is intended to comply with all applicable laws. If we learn during the course of our recruitment process that, due to an applicant’s location, further information about the position is required, including certain salary information, this information in this posting will be supplemented accordingly.

The estimated base salary range for new hires into this role is $120,000 - $150,000 annually + bonus depending on factors such as job-related knowledge, relevant experience, and location. We also offer a competitive benefits package, details of which can be found here.

We are not able to offer sponsorship for this position

Physical Demands

Cvent

About Cvent

Join #CventNation! www.cvent.com/careers

What We Do: Cvent is a global market-leading meetings, events, & hospitality technology provider.

Our Mission: To bring people together, power the human connection and transform the meetings & events industry through innovative technology.

Who We Are: With 5,500+ employees & 30,000 customers, we're one of the largest event technology companies in the world, & we're always looking for the next generation of Cventers to join #CventNation!

Cvent is consistently recognized as a Top Workplace by The Washington Post, Washington Business Journal, & certified as a Great Place to Work in the US, UK, & India, among other regional & international accolades.

**Career Fraud Alert: In today's digital-first world, fraudulent job offers and hiring manager impersonations, often via social media or chat/messaging apps, continue to increase. To protect yourself, verify all communications come from a legitimate Cvent email (name@cvent.com) and be cautious of requests for personal information, unsolicited messages, and offers via chat or social media. If you suspect fraud, contact HRHelp@cvent.com immediately and apply for positions directly through our careers site at https://careers.cvent.com. Your safety and security are our priority.

Industry
IT & Software
Company Size
5,001-10,000 employees
Headquarters
Tysons Corner, VA
Year Founded
1999
Website
cvent.com
Social Media