Job Title: Kafka Engineer/Administrator(w2 requuirement)
Location: St. Louis, MO or Greenwood Village, CO (4 days onsite & 1 day remote per week)
Duration: 6 months base contract; long-term extensions
Work Authorization: Any
Interview: 2 Virtual Interviews; Offer
Job Description:
org: Data Platforms
open to visa candidates
open to relocation candidates
may be 2 openings
They are currently working on migrating all Cloudera Kafka from on prem to AWS. They need this person to handle to workload
Client JD:
Senior Kafka Engineer / Administrator
# About the Role
We operate a portfolio of Apache Kafka clusters across on-premises platforms (Cloudera and Confluent) and *AWS MSK, managed through a **CI/CD pipeline built on Terraform and GitLab*. We are looking for a senior, hands-on Kafka engineer to provide ongoing administration and operational support across these environments, keep our streaming platform healthy and reliable, and coordinate closely with producer and consumer application teams. In parallel, you will contribute to ongoing modernization and workload-transition efforts to AWS MSK.
*This is a collaborative, team-based role.* You will work as part of a cross-functional streaming/data platform team alongside other engineers, SREs, application teams, and the engineering lead. Success depends on strong partnership, shared ownership, knowledge sharing, and clear communication, not working in isolation.
# Key Responsibilities
*Kafka Administration & Ongoing Support (On-Prem + Cloud)*
- Collaborate with the platform team to administer and operate Kafka clusters across Cloudera, Confluent Platform, and AWS MSK.
- Provide day-to-day operational support: incident response, on-call participation, root-cause analysis, and continuous reliability improvements.
- Manage topics, partitions, replication, retention, quotas, ACLs, and consumer groups.
- Tune brokers, producers, and consumers for throughput, latency, and reliability.
- Handle patching, upgrades, capacity planning, and cluster health/performance troubleshooting.
- Configure and maintain security: TLS/SSL, SASL, mTLS, IAM (for MSK), RBAC/ACLs, and encryption at rest and in transit.
AWS MSK Modernization (Parallel Workstream)*
- Partner with the team on ongoing workload-transition efforts from on-prem (Cloudera/Confluent) to AWS MSK.
- Contribute to transition strategies (e.g., MirrorMaker 2 / Confluent Replicator / cluster linking) with minimal downtime and no data loss.
- Validate topic parity, offsets, throughput, and data integrity before and after transitions.
- Help define rollback plans and run dry-runs in lower environments.
*CI/CD & Infrastructure as Code*
- Build and maintain Kafka/MSK infrastructure using *Terraform*.
- Manage automated deployments and configuration through *GitLab CI/CD* pipelines.
- Implement GitOps practices for topic, ACL, and cluster configuration management.
- Participate in code reviews and enforce version control and environment promotion (dev -> test -> prod).
*Observability & Monitoring*
- Build and maintain monitoring, alerting, and dashboards in *Datadog* for Kafka/MSK (broker metrics, consumer lag, partition health, JVM, throughput).
- Work with the team to establish SLAs/SLOs and proactive alerting to reduce incidents and mean-time-to-resolution.
- Correlate Kafka metrics with application and infrastructure telemetry.
*Producer / Consumer Coordination & Teamwork*
- Act as a technical liaison between the Kafka platform team and producer/consumer application teams.
- Advise application teams on client configuration, serialization/schema management, partitioning strategy, idempotence, and exactly-once/at-least-once semantics.
- Support onboarding of new producers/consumers and troubleshoot client-side issues (rebalancing, lag, poison messages).
- Coordinate schedules, testing, and change windows with impacted teams.
- Share knowledge, document runbooks, and contribute to a strong team culture of collaboration and continuous improvement.
# Required Qualifications
- *5+ years in data/platform/infrastructure engineering, with 4+ years operating Apache Kafka in production.*
- Hands-on administration experience with *Cloudera* and *Confluent Platform*.
- Production experience with *AWS MSK* (provisioning, configuration, IAM auth, networking).
- Experience supporting and operating Kafka in production, including on-call and incident response.
- Strong *Terraform* and *GitLab CI/CD* experience for infrastructure automation.
- Hands-on *Datadog* experience for monitoring, dashboards, and alerting.
- Deep understanding of Kafka internals: partitions, replication, ISR, offsets, consumer groups, and delivery semantics.
- Solid AWS fundamentals: VPC, security groups, IAM, KMS, CloudWatch, PrivateLink.
- Strong Linux, networking, and scripting skills (Bash, Python).
- Excellent communication and collaboration skills, with a track record of working effectively in a team.
# Preferred / Nice to Have
- Experience with *MirrorMaker 2*, Confluent Replicator, or cluster linking for replication and workload transition.
- Schema Registry, Kafka Connect, and ksqlDB experience.
- Kafka Streams or Flink familiarity.
- Kubernetes / EKS experience (e.g., Strimzi).
- Certifications: Confluent Certified Administrator, AWS Solutions Architect/DevOps.
- Experience in a regulated or large enterprise environment.
# Success in the First 90 Days
- Ramp up on existing on-prem and MSK clusters and contribute to a health assessment of the current platform.
- Help standardize Datadog dashboards and alerting across all environments.
- Support ongoing operations and strengthen the Terraform + GitLab pipeline for provisioning and topic/ACL management.
- Partner with the team on the parallel MSK modernization effort, including validated transition and rollback plans.