Home/ Integration & Security/ Kafka Support
Integration · Kafka Support

Apache Kafka support for the clusters your business runs on.

Kafka support is the help an organization needs to keep its Kafka clusters, and the applications that depend on them, running in production: incident response, root-cause analysis, upgrades, tuning, and monitoring. Duczer East provides enterprise Kafka support and Kafka consulting for Apache Kafka, Kafka on Cloudera, and Kafka on AWS, including Amazon MSK, and can run the platform for you as an outsourced team.

Trusted By

Proven with Industry Leaders

DTCC
Koch
ExxonMobil
HCA Healthcare
Ryder
AAA Mid-Atlantic
DTCC
Koch
ExxonMobil
HCA Healthcare
Ryder
AAA Mid-Atlantic
Where your Kafka runs

Apache Kafka, Kafka on Cloudera, and Kafka on AWS

What support has to cover depends on who operates the brokers. On a self-managed cluster that is your team. On Amazon MSK it is AWS, and the work moves to configuration, clients, connectors, and the applications around the cluster. We support all three deployment models.

Apache Kafka

Self-managed open-source Kafka, on-premises or on cloud virtual machines. You own the brokers, the configuration, and the upgrade path, and there is no vendor to call when a cluster misbehaves. We support the cluster as it runs today.

Kafka on Cloudera

Kafka running inside a Cloudera platform, managed alongside the rest of your data estate. We support the Kafka layer as part of the platform it lives in. Platform-wide Cloudera support is covered by our Cloudera practice.

Cloudera support

Kafka on AWS

Amazon MSK, where AWS operates the brokers, and self-managed Kafka on EC2, where you do. On MSK we support what stays yours: cluster configuration, topics, clients, connectors, monitoring, and the applications that depend on them. On EC2 we support the whole cluster.

What support covers

Kafka support services: the brokers and everything that depends on them

Most Kafka problems show up in an application before they show up on a broker: a consumer falls behind, a connector stops, a producer times out. Our support covers the cluster and the producers, consumers, and connectors around it, because that is where incidents start and where they have to be fixed. Monitoring and upgrades have their own sections below.

Incident response and root cause

When a production cluster breaks, we respond, restore service, and find the cause: broker failures, under-replicated partitions, consumer lag that will not clear, rebalancing storms, disk pressure, and certificate or authentication problems. Root-cause analysis comes with the fix, so the same failure does not return.

Performance and capacity

Throughput that falls short, latency that creeps up, and clusters that are running out of room. We tune brokers, producers, consumers, and partition layouts, and plan capacity against the load you expect rather than the load you had.

Kafka monitoring

24x7 Kafka monitoring, consumer lag, and alerting from our NOC

We monitor the clusters we support from our network operations center, using Grafana dashboards and alerting, with coverage up to 24x7 as agreed in each contract. The aim is to see a problem while it is still a trend, before a consumer falls far enough behind that an application downstream notices.

Consumer lag monitoring

Consumer lag is the gap between the newest message in a partition and the last one a consumer group has processed. Rising lag is usually the first sign that an application downstream is falling behind. We track lag per consumer group and partition, and alert on the trend as well as the number.

Kafka metrics

Broker health, under-replicated and offline partitions, request latency, throughput in and out, disk usage, and controller state. We collect the Kafka metrics that predict an incident, not every metric the cluster can emit.

Kafka alerting

Alerts are tuned to what each consumer can tolerate, so a reporting job that can fall an hour behind does not page anyone, and a payments consumer that cannot fall a minute behind does. Each alert routes to an engineer who knows the estate.

Upgrades and migrations

Kafka upgrades without taking the cluster down

A Kafka upgrade touches the brokers, every client that talks to them, and the connectors and streaming applications in between. Older clusters can also face a required migration before they can reach a current release. We plan and run upgrades and migrations as staged changes, so the cluster keeps serving traffic while it moves.

Rolling upgrades

Brokers are upgraded one at a time with the cluster serving traffic throughout, and each step is checked before the next one starts.

Client and connector compatibility

We check producers, consumers, Kafka Streams applications, and connectors against the target version before the brokers move, so the upgrade does not break an application.

A tested rollback

Every upgrade plan includes a rollback path rehearsed in a lower environment, and a point beyond which the change is committed.

Kafka consulting

Kafka consulting from the team that supports it

Kafka consulting is most useful when the people giving the advice also have to live with it. Our consultants design, review, and build on Kafka, and the same team supports the result in production.

Architecture and design

Cluster topology, topic and partition design, retention, replication across sites, and how Kafka fits with your API gateways, message queues, and data platform.

Health checks

A review of a running cluster: configuration, capacity, security settings, monitoring, and the operating practices around it, with a prioritized list of what to fix.

Streaming builds

Event-driven and streaming designs, change data capture from legacy systems, and stream processing, delivered by Kafka consultants who also support what they build.

How to engage

Support, an outsourced Kafka team, or capacity for a project

There are three ways to work with us on Kafka: support and managed services for the clusters you run, an outsourced team that manages and develops the platform for you, or delivery capacity: experienced Kafka developers and engineers for a specific piece of work such as an upgrade, a migration, or a new streaming build. The outsourced team is built the same way on every platform we run.

A persistent core team

Senior engineers stay on your Kafka estate for the life of the contract. They know your clusters, topics, producers, consumers, and release process, and they carry the run.

Capacity on an agreed lead time

When an initiative needs more hands, we add engineers within a lead time agreed in the contract, typically weeks rather than months, then step back down when the work lands.

A knowledge management system

Runbooks, architecture decisions, environment knowledge, and documentation for every project live in a shared knowledge base, so engineers added for a surge work the same way as the core team.

Monitoring from our NOC

Clusters, connectors, and message flows are monitored from our network operations center, with Grafana dashboards and alerting tuned to each consumer.

In practice

A property and casualty insurer

For a P&C insurer, a core team of two manages and develops a platform built on WSO2 API Manager, RabbitMQ, and Kafka, and has ramped to ten engineers to deliver multiple initiatives. The full model is set out under outsourced WSO2 integration team.

The same team runs Kafka alongside API gateways, message queues, and integration platforms; see managed integration services.

Coverage

Coverage set by what the cluster carries

A cluster that carries claims, payments, or customer events needs a different arrangement than one that feeds a nightly report. Coverage hours and response targets are agreed per contract, built around the topics and flows that sit in the path of customers, revenue, or regulatory reporting.

Before agreeing a model, list which producers and consumers depend on each cluster and what happens downstream when they stop. That list sets the coverage requirement.

Pricing

Quoted per deployment

Kafka support is quoted per deployment. The quote depends on the number and size of clusters, the deployment model, the coverage you need, and whether you want support only, managed services, or an outsourced team.

We scope it after a short review of your estate, so the quote reflects what you run rather than a rate card.

Choosing a provider

How to evaluate a Kafka support provider

These are the criteria worth testing before you sign, whoever the provider is.

  • Your deployment, not a reference cluster.

    Ask whether the provider supports what you run: open-source Kafka, Kafka on Cloudera, Amazon MSK, or Kafka on EC2, at the versions you have in production.

  • The applications as well as the brokers.

    Most Kafka incidents involve a producer, a consumer, or a connector. Check that the provider will work on those, not only on broker health.

  • Coverage that matches criticality.

    Confirm who answers outside business hours, how escalation works, and which topics and flows the coverage is built around.

  • Upgrades in scope.

    Check whether the team that answers incidents also plans and runs upgrades, or whether upgrades become a separate engagement.

  • Evidence for audit.

    In a regulated environment, ask what incident reports, root-cause analyses, and change records you will receive.

  • Monitoring you can see.

    Ask what the provider monitors, what alerts on what, and whether your team can see the same dashboards.

For streaming architecture and new builds beyond support, see our system integration services. For how real-time data feeds governed AI, see our Insights on Private AI & Data Architecture.

Questions

Kafka support, answered plainly.

Kafka support

What is Kafka support?

Kafka support is the incident response, root-cause analysis, upgrades, tuning, and monitoring that keep Kafka clusters and the applications that use them running in production. It covers the brokers, the clients that produce and consume messages, and the connectors and tools around them.

Is there commercial support for Apache Kafka?

Apache Kafka is open source, and the Apache Software Foundation does not offer a support desk or a service agreement. Commercial and enterprise Kafka support comes from vendors and from partners such as Duczer East, who support the open-source distribution and the platforms it runs on under a contract with agreed coverage.

Which Kafka distributions do you support?

Open-source Apache Kafka, Kafka on Cloudera, and Kafka on AWS, both Amazon MSK and self-managed Kafka on EC2.

Do you support Amazon MSK?

Yes. On Amazon MSK, AWS operates the brokers and the underlying infrastructure. We support the parts that remain yours: cluster configuration, topics and partitions, security settings, producers and consumers, connectors, monitoring, and the applications that depend on the cluster. For self-managed Kafka on EC2 we support the whole cluster.

Which parts of the Kafka ecosystem do you support?

The full ecosystem: the brokers, Kafka Connect and its connectors, Schema Registry, MirrorMaker 2 for replication between clusters, Kafka Streams applications, and ksqlDB.

Can you migrate our cluster from ZooKeeper to KRaft?

Yes. Apache Kafka 4.0 removed ZooKeeper mode entirely, so a cluster still running on ZooKeeper has to migrate to KRaft on Kafka 3.9 before it can move to 4.x. We plan and run that migration with a tested rollback, then complete the upgrade.

How do you monitor Kafka?

From our network operations center, using Grafana dashboards and alerting. We watch broker health, under-replicated partitions, consumer lag, throughput, and disk usage, and route problems to the engineers who know your estate.

What is Kafka consumer lag, and how should it be monitored?

Consumer lag is the number of messages a consumer group has not yet processed, measured per partition as the gap between the newest offset and the group’s committed offset. A single reading means little. What matters is whether lag is growing, and how long it has grown relative to what that consumer can tolerate. Monitor it per group and partition, and alert on sustained growth rather than on a fixed number.

Do you offer Kafka consulting?

Yes. Our Kafka consultants work on cluster architecture and sizing, topic and partition design, health checks of running clusters, and streaming builds. The same team supports what it designs, so advice does not stop at a recommendation.

Kafka logging

Do you support Kafka logging?

Yes, in both senses. We configure and troubleshoot the logs Kafka itself writes: broker, controller, and client logs, their levels, and their retention. We also build and support Kafka as a log pipeline, carrying application and system logs to platforms such as Splunk or Elastic.

What is the difference between Kafka logs and Kafka application logs?

Kafka stores the messages in each topic partition as a log on disk, so the word refers to the data itself. The application logs are the separate diagnostic files the brokers and clients write about what they are doing. Retention for topic data and retention for diagnostic logs are set separately, and confusing the two is a common cause of full disks.

Engagement and cost

What hours does Kafka support cover?

Coverage hours and response targets are agreed per contract, including up to 24x7 monitoring from our network operations center, based on which clusters and flows sit in the path of customers, revenue, or regulatory reporting.

What does Kafka support cost?

It is quoted per deployment, based on the number and size of clusters, the coverage you need, and whether you want support only, managed services, or an outsourced team. We scope it after a short review of your estate.

Can you run our Kafka platform as an outsourced team?

Yes. A persistent core of senior engineers manages and develops the platform for the life of the contract, more engineers are added within an agreed lead time when an initiative needs them, and a shared knowledge base keeps the work consistent as the team changes. Monitoring runs through our network operations center.

Do you support Kafka alongside other messaging and integration platforms?

Yes. Kafka rarely runs alone. We support it alongside RabbitMQ and other message queue products, API gateways, and integration platforms, and the same team can cover the estate end to end.

Start here

Tell us what your Kafka estate looks like.

A cluster that needs coverage, an upgrade you have been deferring, or a platform that needs a team. The first conversation is with an engineer, not a salesperson.