Home/ Cloudera/ Streaming/ Debezium CDC
Cloudera Streaming · Change Data Capture

Debezium CDC on Cloudera, from your core systems into Kafka.

Debezium change data capture reads every insert, update, and delete from a database’s log and publishes it to Kafka as it happens. Cloudera ships Debezium connectors for Oracle, Db2, SQL Server, PostgreSQL, and MySQL, and leaves configuring and running them to you. Duczer East, a Cloudera Premier Partner, does that work, and has delivered CDC from all five databases.

Trusted By

Proven with Industry Leaders

DTCC
Koch
ExxonMobil
HCA Healthcare
Ryder
AAA Mid-Atlantic
DTCC
Koch
ExxonMobil
HCA Healthcare
Ryder
AAA Mid-Atlantic
Debezium on Cloudera

What Cloudera ships, and what it leaves to you

Cloudera Runtime includes Debezium source connectors that run in Kafka Connect on the Cloudera platform, alongside Kafka, Schema Registry, and the rest of Streams Messaging. That saves you sourcing and packaging the connectors yourself.

Cloudera documents the setup steps and stops there. Its documentation states that it does not cover how to use or configure the Debezium connectors. Connecting them to production databases, planning snapshots, sizing log retention, governing schemas, and running the pipelines is the part a partner provides.

How Debezium captures changes from each database Cloudera ships a connector for, and what each database needs
Database How Debezium reads changes What the database needs
Oracle Reads the redo and archive logs through LogMiner or XStream. Archive log mode, supplemental logging at the database level and on each captured table, and a dedicated connector user.
Db2 Reads the change-data tables that Db2 populates for each captured table. Change-data capture tables set up for every table in scope, with the capture process running.
SQL Server Reads the change tables that SQL Server’s own CDC feature maintains. CDC enabled on the database and on each table you want to capture.
PostgreSQL Reads the write-ahead log through logical decoding. Logical replication enabled, a replication slot, and a user with replication rights.
MySQL Reads the binary log. Row-based binary logging enabled and retained long enough to cover connector downtime.

Prerequisites vary by database version and edition. We confirm them against your environment before a build starts.

How it works

How a Debezium CDC pipeline works

Kafka CDC with Debezium has four stages. Each one has decisions that matter in production.

01

Snapshot

On first start, the connector reads a consistent snapshot of the tables in scope, so Kafka begins with the current state of each table.

02

Stream

It then reads committed changes from the database log or change tables and writes each insert, update, and delete to a Kafka topic, in order, per key.

03

Govern

Schema Registry records the structure of each change event, so producers and consumers stay in agreement when a table changes.

04

Consume

Downstream systems, analytics, NiFi flows, Flink jobs, and AI workloads read the changes from Kafka without querying the source database.

What goes wrong

Where CDC projects get stuck

Debezium is reliable once it is running. The failures come from the parts around it.

Log retention gaps

If the connector is down longer than the database keeps its logs, changes are lost and the table has to be snapshotted again. Retention and connector monitoring have to be sized together.

Large initial snapshots

Snapshotting a large core-system table can run for hours and load the source database. It has to be planned for a window, or run incrementally.

Schema changes

A column added or changed in the source flows through as a new event structure. Without schema governance, consumers break silently.

Source system load and access

DBAs and security teams need to approve log access, supplemental logging, and connector credentials. On core banking and policy systems, that approval is often the longest step.

Why it matters

Change data capture for core banking and insurance systems

Core banking, policy, and claims systems hold the data every other system needs, and they are the systems no one wants to query harder. CDC reads changes from the log rather than the tables, so downstream systems stay current without adding query load to the core.

The same stream of changes is what AI initiatives need: models and agents working from data that is minutes old rather than a day old, and a retained record of what changed and when, which model-risk and audit reviews ask for. On Cloudera, that stream stays inside your perimeter, next to private AI on Cloudera.

How CDC fits with Kafka, NiFi, and Flink on the platform is covered on Cloudera streaming.

Beyond Cloudera

Running Kafka somewhere else?

Cloudera is where we focus, and the same CDC work applies on other Kafka platforms. The engineering questions do not change with the platform: getting log access approved on the source, planning snapshots, preserving order, handling schema changes, and monitoring lag before it becomes data loss.

If your Kafka runs as open-source Apache Kafka or on AWS, see Kafka support. For how CDC fits alongside message queues and APIs, see event-driven architecture.

Support and operations

Built, then run by the same team

We design and build CDC pipelines, then support them in production. Connectors, lag between source and Kafka, and log retention headroom are monitored from our network operations center with Grafana dashboards and alerting, with coverage up to 24x7 as agreed in each contract.

On Cloudera, CDC support sits within our Cloudera support practice, alongside the rest of the platform.

Questions

Debezium CDC, answered plainly.

Debezium on Cloudera

What is Debezium CDC?

Debezium is an open-source change data capture tool. Its connectors run in Kafka Connect, read changes from a database’s transaction log or change tables, and publish every insert, update, and delete to Kafka topics, so other systems see changes as they happen without querying the database.

Does Cloudera include Debezium?

Yes. Cloudera Runtime ships Debezium source connectors for Oracle, Db2, SQL Server, PostgreSQL, and MySQL, which run in Kafka Connect on the Cloudera platform. Cloudera’s documentation covers the setup steps, and states that it does not document how to use or configure the Debezium connectors themselves.

Who configures and supports Debezium on Cloudera?

Cloudera supports the platform and ships the connectors. Configuring the connectors against your databases, sizing snapshots and log retention, governing schemas, and operating the pipelines is work your team or a partner does. Duczer East does that work as a Cloudera Premier Partner.

Databases

Can Debezium capture changes from Oracle?

Yes. The Debezium Oracle connector reads changes through LogMiner or XStream. The database needs archive log mode and supplemental logging enabled, at the database level and on each captured table, plus a dedicated connector user.

Can Debezium capture changes from Db2 and SQL Server?

Yes. For Db2, the connector reads the change-data tables Db2 maintains for each captured table. For SQL Server, CDC must be enabled on the database and on each table, and the connector reads the resulting change tables.

Which databases have you implemented CDC from?

We have delivered change data capture into Kafka from Oracle, Db2, SQL Server, PostgreSQL, and MySQL.

CDC beyond Cloudera

Do you do CDC work outside Cloudera?

Yes. Cloudera is our focus, and we also deliver CDC on other Kafka platforms. The design questions are the same wherever the pipeline runs: log access on the source, snapshot planning, ordering, schema changes, and monitoring. For Kafka outside Cloudera, see our Kafka support page.

Why use CDC instead of batch extracts?

Batch extracts copy data on a schedule, so downstream systems are always behind and the extract loads the source system each time it runs. CDC reads only what changed, as it changes, from the database log, which keeps consumers current and keeps load on the core system low.

Support and cost

How do you monitor CDC pipelines?

From our network operations center with Grafana dashboards and alerting, with coverage up to 24x7 as agreed in each contract. We watch connector health, lag between the source database and Kafka, and log retention headroom, so a stalled connector is caught before the database logs roll over.

What does a Debezium CDC engagement cost?

It is quoted per deployment, based on the number of source databases and tables, the volume of change, and whether you want a build, ongoing support, or both. We scope it after a short review of the sources and the target platform.

Start here

Tell us which systems you need to stream from.

An Oracle or Db2 core system, a SQL Server estate, or a CDC pipeline that keeps falling behind. The first conversation is with an engineer, not a salesperson.