Proven with Industry Leaders
What Cloudera ships, and what it leaves to you
Cloudera Runtime includes Debezium source connectors that run in Kafka Connect on the Cloudera platform, alongside Kafka, Schema Registry, and the rest of Streams Messaging. That saves you sourcing and packaging the connectors yourself.
Cloudera documents the setup steps and stops there. Its documentation states that it does not cover how to use or configure the Debezium connectors. Connecting them to production databases, planning snapshots, sizing log retention, governing schemas, and running the pipelines is the part a partner provides.
| Database | How Debezium reads changes | What the database needs |
|---|---|---|
| Oracle | Reads the redo and archive logs through LogMiner or XStream. | Archive log mode, supplemental logging at the database level and on each captured table, and a dedicated connector user. |
| Db2 | Reads the change-data tables that Db2 populates for each captured table. | Change-data capture tables set up for every table in scope, with the capture process running. |
| SQL Server | Reads the change tables that SQL Server’s own CDC feature maintains. | CDC enabled on the database and on each table you want to capture. |
| PostgreSQL | Reads the write-ahead log through logical decoding. | Logical replication enabled, a replication slot, and a user with replication rights. |
| MySQL | Reads the binary log. | Row-based binary logging enabled and retained long enough to cover connector downtime. |
Prerequisites vary by database version and edition. We confirm them against your environment before a build starts.
How a Debezium CDC pipeline works
Kafka CDC with Debezium has four stages. Each one has decisions that matter in production.
Snapshot
On first start, the connector reads a consistent snapshot of the tables in scope, so Kafka begins with the current state of each table.
Stream
It then reads committed changes from the database log or change tables and writes each insert, update, and delete to a Kafka topic, in order, per key.
Govern
Schema Registry records the structure of each change event, so producers and consumers stay in agreement when a table changes.
Consume
Downstream systems, analytics, NiFi flows, Flink jobs, and AI workloads read the changes from Kafka without querying the source database.
Where CDC projects get stuck
Debezium is reliable once it is running. The failures come from the parts around it.
Log retention gaps
If the connector is down longer than the database keeps its logs, changes are lost and the table has to be snapshotted again. Retention and connector monitoring have to be sized together.
Large initial snapshots
Snapshotting a large core-system table can run for hours and load the source database. It has to be planned for a window, or run incrementally.
Schema changes
A column added or changed in the source flows through as a new event structure. Without schema governance, consumers break silently.
Source system load and access
DBAs and security teams need to approve log access, supplemental logging, and connector credentials. On core banking and policy systems, that approval is often the longest step.
Running Kafka somewhere else?
Cloudera is where we focus, and the same CDC work applies on other Kafka platforms. The engineering questions do not change with the platform: getting log access approved on the source, planning snapshots, preserving order, handling schema changes, and monitoring lag before it becomes data loss.
If your Kafka runs as open-source Apache Kafka or on AWS, see Kafka support. For how CDC fits alongside message queues and APIs, see event-driven architecture.
Built, then run by the same team
We design and build CDC pipelines, then support them in production. Connectors, lag between source and Kafka, and log retention headroom are monitored from our network operations center with Grafana dashboards and alerting, with coverage up to 24x7 as agreed in each contract.
On Cloudera, CDC support sits within our Cloudera support practice, alongside the rest of the platform.
Debezium CDC, answered plainly.
Debezium on Cloudera
What is Debezium CDC?
Debezium is an open-source change data capture tool. Its connectors run in Kafka Connect, read changes from a database’s transaction log or change tables, and publish every insert, update, and delete to Kafka topics, so other systems see changes as they happen without querying the database.
Does Cloudera include Debezium?
Yes. Cloudera Runtime ships Debezium source connectors for Oracle, Db2, SQL Server, PostgreSQL, and MySQL, which run in Kafka Connect on the Cloudera platform. Cloudera’s documentation covers the setup steps, and states that it does not document how to use or configure the Debezium connectors themselves.
Who configures and supports Debezium on Cloudera?
Cloudera supports the platform and ships the connectors. Configuring the connectors against your databases, sizing snapshots and log retention, governing schemas, and operating the pipelines is work your team or a partner does. Duczer East does that work as a Cloudera Premier Partner.
Databases
Can Debezium capture changes from Oracle?
Yes. The Debezium Oracle connector reads changes through LogMiner or XStream. The database needs archive log mode and supplemental logging enabled, at the database level and on each captured table, plus a dedicated connector user.
Can Debezium capture changes from Db2 and SQL Server?
Yes. For Db2, the connector reads the change-data tables Db2 maintains for each captured table. For SQL Server, CDC must be enabled on the database and on each table, and the connector reads the resulting change tables.
Which databases have you implemented CDC from?
We have delivered change data capture into Kafka from Oracle, Db2, SQL Server, PostgreSQL, and MySQL.
CDC beyond Cloudera
Do you do CDC work outside Cloudera?
Yes. Cloudera is our focus, and we also deliver CDC on other Kafka platforms. The design questions are the same wherever the pipeline runs: log access on the source, snapshot planning, ordering, schema changes, and monitoring. For Kafka outside Cloudera, see our Kafka support page.
Why use CDC instead of batch extracts?
Batch extracts copy data on a schedule, so downstream systems are always behind and the extract loads the source system each time it runs. CDC reads only what changed, as it changes, from the database log, which keeps consumers current and keeps load on the core system low.
Support and cost
How do you monitor CDC pipelines?
From our network operations center with Grafana dashboards and alerting, with coverage up to 24x7 as agreed in each contract. We watch connector health, lag between the source database and Kafka, and log retention headroom, so a stalled connector is caught before the database logs roll over.
What does a Debezium CDC engagement cost?
It is quoted per deployment, based on the number of source databases and tables, the volume of change, and whether you want a build, ongoing support, or both. We scope it after a short review of the sources and the target platform.