Home/ Cloudera/ Streaming
Cloudera Practice · Streaming

Cloudera streaming: Kafka, NiFi, and Flink, inside your perimeter.

Cloudera streaming is how data moves in real time on the Cloudera platform: Kafka for event streaming, change data capture from core databases, NiFi for flow-based ingest, and Flink for stream processing. Duczer East, a Cloudera Premier Partner, designs, builds, supports, and operates it on Private Cloud Base and Public Cloud, and connects it to the AI your business is putting into production.

Trusted By

Proven with Industry Leaders

DTCC
Koch
ExxonMobil
HCA Healthcare
Ryder
AAA Mid-Atlantic
DTCC
Koch
ExxonMobil
HCA Healthcare
Ryder
AAA Mid-Atlantic
The streaming stack

Kafka, NiFi, and Flink on Cloudera, and what each one is for

Cloudera packages its streaming tools so they share the platform’s security, lineage, and access controls. We support every component below, sized to what your workloads need, on Cloudera Private Cloud Base and Public Cloud.

Streams Messaging

Kafka on Cloudera

Cloudera’s Streams Messaging bundles Apache Kafka with Schema Registry for managing schemas, Streams Messaging Manager for managing and monitoring Kafka, Streams Replication Manager for replicating Kafka data between clusters, and Cruise Control for balancing load across brokers.

Kafka Connect

Change data capture with Debezium

Cloudera ships Debezium source connectors for Oracle, Db2, SQL Server, PostgreSQL, and MySQL, which capture row-level inserts, updates, and deletes from those databases into Kafka topics.

Cloudera DataFlow

NiFi on Cloudera

Flow-based ingest, routing, and transformation, for moving data between files, APIs, databases, and Kafka with provenance on every flow.

Cloudera Streaming Analytics

Flink and SQL Stream Builder

Stream processing on Apache Flink, including SQL Stream Builder for writing continuous queries over Kafka topics in SQL.

Streaming and AI

Why AI initiatives need real-time data

Most AI projects in regulated firms stall on data, not on models. The model is ready; the data it needs is a day old, copied somewhere it should not be, or impossible to trace afterwards. Streaming on Cloudera addresses all three.

Fresh data for models and retrieval

A model or a retrieval pipeline is only as current as the data it reads. Change data capture into Kafka keeps it in step with core systems, instead of with last night’s batch copy.

Agents that act on events

AI agents can react to events as they happen, such as a claim filed or a payment flagged, instead of polling systems on a schedule and missing what changed in between.

Replay for audit and model risk

Kafka keeps an ordered, retained record of events. That lets you reconstruct what data a model or agent saw and when, which is the evidence model-risk and audit reviews ask for.

Inside the perimeter

On Cloudera, the streaming layer runs in your own environment, next to governed data and in-perimeter model serving. Regulated data does not leave your boundary on its way to the model.

In practice

For a property and casualty insurer, our team runs the Kafka and RabbitMQ platform and delivered AI that acts on events as they happen, bringing real-time data to cross-selling across the insurer’s omnichannel touchpoints.

The models themselves run in-perimeter on private AI on Cloudera. How agents get access to that data, and where their limits are enforced, is covered in AI and agent governance.

Change data capture

Kafka CDC with Debezium, from core systems

Banks and insurers keep their most important data in Oracle, Db2, and SQL Server. Change data capture streams every insert, update, and delete from those databases into Kafka, so other systems and AI workloads see changes as they happen without querying the core system.

Cloudera ships Debezium connectors for Oracle, Db2, SQL Server, PostgreSQL, and MySQL. Its documentation covers setup and stops there: configuring and operating the connectors is left to you. We have delivered CDC into Kafka from all five databases.

Replication and resilience

Kafka replication and disaster recovery

Streams Replication Manager replicates Kafka topics, consumer offsets, and configuration between clusters, across data centers or between on-premises and cloud. Schema Registry keeps producers and consumers agreed on message formats, and Cruise Control rebalances load as clusters grow.

Regulated firms have to show that streaming data survives the loss of a site. We design the replication topology, test failover, and document the recovery steps.

Support and delivery

Cloudera streaming support from build to run

The team that designs your streaming layer also supports it in production, as part of our Cloudera support practice.

Build and design

Cluster topology, topic and partition design, schemas, NiFi flows, Flink jobs, and CDC pipelines sized to your workloads.

L3 support and operations

Escalation-tier support and managed operations for the streaming stack, from the brokers and connectors to the flows and jobs that depend on them.

Monitoring from our NOC

Monitoring from our network operations center with Grafana dashboards and alerting, with coverage up to 24x7 as agreed in each contract.

Upgrades and migrations

Upgrades of the streaming stack with rollback planning and staged cutover, and moves from legacy CDH or HDP Kafka onto current Cloudera releases.

Running Kafka outside Cloudera, on open-source Apache Kafka or on AWS? See Kafka support. For how Kafka fits with RabbitMQ and other message queues, see event-driven architecture.

Questions

Cloudera streaming, answered plainly.

Cloudera streaming

What is Cloudera streaming?

Cloudera streaming is the set of Cloudera components for moving and processing data in real time: Streams Messaging, which bundles Kafka with Schema Registry, Streams Messaging Manager, Streams Replication Manager, and Cruise Control; Cloudera DataFlow, built on NiFi; and Cloudera Streaming Analytics, built on Flink with SQL Stream Builder. Kafka Connect, including Debezium change data capture connectors, connects databases and other systems to Kafka.

Which Cloudera streaming components do you support?

All of them, sized to need: Kafka, Schema Registry, Streams Messaging Manager, Streams Replication Manager, Cruise Control, Kafka Connect with Debezium CDC connectors, NiFi and Cloudera DataFlow, and Flink with SQL Stream Builder, on Cloudera Private Cloud Base and Public Cloud.

What is the difference between Kafka on Cloudera and open-source Kafka?

Kafka on Cloudera is Apache Kafka packaged with the tools around it: schema management, a management and monitoring console, cross-cluster replication, and automated load balancing, all governed by Cloudera’s security and lineage controls. Open-source Kafka gives you the brokers, and you assemble and operate the rest yourself. We support both; see our Kafka support page for open-source Kafka and Kafka on AWS.

CDC and Debezium

Does Cloudera support change data capture with Debezium?

Yes. Cloudera Runtime ships Debezium source connectors for Oracle, Db2, SQL Server, PostgreSQL, and MySQL that stream row-level changes into Kafka. Cloudera’s documentation covers setup, and states that it does not document how to use or configure the Debezium connectors themselves. That configuration and operation is the work we do.

Have you implemented CDC from Oracle, Db2, or SQL Server?

Yes. We have delivered change data capture into Kafka from Oracle, Db2, SQL Server, PostgreSQL, and MySQL.

Streaming and AI

Why does real-time streaming matter for AI?

Because models, retrieval pipelines, and AI agents act on whatever data they can see. Streaming keeps that data current, lets agents respond to events as they happen, and keeps a replayable record of what the AI saw, which model-risk and audit reviews need. For a P&C insurer, we delivered AI that acts on events in real time to support cross-selling across its channels.

Can streaming data feed private AI on Cloudera?

Yes. Kafka, NiFi, and Flink run on the same Cloudera platform as Cloudera AI and in-perimeter model serving, so real-time data can reach models and agents without leaving your environment.

Replication, support and cost

How do you handle Kafka replication and disaster recovery on Cloudera?

With Streams Replication Manager, which replicates Kafka topics, consumer offsets, and configuration between clusters. We design the replication topology, test failover, and document the recovery steps, which regulated firms need to show their regulators.

Do you provide 24x7 support for Cloudera streaming?

Coverage is agreed per contract, including up to 24x7 monitoring from our network operations center, and sits within our Cloudera L3 support and managed services.

What does Cloudera streaming support cost?

It is quoted per deployment, based on the components and clusters you run, the coverage you need, and whether you want support, managed operations, or delivery capacity. Your Cloudera subscription is separate, and we resell it if that helps.

Start here

Tell us where your data needs to move in real time.

A CDC pipeline from a core system, a Kafka estate that needs support, or an AI initiative waiting on fresh data. The first conversation is with an engineer, not a salesperson.