Skip to main content

Databricks Zerobus Sink Node

Quick Reference​

Zerobus Table Select an existing Databricks Zerobus table data asset. The asset supplies the fully-qualified table name and the table's Avro schema, and — through its connection profile — the ingest endpoint and credentials.

Encoding The wire format for records written to the Zerobus stream: Protobuf or JSON.

Enable Store and Forward Buffer records to local disk during a connectivity outage and replay them when the connection is restored.

Overview​

The Databricks Zerobus Sink Node streams processed records directly into a Databricks Unity Catalog managed Delta table over the Zerobus gRPC ingest API. Unlike the Databricks Sink, it does not stage files or run a COPY INTO, so there is no SQL Warehouse and no staging volume to configure.

How It Works​

  1. Connect: The node opens a single long-lived gRPC stream to the Zerobus endpoint for the selected table.
  2. Encode: Each record is encoded as Protobuf (built from the table's Avro schema) or as JSON, depending on the Encoding setting.
  3. Ingest: Records are pushed onto the stream in batches of up to 10,000.
  4. Confirm: The node waits for Zerobus to confirm the last record of the batch is durable. Because Zerobus delivers in order per stream, that confirms the whole batch.

Delivery is at-least-once. After a reconnect the stream re-sends unacknowledged records, so downstream consumers should deduplicate if they need exactly-once.

Configuration​

Field NameDescriptionRequired?Default
Zerobus TableSelect an existing Databricks Zerobus table data asset (e.g., prod.telemetry.device_events). You can create or import tables in the Data Assets section. The asset's Avro schema is shown below the selection once a table is chosen.YesN/A
EncodingWire format for records written to the Zerobus stream. Options: Protobuf, JSON. Protobuf is built from the table's Avro schema; JSON does not use a schema.YesProtobuf
Enable Store and ForwardBuffers records to a local durable queue during a connectivity outage and replays them automatically once the connection is restored. Prevents data loss and stops the pipeline from stalling. See below.NoOff

Prerequisites​

Everything about where data lands comes from a Databricks Zerobus Table data asset, so you need one before adding the node.

The asset must be linked to a Databricks Zerobus connection profile — a standard Databricks profile is rejected, because it carries no Zerobus ingest endpoint. That profile supplies the workspace Host URL, the Ingest Endpoint, and the Client ID and Client Secret of an OAuth service principal. Any of these missing produces a configuration error at deploy time naming the field to fix.

The asset must also carry an Avro schema. Unlike the Databricks Sink, Zerobus writes typed records directly, so deployment fails with a clear error if the asset has no schema.

Encoding​

Protobuf (default) builds a protobuf descriptor from the selected table's Avro schema and encodes each record against it. A record that does not match the schema is rejected at encode time and reported as an error for that record; the rest of the batch still commits. This is where a field missing from the schema, or a value whose type cannot be converted, surfaces.

JSON opens the stream without a descriptor and serializes each record as JSON. No schema is used, so schema mismatches are not caught by the node before the record is sent.

Store and Forward​

Turning on Enable Store and Forward makes a connectivity failure buffer records to a local durable queue instead of dropping them. A background forwarder replays them once the connection recovers, reopening the stream as needed.

  • Only transient connectivity failures are buffered — network errors, timeouts, and the stream errors Databricks marks as recoverable. A permanent failure (schema, authentication, or permissions) is not buffered and goes through the normal error path.
  • The local queue is capped at 1 GiB by default. Once the cap is reached, incoming records are dropped.
  • The forwarder retries draining the queue every 5 seconds during an outage.

An outage is detected quickly: the node waits at most 20 seconds for a server acknowledgment and at most 60 seconds for a flush, well below the Databricks SDK defaults of 60 seconds and 5 minutes.

When to use the Zerobus Sink​

  • Choose Zerobus Sink for continuous, low-latency streaming into a Unity Catalog managed table, when you do not want to run a SQL Warehouse or manage a staging volume.
  • Choose Databricks Sink when you need the staged COPY INTO path — for example to pass COPY INTO options, or to write to a table that is not a Zerobus target.