DataSinkSharedOptions
laktory.models.datasinks.DataSinkSharedOptions
¤
Bases: BaseModel, PipelineChild
Options for a sink shared by multiple writers: other nodes of the same pipeline and/or other pipelines.
Each writer owns its rows, and a full refresh of a writer only deletes and reprocesses its own rows, leaving the data of the other writers untouched. Each writer can therefore be executed and refreshed independently, in any order, possibly in parallel. The rows of a writer are identified either by:
- a writer column (default): each written row carries the identifier of its writer in
column,{pipeline_name}.{node_name}by default; - a SQL predicate (
where), e.g.client_id = 23, matching the rows written by the writer: no column is added.
Required on every sink of a target written by several nodes of a pipeline or by several
pipelines. shared: true is equivalent to shared: {}.
Examples:
import laktory as lk
sink = lk.models.UnityCatalogDataSink(
schema_name="finance",
table_name="pooled_prices",
mode="APPEND",
shared={"writer_id": "client_a"},
)
print(sink.shared.writer_id)
# > client_a
sink = lk.models.UnityCatalogDataSink(
schema_name="finance",
table_name="pooled_prices",
mode="APPEND",
shared={"where": "client_id = 23"},
)
print(sink.shared.uses_writer_column)
# > False
| PARAMETER | DESCRIPTION |
|---|---|
column
|
Name of the column storing the writer identifier.
TYPE:
|
where
|
SQL predicate matching the rows owned by the writer, e.g.
TYPE:
|
writer_id_
|
Identifier stored in
TYPE:
|
| ATTRIBUTE | DESCRIPTION |
|---|---|
uses_writer_column |
TYPE:
|
uses_writer_column
property
¤
True if written rows carry the writer identifier in column.