Reading from Postgres using DBReader¶
DBReader supports strategy for incremental data reading, but does not support custom queries, like JOIN.
Warning
Please take into account Postgres types
Supported DBReader features¶
- ✅︎
columns - ✅︎
where - ✅︎
hwm, supported strategies: - ❌
hint(is not supported by Postgres) - ❌
df_schema - ✅︎
options(see Postgres.ReadOptions)
Examples¶
Snapshot strategy:
from onetl.connection import Postgres
from onetl.db import DBReader
postgres = Postgres(...)
reader = DBReader(
connection=postgres,
source="schema.table",
columns=["id", "key", "CAST(value AS text) value", "updated_dt"],
where="key = 'something'",
options=Postgres.ReadOptions(partitionColumn="id", numPartitions=10),
)
df = reader.run()
Incremental strategy:
from onetl.connection import Postgres
from onetl.db import DBReader
from onetl.strategy import IncrementalStrategy
postgres = Postgres(...)
reader = DBReader(
connection=postgres,
source="schema.table",
columns=["id", "key", "CAST(value AS text) value", "updated_dt"],
where="key = 'something'",
hwm=DBReader.AutoDetectHWM(name="postgres_hwm", expression="updated_dt"),
options=Postgres.ReadOptions(partitionColumn="id", numPartitions=10),
)
with IncrementalStrategy():
df = reader.run()
Recommendations¶
Select only required columns¶
Instead of passing "*" in DBReader(columns=[...]) prefer passing exact column names. This reduces the amount of data passed from Postgres to Spark.
Pay attention to where value¶
Instead of filtering data on Spark side using df.filter(df.column == 'value') pass proper DBReader(where="column = 'value'") clause.
This both reduces the amount of data send from Postgres to Spark, and may also improve performance of the query.
Especially if there are indexes or partitions for columns used in where clause.
Options¶
PostgresReadOptions
¶
Spark JDBC reading options.
Added in 0.5.0
Replace Postgres.Options → Postgres.ReadOptions
Examples¶
Note
You can pass any value
supported by Spark,
even if it is not mentioned in this documentation. Option names should be in camelCase!
The set of supported options depends on Spark version.
from onetl.connection import Postgres
options = Postgres.ReadOptions(
partitioning_mode="range",
partitionColumn="reg_id",
numPartitions=10,
customSparkOption="value",
)
Parameters:
-
(query_timeout¶int | None, default:None) –The number of seconds the driver will wait for a statement to execute. Zero means there is no limit.
This option depends on driver implementation, some drivers can check the timeout of each query instead of an entire JDBC batch.
-
(fetchsize¶int, default:100000) –Fetch N rows from an opened cursor per one read round.
Tuning this option can influence performance of reading.
Warning
Default value is different from Spark.
Spark uses driver's own value, and it may be different in different drivers, and even versions of the same driver. For example, Oracle has default
fetchsize=10, which is absolutely not usable.Thus we've overridden default value with
100_000, which should increase reading performance.Changed in 0.2.0
Set explicit default value to
100_000 -
(partition_column¶str | None, default:None) –Column used to parallelize reading from a table.
Warning
It is highly recommended to use primary key, or column with an index to avoid performance issues.
Note
Column type depends on
partitioning_mode.partitioning_mode="range"requires column to be an integer, date or timestamp (can be NULL, but not recommended).partitioning_mode="hash"accepts any column type (NOT NULL).partitioning_mode="mod"requires column to be an integer (NOT NULL).
See documentation for
partitioning_modefor more details -
(num_partitions¶int, default:1) –Number of jobs created by Spark to read the table content in parallel. See documentation for
partitioning_modefor more details -
(lower_bound¶int | None, default:None) –See documentation for
partitioning_modefor more details -
(upper_bound¶int | None, default:None) –See documentation for
partitioning_modefor more details -
(session_init_statement¶str | None, default:None) –After each database session is opened to the remote DB and before starting to read data, this option executes a custom SQL statement (or a PL/SQL block).
Use this to implement session initialization code.
Example:
sessionInitStatement = """ BEGIN execute immediate 'alter session set "_serial_direct_read"=true'; END; """ -
(partitioning_mode¶JDBCPartitioningMode, default:<JDBCPartitioningMode.RANGE: 'range'>) –Defines how Spark will parallelize reading from table.
Possible values:
-
range(default) Allocate each executor a range of values from column passed intopartition_column.Spark generates for each executor an SQL query
Executor 1:
Executor 2:SELECT ... FROM table WHERE (partition_column >= lowerBound OR partition_column IS NULL) AND partition_column < (lowerBound + stride)...SELECT ... FROM table WHERE partition_column >= (lowerBound + stride) AND partition_column < (lowerBound + 2 * stride)Executor N:
WhereSELECT ... FROM table WHERE partition_column >= (lowerBound + (N-1) * stride) AND partition_column <= upperBoundstride=(upperBound - lowerBound) / numPartitions.Column type must be integer, date or timestamp.
Note
lower_bound,upper_boundandnum_partitionsare used just to calculate the partition stride, NOT for filtering the rows in table. So all rows in the table will be returned (unlike Incremental strategy).Note
All queries are executed in parallel. To execute them sequentially, use Batch strategy.
-
hashAllocate each executor a set of values based on hash of thepartition_columncolumn.Spark generates for each executor an SQL query
Executor 1:
Executor 2:SELECT ... FROM table WHERE (some_hash(partition_column) mod num_partitions) = 0 -- lower_bound...SELECT ... FROM table WHERE (some_hash(partition_column) mod num_partitions) = 1 -- lower_bound + 1Executor N:
SELECT ... FROM table WHERE (some_hash(partition_column) mod num_partitions) = num_partitions-1 -- upper_boundNote
The hash function implementation depends on RDBMS. It can be
MD5or any other fast hash function, or expression based on this function call. Usually such functions accepts any column type as an input. -
modAllocate each executor a set of values based on modulus of thepartition_columncolumn.Spark generates for each executor an SQL query
Executor 1:
Executor 2:SELECT ... FROM table WHERE (partition_column mod num_partitions) = 0 -- lower_boundExecutor N:SELECT ... FROM table WHERE (partition_column mod num_partitions) = 1 -- lower_bound + 1SELECT ... FROM table WHERE (partition_column mod num_partitions) = num_partitions-1 -- upper_boundNote
Can be used only with columns of integer type.
Added in 0.5.0
Examples¶
Read data in 10 parallel jobs by range of values in
id_columncolumn:Read data in 10 parallel jobs by hash of values inReadOptions( partitioning_mode="range", # default mode, can be omitted partitionColumn="id_column", numPartitions=10, # Options below can be discarded because they are # calculated automatically as MIN and MAX values of `partitionColumn` lowerBound=0, upperBound=100_000, )some_columncolumn:Read data in 10 parallel jobs by modulus of values inReadOptions( partitioning_mode="hash", partitionColumn="some_column", numPartitions=10, # lowerBound and upperBound are automatically set to `0` and `9` )id_columncolumn:ReadOptions( partitioning_mode="mod", partitionColumn="id_column", numPartitions=10, # lowerBound and upperBound are automatically set to `0` and `9` ) -
parse(options)
¶
If a parameter inherited from the ReadOptions class was passed, then it will be returned unchanged. If a Dict object was passed it will be converted to ReadOptions.
Otherwise, an exception will be raised