Compare commits

..
4 Commits
55 changed files with 856 additions and 2009 deletions
-3
View File
@@ -51,6 +51,3 @@ scm-source.json
docs/build/
docs/source/_static/
docs/source/_templates/
# Pycharm IDE
.idea/
+4 -11
View File
@@ -11,13 +11,8 @@ Global/Universal
- **PATRONI\_NAME**: name of the node where the current instance of Patroni is running. Must be unique for the cluster.
- **PATRONI\_NAMESPACE**: path within the configuration store where Patroni will keep information about the cluster. Default value: "/service"
- **PATRONI\_SCOPE**: cluster name
- **PATRONI\_LOG\_LEVEL**: sets the general logging level. Default value is **INFO** (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
- **PATRONI\_LOG\_FORMAT**: sets the log formatting string. Default value is **%(asctime)s %(levelname)s: %(message)s** (see `the LogRecord attributes <https://docs.python.org/3.6/library/logging.html#logrecord-attributes>`_)
- **PATRONI\_LOG\_DATEFORMAT**: sets the datetime formatting string. (see the `formatTime() documentation <https://docs.python.org/3.6/library/logging.html#logging.Formatter.formatTime>`_)
- **PATRONI\_LOG\_DIR**: Directory to write application logs to. The directory must exist and be writable by the user executing Patroni. If you set this env variable, the application will retain 4 25MB logs by default. You can tune those retention values with `PATRONI_LOG_FILE_NUM` and `PATRONI_LOG_FILE_SIZE` (see below).
- **PATRONI\_LOG\_FILE\_NUM**: The number of application logs to retain.
- **PATRONI\_LOG\_FILE\_SIZE**: Size of patroni.log file (in bytes) that triggers a log rolling.
- **PATRONI\_LOG\_LOGGERS**: Redefine logging level per python module. Example ``PATRONI_LOG_LOGGERS="{patroni.postmaster: WARNING, urllib3: DEBUG}"``
- **PATRONI\_LOGLEVEL**: sets the general logging level (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
- **PATRONI\_REQUESTS_LOGLEVEL**: sets the logging level for all HTTP requests e.g. Kubernetes API calls (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
Bootstrap configuration
-----------------------
@@ -41,13 +36,11 @@ Consul
- **PATRONI\_CONSUL\_KEY**: (optional) File with the client key. Can be empty if the key is part of certificate.
- **PATRONI\_CONSUL\_DC**: (optional) Datacenter to communicate with. By default the datacenter of the host is used.
- **PATRONI\_CONSUL\_CHECKS**: (optional) list of Consul health checks used for the session. If not specified Consul will use "serfHealth" in additional to the TTL based check created by Patroni. Additional checks, in particular the "serfHealth", may cause the leader lock to expire faster than in `ttl` seconds when the leader instance becomes unavailable.
- **PATRONI\_CONSUL\_REGISTER\_SERVICE**: (optional) whether or not to register a service with the name defined by the scope parameter and the tag master, replica or standby-leader depending on the node's role. Defaults to **false**
- **PATRONI\_CONSUL\_SERVICE\_CHECK\_INTERVAL**: (optional) how often to perform health check against registered url
Etcd
----
- **PATRONI\_ETCD\_HOST**: the host:port for the etcd endpoint.
- **PATRONI\_ETCD\_HOSTS**: list of etcd endpoints in format 'host1:port1','host2:port2',etc...
- **PATRONI\_ETCD\_HOSTS**: list of etcd endpoints in format host1:port1,host2:port2,etc...
- **PATRONI\_ETCD\_URL**: url for the etcd, in format: http(s)://(username:password@)host:port
- **PATRONI\_ETCD\_PROXY**: proxy url for the etcd. If you are connecting to the etcd using proxy, use this parameter instead of **PATRONI\_ETCD\_URL**
- **PATRONI\_ETCD\_SRV**: Domain to search the SRV record(s) for cluster autodiscovery.
@@ -86,7 +79,7 @@ PostgreSQL
- **PATRONI\_SUPERUSER\_PASSWORD**: password for the superuser, set during initialization (initdb).
REST API
--------
--------
- **PATRONI\_RESTAPI\_CONNECT\_ADDRESS**: IP address and port to access the REST API.
- **PATRONI\_RESTAPI\_LISTEN**: IP address and port that Patroni will listen to, to provide health-check information for HAProxy.
- **PATRONI\_RESTAPI\_USERNAME**: Basic-auth username to protect unsafe REST API endpoints.
+2 -25
View File
@@ -10,18 +10,6 @@ Global/Universal
- **namespace**: path within the configuration store where Patroni will keep information about the cluster. Default value: "/service"
- **scope**: cluster name
Log
---
- **level**: sets the general logging level. Default value is **INFO** (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
- **format**: sets the log formatting string. Default value is **%(asctime)s %(levelname)s: %(message)s** (see `the LogRecord attributes <https://docs.python.org/3.6/library/logging.html#logrecord-attributes>`_)
- **dateformat**: sets the datetime formatting string. (see the `formatTime() documentation <https://docs.python.org/3.6/library/logging.html#logging.Formatter.formatTime>`_)
- **dir**: Directory to write application logs to. The directory must exist and be writable by the user executing Patroni. If you set this value, the application will retain 4 25MB logs by default. You can tune those retention values with `file_num` and `file_size` (see below).
- **file\_num**: The number of application logs to retain.
- **file\_size**: Size of patroni.log file (in bytes) that triggers a log rolling.
- **loggers**: This section allows redefining logging level per python module
- **patroni.postmaster: WARNING**
- **urllib3: DEBUG**
Bootstrap configuration
-----------------------
- **dcs**: This section will be written into `/<namespace>/<scope>/config` of a given configuration store after initializing of new cluster. This is the global configuration for the cluster. If you want to change some parameters for all cluster nodes - just do it in DCS (or via Patroni API) and all nodes will apply this configuration.
@@ -45,11 +33,6 @@ Bootstrap configuration
- **restore\_command**: command to restore WAL records from the remote master to standby leader, can be different from the list defined in :ref:`postgresql_settings`
- **archive\_cleanup\_command**: cleanup command for standby leader
- **recovery\_min\_apply\_delay**: how long to wait before actually apply WAL records on a standby leader
- **slots**: define permanent replication slots. These slots will be preserved during switchover/failover. Patroni will try to create slots before opening connections to the cluster.
- **my_slot_name**: the name of replication slot. It is the responsibility of the operator to make sure that there are no clashes in names between replication slots automatically created by Patroni for members and permanent replication slots.
- **type**: slot type. Could be ``physical`` or ``logical``. If the slot is logical, you have to additionally define ``database`` and ``plugin``.
**database**: the database name where logical slots should be created.
**plugin**: the plugin name for the logical slot.
- **method**: custom script to use for bootstrapping this cluster.
See :ref:`custom bootstrap methods documentation <custom_bootstrap>` for details.
When ``initdb`` is specified revert to the default ``initdb`` command. ``initdb`` is also triggered when no ``method``
@@ -86,8 +69,6 @@ Most of the parameters are optional, but you have to specify one of the **host**
- **key**: (optional) file with the client key. Can be empty if the key is part of **cert**.
- **dc**: (optional) Datacenter to communicate with. By default the datacenter of the host is used.
- **checks**: (optional) list of Consul health checks used for the session. If not specified Consul will use "serfHealth" in additional to the TTL based check created by Patroni. Additional checks, in particular the "serfHealth", may cause the leader lock to expire faster than in `ttl` seconds when the leader instance becomes unavailable
- **register\_service**: (optional) whether or not to register a service with the name defined by the scope parameter and the tag master, replica or standby-leader depending on the node's role. Defaults to **false**
- **service\_check\_interval**: (optional) how often to perform health check against registered url
Etcd
----
@@ -163,10 +144,8 @@ PostgreSQL
REST API
--------
- **connect\_address**: IP address (or hostname) and port, to access the Patroni's REST API. All the members of the cluster must be able to connect to this address, so unless the Patroni setup is intended for a demo inside the localhost, this address must be a non "localhost" or loopback addres (ie: "localhost" or "127.0.0.1"). It can serve as a endpoint for HTTP health checks (read below about the "listen" REST API parameter), and also for user queries (either directly or via the REST API), as well as for the health checks done by the cluster members during leader elections (for example, to determine whether the master is still running, or if there is a node which has a WAL position that is ahead of the one doing the query; etc.) The connect_address is put in the member key in DCS, making it possible to translate the member name into the address to connect to its REST API.
- **listen**: IP address (or hostname) and port that Patroni will listen to for the REST API - to provide also the same health checks and cluster messaging between the participating nodes, as described above. to provide health-check information for HAProxy (or any other load balancer capable of doing a HTTP "OPTION" or "GET" checks).
- **connect\_address**: IP address (or hostname) and port, to access the Patroni's REST API. It can serve as a endpoint for HTTP health checks (read below about the "listen" REST API parameter), and also for user queries (either directly or via the REST API), as well as for the health checks done by the cluster members during leader elections (for example, to determine whether the master is still running, or if there is a node which has a WAL position that is ahead of the one doing the query; etc.) The connect_address is put in the member key in DCS, making it possible to translate the member name into the address to connect to its REST API.
- **listen**: IP address (or hostname) and port that Patroni will listen to for the REST API - to provide also the same health checks and cluster messaging between the participating nodes, as described above. to provide health-check information for HAProxy (or any other load balancer capable of doing a HTTP "OPTION" or "GET" checks)
- **Optional**:
- **authentication**:
- **username**: Basic-auth username to protect unsafe REST API endpoints.
@@ -175,8 +154,6 @@ REST API
- **certfile**: Specifies the file with the certificate in the PEM format. If the certfile is not specified or is left empty, the API server will work without SSL.
- **keyfile**: Specifies the file with the secret key in the PEM format.
.. _patronictl_settings:
CTL
---
- **Optional**:
-1
View File
@@ -76,7 +76,6 @@ Also, the following Patroni configuration options can be changed only dynamicall
- loop_wait: 10
- retry_timeouts: 10
- maximum_lag_on_failover: 1048576
- check_timeline: false
- postgresql.use_slots: true
Upon changing these options, Patroni will read the relevant section of the configuration stored in DCS and change its
+7 -169
View File
@@ -3,168 +3,6 @@
Release notes
=============
Version 1.5.4
-------------
This version implements flexible logging and fixes a number of bugs.
**New features**
- Improvements in logging infrastructure (Alexander Kukushkin, Lucas Capistrant, Alexander Anikin)
Logging configuration could be configured not only from environment variables but also from Patroni config file. It makes it possible to change logging configuration in runtime by updating config and doing reload or sending SIGHUP to the Patroni process. By default Patroni writes logs to stderr, but now it becomes possible to write logs directly into the file and rotate when it reaches a certain size. In addition to that added support of custom dateformat and the possibility to fine-tune log level for each python module.
- Make it possible to take into account the current timeline during leader elections (Alexander Kukushkin)
It could happen that the node is considering itself as a healthiest one although it is currently not on the latest known timeline. In some cases we want to avoid promoting of such node, which could be achieved by setting `check_timeline` parameter to `true` (default behavior remains unchanged).
- Relaxed requirements on superuser credentials
Libpq allows opening connections without explicitly specifying neither username nor password. Depending on situation it relies either on pgpass file or trust authentication method in pg_hba.conf. Since pg_rewind is also using libpq, it will work the same way.
- Implemented possibility to configure Consul Service registration and check interval via environment variables (Alexander Kukushkin)
Registration of service in Consul was added in the 1.5.0, but so far it was only possible to turn it on via patroni.yaml.
**Stability Improvements**
- Set archive_mode to off during the custom bootstrap (Alexander Kukushkin)
We want to avoid archiving wals and history files until the cluster is fully functional. It really helps if the custom bootstrap involves pg_upgrade.
- Apply five seconds backoff when loading global config on start (Alexander Kukushkin)
It helps to avoid hammering DCS when Patroni just starting up.
- Reduce amount of error messages generated on shutdown (Alexander Kukushkin)
They were harmless but rather annoying and sometimes scary.
- Explicitly secure rw perms for recovery.conf at creation time (Lucas)
We don't want anybody except patroni/postgres user reading this file, because it contains replication user and password.
- Redirect HTTPServer exceptions to logger (Julien Riou)
By default, such exceptions were logged on standard output messing with regular logs.
**Bug fixes**
- Removed stderr pipe to stdout on pg_ctl process (Cody Coons)
Inheriting stderr from the main Patroni process allows all Postgres logs to be seen along with all patroni logs. This is very useful in a container environment as Patroni and Postgres logs may be consumed using standard tools (docker logs, kubectl, etc). In addition to that, this change fixes a bug with Patroni not being able to catch postmaster pid when postgres writing some warnings into stderr.
- Set Consul service check deregister timeout in Go time format (Pavel Kirillov)
Without explicitly mentioned time unit registration was failing.
- Relax checks of standby_cluster cluster configuration (Dmitry Dolgov, Alexander Kukushkin)
It was accepting only strings as valid values and therefore it was not possible to specify the port as integer and create_replica_methods as a list.
Version 1.5.3
-------------
Compatibility and bugfix release.
- Improve stability when running with python3 against zookeeper (Alexander Kukushkin)
Change of `loop_wait` was causing Patroni to disconnect from zookeeper and never reconnect back.
- Fix broken compatibility with postgres 9.3 (Alexander)
When opening a replication connection we should specify replication=1, beacuse 9.3 does not understand replication='database'
- Make sure we refresh Consul session at least once per HA loop and improve handling of consul sessions exceptions (Alexander)
Restart of local consul agent invalidates all sessions related to the node. Not calling session refresh on time and not doing proper handling of session errors was causing demote of the primary.
Version 1.5.2
-------------
Compatibility and bugfix release.
- Compatibility with kazoo-2.6.0 (Alexander Kukushkin)
In order to make sure that requests are performed with an appropriate timeout, Patroni redefines create_connection method from python-kazoo module. The last release of kazoo slightly changed the way how create_connection method is called.
- Fix Patroni crash when Consul cluster loses the leader (Alexander)
The crash was happening due to incorrect implementation of touch_member method, it should return boolean and not raise any exceptions.
Version 1.5.1
-------------
This version implements support of permanent replication slots, adds support of pgBackRest and fixes number of bugs.
**New features**
- Permanent replication slots (Alexander Kukushkin)
Permanent replication slots are preserved on failover/switchover, that is, Patroni on the new primary will create configured replication slots right after doing promote. Slots could be configured with the help of `patronictl edit-config`. The initial configuration could be also done in the :ref:`bootstrap.dcs <settings>`.
- Add pgbackrest support (Yogesh Sharma)
pgBackrest can restore in existing $PGDATA folder, this allows speedy restore as files which have not changed since last backup are skipped, to support this feature new parameter `keep_data` has been introduced. See :ref:`replica creation method <custom_replica_creation>` section for additional examples.
**Bug fixes**
- A few bugfixes in the "standby cluster" workflow (Alexander)
Please see https://github.com/zalando/patroni/pull/823 for more details.
- Fix REST API health check when cluster management is paused and DCS is not accessible (Alexander)
Regression was introduced in https://github.com/zalando/patroni/commit/90cf930036a9d5249265af15d2b787ec7517cf57
Version 1.5.0
-------------
This version enables Patroni HA cluster to operate in a standby mode, introduces experimental support for running on Windows, and provides a new configuration parameter to register PostgreSQL service in Consul.
**New features**
- Standby cluster (Dmitry Dolgov)
One or more Patroni nodes can form a standby cluster that runs alongside the primary one (i.e. in another datacenter) and consists of standby nodes that replicate from the master in the primary cluster. All PostgreSQL nodes in the standby cluster are replicas; one of those replicas elects itself to replicate directly from the remote master, while the others replicate from it in a cascading manner. More detailed description of this feature and some configuration examples can be found at :ref:`here <standby_cluster>`.
- Register Services in Consul (Pavel Kirillov, Alexander Kukushkin)
If `register_service` parameter in the consul :ref:`configuration <consul_settings>` is enabled, the node will register a service with the name `scope` and the tag `master`, `replica` or `standby-leader`.
- Experimental Windows support (Pavel Golub)
From now on it is possible to run Patroni on Windows, although Windows support is brand-new and hasn't received as much real-world testing as its Linux counterpart. We welcome your feedback!
**Improvements in patronictl**
- Add patronictl -k/--insecure flag and support for restapi cert (Wilfried Roset)
In the past if the REST API was protected by the self-signed certificates `patronictl` would fail to verify them. There was no way to disable that verification. It is now possible to configure `patronictl` to skip the certificate verification altogether or provide CA and client certificates in the :ref:`ctl: <patronictl_settings>` section of configuration.
- Exclude members with nofailover tag from patronictl switchover/failover output (Alexander Anikin)
Previously, those members were incorrectly proposed as candidates when performing interactive switchover or failover via patronictl.
**Stability improvements**
- Avoid parsing non-key-value output lines of pg_controldata (Alexander Anikin)
Under certain circuimstances pg_controldata outputs lines without a colon character. That would trigger an error in Patroni code that parsed pg_controldata output, hiding the actual problem; often such lines are emitted in a warning shown by pg_controldata before the regular output, i.e. when the binary major version does not match the one of the PostgreSQL data directory.
- Add member name to the error message during the leader election (Jan Mussler)
During the leader election, Patroni connects to all known members of the cluster and requests their status. Such status is written to the Patroni log and includes the name of the member. Previously, if the member was not accessible, the error message did not indicate its name, containing only the URL.
- Immediately reserve the WAL position upon creation of the replication slot (Alexander Kukushkin)
Starting from 9.6, `pg_create_physical_replication_slot` function provides an additional boolean parameter `immediately_reserve`. When it is set to `false`, which is also the default, the slot doesn't reserve the WAL position until it receives the first client connection, potentially losing some segments required by the client in a time window between the slot creation and the intiial client connection.
- Fix bug in strict synchronous replication (Alexander Kukushkin)
When running with `synchronous_mode_strict: true`, in some cases Patroni puts `*` into the `synchronous_standby_names`, changing the sync state for most of the replication connections to `potential`. Previously, Patroni couldn't pick a synchronous candidate under such curcuimstances, as it only considered those with the state `async`.
Version 1.4.6
-------------
@@ -413,8 +251,8 @@ In addition to using Endpoints, Patroni supports ConfigMaps. You can find more i
- Minimize the amount of SELECT's issued by Patroni on every loop of HA cylce (Alexander Kukushkin)
On every iteration of HA loop Patroni needs to know recovery status and absolute wal position. From now on Patroni will run only single SELECT to get this information instead of two on the replica and three on the master.
On every iteration of HA loop Patroni needs to know recovery status and absolute wal position. From now on Patroni will run only single SELECT to get this information instead of two on the replica and three on the master.
- Remove leader key on shutdown only when we have the lock (Ants)
Unconditional removal was generating unnecessary and missleading exceptions.
@@ -424,7 +262,7 @@ In addition to using Endpoints, Patroni supports ConfigMaps. You can find more i
- Add version command to patronictl (Ants)
It will show the version of installed Patroni and versions of running Patroni instances (if the cluster name is specified).
- Make optional specifying cluster_name argument for some of patronictl commands (Alexander, Ants)
It will work if patronictl is using usual Patroni configuration file with the ``scope`` defined.
@@ -432,11 +270,11 @@ In addition to using Endpoints, Patroni supports ConfigMaps. You can find more i
- Show information about scheduled switchover and maintenance mode (Alexander)
Before that it was possible to get this information only from Patroni logs or directly from DCS.
- Improve ``patronictl reinit`` (Alexander)
Sometimes ``patronictl reinit`` refused to proceed when Patroni was busy with other actions, namely trying to start postgres. `patronictl` didn't provide any commands to cancel such long running actions and the only (dangerous) workarond was removing a data directory manually. The new implementation of `reinit` forcefully cancells other long-running actions before proceeding with reinit.
- Implement ``--wait`` flag in ``patronictl pause`` and ``patronictl resume`` (Alexander)
It will make ``patronictl`` wait until the requested action is acknowledged by all nodes in the cluster.
@@ -483,7 +321,7 @@ Version 1.3.6
After a crash that doesn't clean up postmaster.pid there could be a new process with the same pid, resulting in a false positive for is_running(), which will lead to all kinds of bad behavior.
- Shutdown postgresql before bootstrap when we lost data directory (ainlolcat)
When data directory on the master is forcefully removed, postgres process can still stay alive for some time and prevent the replica created in place of that former master from starting or replicating.
The fix makes Patroni cache the postmaster pid and its start time and let it terminate the old postmaster in case it is still running after the corresponding data directory has been removed.
@@ -553,7 +391,7 @@ Version 1.3.4
- Pass the consul token as a header (Andrew Colin Kissa)
Headers are now the prefered way to pass the token to the consul `API <https://www.consul.io/api/index.html#authentication>`__.
- Advanced configuration for Consul (Alexander Kukushkin)
-36
View File
@@ -65,19 +65,6 @@ for outdated backup files. Some people prefer other backup solutions, such as ``
others, or simply roll their own scripts. In order to accommodate all those use-cases Patroni supports running custom
scripts to clone a new replica. Those are configured in the ``postgresql`` configuration block:
.. code:: YAML
postgresql:
create_replica_methods:
- <method name>
<method name>:
command: <command name>
keep_data: True
no_params: True
no_master: 1
example: wal_e
.. code:: YAML
postgresql:
@@ -92,21 +79,6 @@ example: wal_e
basebackup:
max-rate: '100M'
example: pgbackrest
.. code:: YAML
postgresql:
create_replica_methods:
- pgbackrest
- basebackup
pgbackrest:
command: /usr/bin/pgbackrest --stanza=<scope> --delta restore
keep_data: True
no_params: True
basebackup:
max-rate: '100M'
The ``create_replica_methods`` defines available replica creation methods and the order of executing them. Patroni will
stop on the first one that returns 0. Each method should define a separate section in the configuration file, listing the command
@@ -127,10 +99,6 @@ A special ``no_master`` parameter, if defined, allows Patroni to call the replic
running master or replicas. In that case, an empty string will be passed in a connection string. This is useful for
restoring the formerly running cluster from the binary backup.
A special ``keep_data`` parameter, if defined, will instuct Patroni to not clean PGDATA folder before calling restore.
A special ``no_params`` parameter, if defined, restricts passing parameters to custom command.
A ``basebackup`` method is a special case: it will be used if
``create_replica_methods`` is empty, although it is possible
to list it explicitly among the ``create_replica_methods`` methods. This method initializes a new replica with the
@@ -162,8 +130,6 @@ and
If all replica creation methods fail, Patroni will try again all methods in order during the next event loop cycle.
.. _standby_cluster:
Standby cluster
---------------
@@ -195,5 +161,3 @@ in a patroni configuration:
Note, that these options will be applied only once during cluster bootstrap,
and the only way to change them afterwards is through DCS.
If you use replication slots on the standby cluster, you must also create the corresponding replication slot on the primary cluster. It will not be done automatically by the standby cluster implementation. You can use Patroni's permenant replication slots feature on the primary cluster to maintain a replication slot with the same name as ``primary_slot_name``, or its default value if ``primary_slot_name`` is not provided.
-2
View File
@@ -13,8 +13,6 @@ In asynchronous mode the cluster is allowed to lose some committed transactions
The amount of transactions that can be lost is controlled via ``maximum_lag_on_failover`` parameter. Because the primary transaction log position is not sampled in real time, in reality the amount of lost data on failover is worst case bounded by ``maximum_lag_on_failover`` bytes of transaction log plus the amount that is written in the last ``ttl`` seconds (``loop_wait``/2 seconds in the average case). However typical steady state replication delay is well under a second.
By default, when running leader elections, Patroni does not take into account the current timeline of replicas, what in some cases could be undesirable behavior. You can prevent the node not having the same timeline as a former master become the new leader by changing the value of ``check_timeline`` parameter to ``true``.
PostgreSQL synchronous replication
----------------------------------
+2 -2
View File
@@ -1,11 +1,11 @@
### confd
`confd` directory contains haproxy and pgbouncer template files for the [confd](https://github.com/kelseyhightower/confd) -- lightweight configuration management tool
`confd` directory contains haproxy template files for the [confd](https://github.com/kelseyhightower/confd) -- lightweight configuration management tool
You need to copy content of `confd` directory into /etcd/confd and run confd service:
```bash
$ confd -prefix=/service/$PATRONI_SCOPE -backend etcd -node $PATRONI_ETCD_URL -interval=10
```
It will periodically update haproxy.cfg and pgbouncer.ini with the actual list of Patroni nodes from `etcd` and "reload" haproxy and pgbouncer.ini when it is necessary.
It will periodically update haproxy.cfg with the actual list of Patroni nodes from `etcd` and "reload" haproxy when it is necessary.
### startup-scripts
-12
View File
@@ -1,12 +0,0 @@
[template]
prefix = "/service/batman"
owner = "postgres"
mode = "0644"
src = "pgbouncer.tmpl"
dest = "/etc/pgbouncer/pgbouncer.ini"
reload_cmd = "systemctl reload pgbouncer"
keys = [
"/members/","/leader"
]
-17
View File
@@ -1,17 +0,0 @@
[databases]
{{with get "/leader"}}{{$leader := .Value}}{{$leadkey := printf "/members/%s" $leader}}{{with get $leadkey}}{{$data := json .Value}}{{$hostport := base (replace (index (split $data.conn_url "/") 2) "@" "/" -1)}}{{ $host := base (index (split $hostport ":") 0)}}{{ $port := base (index (split $hostport ":") 1)}}* = host={{ $host }} port={{ $port }} pool_size=10{{end}}{{end}}
[pgbouncer]
logfile = /var/log/postgresql/pgbouncer.log
pidfile = /var/run/postgresql/pgbouncer.pid
listen_addr = *
listen_port = 6432
unix_socket_dir = /var/run/postgresql
auth_type = trust
auth_file = /etc/pgbouncer/userlist.txt
auth_hba_file = /etc/pgbouncer/pg_hba.txt
admin_users = pgbouncer
stats_users = pgbouncer
pool_mode = session
max_client_conn = 100
default_pool_size = 20
+1 -7
View File
@@ -18,14 +18,8 @@ WorkingDirectory=~
# Where to send early-startup messages from the server
# This is normally controlled by the global default set by systemd
#StandardOutput=syslog
# StandardOutput=syslog
# Pre-commands to start watchdog device
# Uncomment if watchdog is part of your patroni setup
#ExecStartPre=-/usr/bin/sudo /sbin/modprobe softdog
#ExecStartPre=-/usr/bin/sudo /bin/chown postgres /dev/watchdog
# Start the patroni process
ExecStart=/bin/patroni /etc/patroni.yml
# Send HUP to reload from patroni.yml
-5
View File
@@ -1,5 +0,0 @@
#!/bin/bash
[[ "$3" == "master" ]] || exit
PGPASSWORD=zalando psql -h localhost -U postgres -p $1 -w -tAc "SELECT slot_name FROM pg_replication_slots WHERE slot_type = 'logical'" >> data/postgres0/label
+9 -19
View File
@@ -117,24 +117,12 @@ class PatroniController(AbstractController):
except IOError:
return None
@staticmethod
def recursive_update(dst, src):
for k, v in src.items():
if k in dst and isinstance(dst[k], dict):
PatroniController.recursive_update(dst[k], v)
else:
dst[k] = v
def update_config(self, custom_config):
def add_tag_to_config(self, tag, value):
with open(self._config) as r:
config = yaml.safe_load(r)
self.recursive_update(config, custom_config)
config['tags']['tag'] = value
with open(self._config, 'w') as w:
yaml.safe_dump(config, w, default_flow_style=False)
self._scope = config.get('scope', 'batman')
def add_tag_to_config(self, tag, value):
self.update_config({'tags': {tag: value}})
def _start(self):
if self.watchdog:
@@ -186,10 +174,13 @@ class PatroniController(AbstractController):
config['bootstrap']['initdb'].extend([{'auth': 'md5'}, {'auth-host': 'md5'}])
if custom_config is not None:
self.recursive_update(config, custom_config)
if config['postgresql'].get('callbacks', {}).get('on_role_change'):
config['postgresql']['callbacks']['on_role_change'] += ' ' + str(self.__PORT)
def recursive_update(dst, src):
for k, v in src.items():
if k in dst and isinstance(dst[k], dict):
recursive_update(dst[k], v)
else:
dst[k] = v
recursive_update(config, custom_config)
with open(patroni_config_path, 'w') as f:
yaml.safe_dump(config, f, default_flow_style=False)
@@ -372,7 +363,6 @@ class ConsulController(AbstractDcsController):
def __init__(self, context):
super(ConsulController, self).__init__(context)
os.environ['PATRONI_CONSUL_HOST'] = 'localhost:8500'
os.environ['PATRONI_CONSUL_REGISTER_SERVICE'] = 'on'
self._client = consul.Consul()
self._config_file = None
+1 -1
View File
@@ -43,7 +43,7 @@ Scenario: check dynamic configuration change via DCS
And I receive a response loop_wait 2
When I issue a GET request to http://127.0.0.1:8008/patroni
Then I receive a response code 200
And I receive a response tags {'new_tag': 'new_value'}
And I receive a response tags {'tag': 'new_value'}
Scenario: check API requests for the primary-replica pair in the pause mode
Given I run patronictl.py pause batman
+4 -22
View File
@@ -1,26 +1,9 @@
Feature: standby cluster
Scenario: check permanent logical slots are preserved on failover/switchover
Given I start postgres1
Then postgres1 is a leader after 10 seconds
And I sleep for 2 seconds
When I issue a PATCH request to http://127.0.0.1:8009/config with {"loop_wait": 2, "slots": {"pm_1": {"type": "physical"}}, "postgresql": {"parameters": {"wal_level": "logical"}}}
Then I receive a response code 200
And Response on GET http://127.0.0.1:8009/config contains slots after 10 seconds
And I sleep for 2 seconds
When I issue a PATCH request to http://127.0.0.1:8009/config with {"slots": {"test_logical": {"type": "logical", "database": "postgres", "plugin": "test_decoding"}}}
Then I receive a response code 200
When I start postgres0 with callback configured
Then "members/postgres0" key in DCS has state=running after 10 seconds
And replication works from postgres1 to postgres0 after 15 seconds
When I shut down postgres1
Then postgres0 is a leader after 10 seconds
And I sleep for 2 seconds
When I issue a GET request to http://127.0.0.1:8008/
Then I receive a response code 200
And there is a label with "test_logical" in postgres0 data directory
Scenario: check replication of a single table in a standby cluster
Given I start postgres1 in a standby cluster batman1 as a clone of postgres0
Given I start postgres0 without slots sync
And I create a replication slot postgres1 on postgres0
And I start postgres1 in a standby cluster batman1 as a clone of postgres0
Then postgres1 is a leader of batman1 after 10 seconds
When I add the table foo to postgres0
Then table foo is present on postgres1 after 20 seconds
@@ -30,5 +13,4 @@ Feature: standby cluster
Scenario: check failover
When I kill postgres1
And I kill postmaster on postgres1
Then postgres2 is replicating from postgres0 after 20 seconds
Then postgres2 is replicating from postgres0 after 20 seconds
+22 -11
View File
@@ -1,4 +1,3 @@
import os
import time
from behave import step
@@ -9,13 +8,19 @@ SELECT * FROM pg_catalog.pg_stat_replication
WHERE application_name = '{0}'
"""
create_replication_slot_query = """
SELECT pg_create_physical_replication_slot('{0}')
"""
@step('I start {name:w} with callback configured')
def start_patroni_with_callbacks(context, name):
@step('I start {name:w} without slots sync')
def start_patroni_without_slots_sync(context, name):
return context.pctl.start(name, custom_config={
"postgresql": {
"callbacks": {
"on_role_change": "features/callback.sh"
"bootstrap": {
"dcs": {
"postgresql": {
"use_slots": False
}
}
}
})
@@ -30,22 +35,19 @@ def start_patroni(context, name, cluster_name):
@step('I start {name:w} in a standby cluster {cluster_name:w} as a clone of {name2:w}')
def start_patroni_stanby_cluster(context, name, cluster_name, name2):
# we need to remove patroni.dynamic.json in order to "bootstrap" standby cluster with existing PGDATA
os.unlink(os.path.join(context.pctl._processes[name]._data_dir, 'patroni.dynamic.json'))
port = context.pctl._processes[name2]._connkwargs.get('port')
context.pctl._processes[name].update_config({
return context.pctl.start(name, custom_config={
"scope": cluster_name,
"bootstrap": {
"dcs": {
"standby_cluster": {
"host": "localhost",
"port": port,
"primary_slot_name": "pm_1",
"primary_slot_name": "postgres1",
}
}
}
})
return context.pctl.start(name)
@step('{pg_name1:w} is replicating from {pg_name2:w} after {timeout:d} seconds')
@@ -65,3 +67,12 @@ def check_replication_status(context, pg_name1, pg_name2, timeout):
time.sleep(1)
return False
@step('I create a replication slot {slot_name:w} on {pg_name:w}')
def create_replication_slot(context, slot_name, pg_name):
return context.pctl.query(
pg_name,
create_replication_slot_query.format(slot_name),
fail_ok=True
)
+10 -12
View File
@@ -1,24 +1,22 @@
FROM postgres:11
FROM postgres:10
MAINTAINER Alexander Kukushkin <[email protected]>
RUN export DEBIAN_FRONTEND=noninteractive \
&& echo 'APT::Install-Recommends "0";\nAPT::Install-Suggests "0";' > /etc/apt/apt.conf.d/01norecommend \
&& apt-get update -y \
&& apt-get upgrade -y \
&& apt-cache depends patroni | sed -n -e 's/.* Depends: \(python3-.\+\)$/\1/p' \
| grep -Ev '^python3-(sphinx|etcd|consul|kazoo|kubernetes)' \
| xargs apt-get install -y vim-tiny curl jq locales git python3-pip python3-wheel \
&& apt-get install -s patroni | sed -n -e '/^Inst patroni /d' -e 's/^Inst \([^ ]\+\) .*$/\1/p' \
| xargs apt-get install -y curl jq locales git python3-pip python3-wheel \
## Make sure we have a en_US.UTF-8 locale available
&& localedef -i en_US -c -f UTF-8 -A /usr/share/locale/locale.alias en_US.UTF-8 \
&& pip3 install setuptools \
&& pip3 install 'git+https://github.com/zalando/patroni.git#egg=patroni[kubernetes]' \
&& PGHOME=/home/postgres \
&& mkdir -p $PGHOME \
&& chown postgres $PGHOME \
&& sed -i "s|/var/lib/postgresql.*|$PGHOME:/bin/bash|" /etc/passwd \
# Set permissions for OpenShift
&& chmod 775 $PGHOME \
&& chmod 664 /etc/passwd \
&& mkdir -p /home/postgres \
&& chown postgres:postgres /home/postgres \
# Clean up
&& apt-get remove -y git python3-pip python3-wheel \
&& apt-get autoremove -y \
@@ -28,7 +26,7 @@ RUN export DEBIAN_FRONTEND=noninteractive \
ADD entrypoint.sh /
EXPOSE 5432 8008
ENV LC_ALL=en_US.UTF-8 LANG=en_US.UTF-8 EDITOR=/usr/bin/editor
ENV LC_ALL=en_US.UTF-8 LANG=en_US.UTF-8
USER postgres
WORKDIR /home/postgres
CMD ["/bin/bash", "/entrypoint.sh"]
+189
View File
@@ -0,0 +1,189 @@
#!/bin/bash
# here you can defind your own alias to minikube
lkubectl='kubectl --context=local'
shopt -s expand_aliases
alias lkubectl="$lkubectl"
labels=application=patroni,cluster-name=patronidemo
# clean up objects from previous attempts
lkubectl delete statefulset,pods,service,endpoints -l $labels
# clear old panes
for pane in {3..1}; do
tmux send-keys -t 0.$pane C-c
tmux kill-pane -t 0.$pane
done
tmux send-keys -t 0.0 C-c
tmux send-keys -t 0.0 C-l
# split window into 3 panes
tmux split-window -h -t 0.0
tmux split-window -v -t 0.0
tmux split-window -v -t 0.1
tmux send-keys -t 0.3 "export PGPASSWORD=zalando" Enter
tmux send-keys -t 0.3 "export PGCONNECT_TIMEOUT=1" Enter
for pane in {0..3}; do
tmux send-keys -t 0.$pane "shopt -s expand_aliases" Enter
tmux send-keys -t 0.$pane "alias lkubectl='$lkubectl'" Enter
tmux send-keys -t 0.$pane C-l
done
# run
tmux resize-pane -Z -t 0.3
tmux send-keys -t 0.3 "lkubectl apply -f patroni_k8s.yaml" Enter
tmux send-keys -t 0.3 "lkubectl get pods -w -l $labels -L role -o wide" Enter
read
tmux send-keys -t 0.3 C-c
#tmux send-keys -t 0.3 C-l
#
#tmux send-keys -t 0.3 "lkubectl get pods -l $labels -L role -o wide" Enter
#
#read
# tail logs from patroni pods
for pane in {0..2}; do
tmux send-keys -t 0.$pane "lkubectl logs patronidemo-$pane -f" Enter
done
tmux resize-pane -Z -t 0.3
read
tmux resize-pane -Z -t 0.0
read
tmux resize-pane -Z -t 0.0
read
tmux resize-pane -Z -t 0.1
read
tmux resize-pane -Z -t 0.1
read
tmux resize-pane -Z -t 0.3
tmux send-keys -t 0.3 C-l
tmux send-keys -t 0.3 "lkubectl get pods -l $labels -L role -o wide" Enter
read
tmux send-keys -t 0.3 Enter
tmux send-keys -t 0.3 "lkubectl get services -l $labels" Enter
read
tmux send-keys -t 0.3 Enter
tmux send-keys -t 0.3 "lkubectl get endpoints -l $labels" Enter
read
tmux send-keys -t 0.3 C-l
tmux send-keys -t 0.3 "lkubectl get pods -l $labels -L role -o wide" Enter
tmux send-keys -t 0.3 Enter
tmux send-keys -t 0.3 "lkubectl get endpoints patronidemo -o yaml" Enter
read
tmux send-keys -t 0.3 C-l
tmux send-keys -t 0.3 "lkubectl get pods -l $labels -L role -o wide" Enter
tmux send-keys -t 0.3 Enter
tmux send-keys -t 0.3 "lkubectl describe service patronidemo-repl" Enter
read
tmux send-keys -t 0.3 C-l
tmux send-keys -t 0.3 "lkubectl get endpoints patronidemo-config -o yaml" Enter
read
MASTERIP=$(lkubectl get service patronidemo -o jsonpath='{.spec.clusterIP}')
tmux resize-pane -Z -t 0.3
tmux send-keys -t 0.3 "watch -n 1 \"psql -h $MASTERIP -U postgres -c \\\"select txid_current(), array_agg(application_name||'=>'|| sync_state) AS \\\\\\\"array_agg(application_name||'=>'||sync_state)\\\\\\\" from pg_stat_replication\\\"\"" Enter
read
# kill master node
lkubectl delete pod patronidemo-0
read
tmux resize-pane -Z -t 0.1
tmux send-keys -t 0.1 C-b PageUp
read
tmux send-keys -t 0.1 Enter
tmux resize-pane -Z -t 0.1
read
tmux resize-pane -Z -t 0.2
tmux send-keys -t 0.2 PageUp
read
tmux send-keys -t 0.2 Enter
tmux resize-pane -Z -t 0.2
read
tmux send-keys -t 0.0 C-c
tmux send-keys -t 0.0 "lkubectl logs patronidemo-0 -f" Enter
read
tmux resize-pane -Z -t 0.3
tmux send-keys -t 0.3 C-c
tmux send-keys -t 0.3 C-l
tmux send-keys -t 0.3 "lkubectl get pods -l $labels -L role -o wide" Enter
tmux send-keys -t 0.3 Enter
tmux send-keys -t 0.3 "lkubectl get endpoints -l $labels" Enter
read
tmux send-keys -t 0.3 C-l
tmux send-keys -t 0.3 "lkubectl get endpoints patronidemo -o yaml" Enter
read
tmux resize-pane -Z -t 0.3
tmux send-keys -t 0.3 C-c
tmux send-keys -t 0.3 C-l
tmux send-keys -t 0.3 "lkubectl exec -ti patronidemo-0 patronictl switchover patronidemo" Enter
read
tmux send-keys -t 0.3 Enter
read
tmux send-keys -t 0.3 "patronidemo-0" Enter
read
tmux send-keys -t 0.3 Enter
read
tmux send-keys -t 0.3 "y" Enter
read
tmux send-keys -t 0.3 "lkubectl exec -ti patronidemo-0 patronictl list patronidemo" Enter
+2 -18
View File
@@ -1,12 +1,5 @@
#!/bin/bash
if [[ $UID -ge 10000 ]]; then
GID=$(id -g)
sed -e "s/^postgres:x:[^:]*:[^:]*:/postgres:x:$UID:$GID:/" /etc/passwd > /tmp/passwd
cat /tmp/passwd > /etc/passwd
rm /tmp/passwd
fi
cat > /home/postgres/patroni.yml <<__EOF__
bootstrap:
dcs:
@@ -20,20 +13,11 @@ bootstrap:
- data-checksums
pg_hba:
- host all all 0.0.0.0/0 md5
- host replication ${PATRONI_REPLICATION_USERNAME} ${PATRONI_KUBERNETES_POD_IP}/16 md5
- host replication ${PATRONI_REPLICATION_USERNAME} ${PATRONI_KUBERNETES_POD_IP}/16 md5
restapi:
connect_address: '${PATRONI_KUBERNETES_POD_IP}:8008'
postgresql:
connect_address: '${PATRONI_KUBERNETES_POD_IP}:5432'
authentication:
superuser:
password: '${PATRONI_SUPERUSER_PASSWORD}'
replication:
password: '${PATRONI_REPLICATION_PASSWORD}'
__EOF__
unset PATRONI_SUPERUSER_PASSWORD PATRONI_REPLICATION_PASSWORD
export KUBERNETES_NAMESPACE=$PATRONI_KUBERNETES_NAMESPACE
export POD_NAME=$PATRONI_NAME
exec /usr/bin/python3 /usr/local/bin/patroni /home/postgres/patroni.yml
exec /usr/bin/python3 /usr/local/bin/patroni /home/postgres/patroni.yml
-49
View File
@@ -1,49 +0,0 @@
# Patroni OpenShift Configuration
Patroni can be run in OpenShift. Based on the kubernetes configuration, the Dockerfile and Entrypoint has been modified to support the dynamic UID/GID configuration that is applied in OpenShift. This can be run under the standard `restricted` SCC.
# Examples
## Create test project
```
oc new-project patroni-test
```
## Build the image
Note: Update the references when merged upstream.
Note: If deploying as a template for multiple users, the following commands should be performed in a shared namespace like `openshift`.
```
oc import-image postgres:10 --confirm -n openshift
oc new-build https://github.com/zalando/patroni --context-dir=kubernetes -n openshift
```
## Deploy the Image
Two configuration templates exist in [templates](templates) directory:
- Patroni Ephemeral
- Patroni Persistent
The only difference is whether or not the statefulset requests persistent storage.
## Create the Template
Install the template into the `openshift` namespace if this should be shared across projects:
```
oc create -f templates/template_patroni_ephemeral.yml -n openshift
```
Then, from your own project:
```
oc new-app patroni-pgsql-ephemeral
```
Once the pods are running, two configmaps should be available:
```
$ oc get configmap
NAME DATA AGE
patroniocp-config 0 1m
patroniocp-leader 0 1m
```
@@ -1,287 +0,0 @@
apiVersion: v1
kind: Template
metadata:
name: patroni-pgsql-ephemeral
annotations:
description: |-
Patroni Postgresql database cluster, without persistent storage.
WARNING: Any data stored will be lost upon pod destruction. Only use this template for testing.
iconClass: icon-postgresql
openshift.io/display-name: Patroni Postgresql (Ephemeral)
openshift.io/long-description: This template deploys a a patroni postgresql HA cluster without persistent storage.
tags: postgresql
objects:
- apiVersion: v1
kind: Service
metadata:
creationTimestamp: null
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
name: ${PATRONI_CLUSTER_NAME}
spec:
ports:
- port: 5432
protocol: TCP
targetPort: 5432
sessionAffinity: None
type: ClusterIP
status:
loadBalancer: {}
- apiVersion: v1
kind: Service
metadata:
creationTimestamp: null
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
name: ${PATRONI_MASTER_SERVICE_NAME}
spec:
ports:
- port: 5432
protocol: TCP
targetPort: 5432
selector:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
role: master
sessionAffinity: None
type: ClusterIP
status:
loadBalancer: {}
- apiVersion: v1
kind: Secret
metadata:
name: ${PATRONI_CLUSTER_NAME}
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
stringData:
superuser-password: ${PATRONI_SUPERUSER_PASSWORD}
replication-password: ${PATRONI_REPLICATION_PASSWORD}
- apiVersion: v1
kind: Service
metadata:
creationTimestamp: null
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
name: ${PATRONI_REPLICA_SERVICE_NAME}
spec:
ports:
- port: 5432
protocol: TCP
targetPort: 5432
selector:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
role: replica
sessionAffinity: None
type: ClusterIP
status:
loadBalancer: {}
- apiVersion: apps/v1
kind: StatefulSet
metadata:
creationTimestamp: null
generation: 3
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
name: ${APPLICATION_NAME}
spec:
podManagementPolicy: OrderedReady
replicas: 3
revisionHistoryLimit: 10
selector:
matchLabels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
serviceName: ${APPLICATION_NAME}
template:
metadata:
creationTimestamp: null
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
spec:
containers:
- env:
- name: PATRONI_KUBERNETES_POD_IP
valueFrom:
fieldRef:
apiVersion: v1
fieldPath: status.podIP
- name: PATRONI_KUBERNETES_NAMESPACE
valueFrom:
fieldRef:
apiVersion: v1
fieldPath: metadata.namespace
- name: PATRONI_KUBERNETES_LABELS
value: '{application: ${APPLICATION_NAME}, cluster-name: ${PATRONI_CLUSTER_NAME}}'
- name: PATRONI_SUPERUSER_USERNAME
value: ${PATRONI_SUPERUSER_USERNAME}
- name: PATRONI_SUPERUSER_PASSWORD
valueFrom:
secretKeyRef:
key: superuser-password
name: ${PATRONI_CLUSTER_NAME}
- name: PATRONI_REPLICATION_USERNAME
value: ${PATRONI_REPLICATION_USERNAME}
- name: PATRONI_REPLICATION_PASSWORD
valueFrom:
secretKeyRef:
key: replication-password
name: ${PATRONI_CLUSTER_NAME}
- name: PATRONI_SCOPE
value: ${PATRONI_CLUSTER_NAME}
- name: PATRONI_NAME
valueFrom:
fieldRef:
apiVersion: v1
fieldPath: metadata.name
- name: PATRONI_POSTGRESQL_DATA_DIR
value: /home/postgres/pgdata/pgroot/data
- name: PATRONI_POSTGRESQL_PGPASS
value: /tmp/pgpass
- name: PATRONI_POSTGRESQL_LISTEN
value: 0.0.0.0:5432
- name: PATRONI_RESTAPI_LISTEN
value: 0.0.0.0:8008
image: docker-registry.default.svc:5000/${NAMESPACE}/patroni:latest
imagePullPolicy: IfNotPresent
name: ${APPLICATION_NAME}
ports:
- containerPort: 8008
protocol: TCP
- containerPort: 5432
protocol: TCP
resources: {}
terminationMessagePath: /dev/termination-log
terminationMessagePolicy: File
volumeMounts:
- mountPath: /home/postgres/pgdata
name: pgdata
dnsPolicy: ClusterFirst
restartPolicy: Always
schedulerName: default-scheduler
securityContext: {}
serviceAccount: ${SERVICE_ACCOUNT}
serviceAccountName: ${SERVICE_ACCOUNT}
terminationGracePeriodSeconds: 0
volumes:
- name: pgdata
emptyDir: {}
updateStrategy:
type: OnDelete
- apiVersion: v1
kind: Endpoints
metadata:
name: ${APPLICATION_NAME}
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
subsets: []
- apiVersion: v1
kind: ServiceAccount
metadata:
name: ${SERVICE_ACCOUNT}
- apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: ${SERVICE_ACCOUNT}
rules:
- apiGroups:
- ""
resources:
- configmaps
verbs:
- create
- get
- list
- patch
- update
- watch
# delete is required only for 'patronictl remove'
- delete
- apiGroups:
- ""
resources:
- endpoints
verbs:
- get
- patch
- update
# the following three privileges are necessary only when using endpoints
- create
- list
- watch
# delete is required only for for 'patronictl remove'
- delete
- apiGroups:
- ""
resources:
- pods
verbs:
- get
- list
- patch
- update
- watch
- apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: ${SERVICE_ACCOUNT}
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: ${SERVICE_ACCOUNT}
subjects:
- kind: ServiceAccount
name: ${SERVICE_ACCOUNT}
parameters:
- description: The name of the application for labelling all artifacts.
displayName: Application Name
name: APPLICATION_NAME
value: patroni-ephemeral
- description: The name of the patroni-pgsql cluster.
displayName: Cluster Name
name: PATRONI_CLUSTER_NAME
value: patroni-ephemeral
- description: The name of the OpenShift Service exposed for the patroni-ephemeral-master container.
displayName: Master service name.
name: PATRONI_MASTER_SERVICE_NAME
value: patroni-ephemeral-master
- description: The name of the OpenShift Service exposed for the patroni-ephemeral-replica containers.
displayName: Replica service name.
name: PATRONI_REPLICA_SERVICE_NAME
value: patroni-ephemeral-replica
- description: Maximum amount of memory the container can use.
displayName: Memory Limit
name: MEMORY_LIMIT
value: 512Mi
- description: The OpenShift Namespace where the patroni and postgresql ImageStream resides.
displayName: ImageStream Namespace
name: NAMESPACE
value: openshift
- description: Username of the superuser account for initialization.
displayName: Superuser Username
name: PATRONI_SUPERUSER_USERNAME
value: postgres
- description: Password of the superuser account for initialization.
displayName: Superuser Passsword
name: PATRONI_SUPERUSER_PASSWORD
value: postgres
- description: Username of the replication account for initialization.
displayName: Replication Username
name: PATRONI_REPLICATION_USERNAME
value: postgres
- description: Password of the replication account for initialization.
displayName: Repication Passsword
name: PATRONI_REPLICATION_PASSWORD
value: postgres
- description: Service account name used for pods and rolebindings to form a cluster in the project.
displayName: Service Account
name: SERVICE_ACCOUNT
value: patroniocp
@@ -1,303 +0,0 @@
apiVersion: v1
kind: Template
metadata:
name: patroni-pgsql-persistent
annotations:
description: |-
Patroni Postgresql database cluster, with persistent storage.
WARNING: Any data stored will be lost upon pod destruction. Only use this template for testing.
iconClass: icon-postgresql
openshift.io/display-name: Patroni Postgresql (Persistent)
openshift.io/long-description: This template deploys a a patroni postgresql HA cluster without persistent storage.
tags: postgresql
objects:
- apiVersion: v1
kind: Service
metadata:
creationTimestamp: null
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
name: ${PATRONI_CLUSTER_NAME}
spec:
ports:
- port: 5432
protocol: TCP
targetPort: 5432
sessionAffinity: None
type: ClusterIP
status:
loadBalancer: {}
- apiVersion: v1
kind: Service
metadata:
creationTimestamp: null
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
name: ${PATRONI_MASTER_SERVICE_NAME}
spec:
ports:
- port: 5432
protocol: TCP
targetPort: 5432
selector:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
role: master
sessionAffinity: None
type: ClusterIP
status:
loadBalancer: {}
- apiVersion: v1
kind: Secret
metadata:
name: ${PATRONI_CLUSTER_NAME}
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
stringData:
superuser-password: ${PATRONI_SUPERUSER_PASSWORD}
replication-password: ${PATRONI_REPLICATION_PASSWORD}
- apiVersion: v1
kind: Service
metadata:
creationTimestamp: null
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
name: ${PATRONI_REPLICA_SERVICE_NAME}
spec:
ports:
- port: 5432
protocol: TCP
targetPort: 5432
selector:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
role: replica
sessionAffinity: None
type: ClusterIP
status:
loadBalancer: {}
- apiVersion: apps/v1
kind: StatefulSet
metadata:
creationTimestamp: null
generation: 3
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
name: ${APPLICATION_NAME}
spec:
podManagementPolicy: OrderedReady
replicas: 3
revisionHistoryLimit: 10
selector:
matchLabels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
serviceName: ${APPLICATION_NAME}
template:
metadata:
creationTimestamp: null
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
spec:
containers:
- env:
- name: PATRONI_KUBERNETES_POD_IP
valueFrom:
fieldRef:
apiVersion: v1
fieldPath: status.podIP
- name: PATRONI_KUBERNETES_NAMESPACE
valueFrom:
fieldRef:
apiVersion: v1
fieldPath: metadata.namespace
- name: PATRONI_KUBERNETES_LABELS
value: '{application: ${APPLICATION_NAME}, cluster-name: ${PATRONI_CLUSTER_NAME}}'
- name: PATRONI_SUPERUSER_USERNAME
value: ${PATRONI_SUPERUSER_USERNAME}
- name: PATRONI_SUPERUSER_PASSWORD
valueFrom:
secretKeyRef:
key: superuser-password
name: ${PATRONI_CLUSTER_NAME}
- name: PATRONI_REPLICATION_USERNAME
value: ${PATRONI_REPLICATION_USERNAME}
- name: PATRONI_REPLICATION_PASSWORD
valueFrom:
secretKeyRef:
key: replication-password
name: ${PATRONI_CLUSTER_NAME}
- name: PATRONI_SCOPE
value: ${PATRONI_CLUSTER_NAME}
- name: PATRONI_NAME
valueFrom:
fieldRef:
apiVersion: v1
fieldPath: metadata.name
- name: PATRONI_POSTGRESQL_DATA_DIR
value: /home/postgres/pgdata/pgroot/data
- name: PATRONI_POSTGRESQL_PGPASS
value: /tmp/pgpass
- name: PATRONI_POSTGRESQL_LISTEN
value: 0.0.0.0:5432
- name: PATRONI_RESTAPI_LISTEN
value: 0.0.0.0:8008
image: docker-registry.default.svc:5000/${NAMESPACE}/patroni:latest
imagePullPolicy: IfNotPresent
name: ${APPLICATION_NAME}
ports:
- containerPort: 8008
protocol: TCP
- containerPort: 5432
protocol: TCP
resources: {}
terminationMessagePath: /dev/termination-log
terminationMessagePolicy: File
volumeMounts:
- mountPath: /home/postgres/pgdata
name: ${APPLICATION_NAME}
dnsPolicy: ClusterFirst
restartPolicy: Always
schedulerName: default-scheduler
securityContext: {}
serviceAccount: ${SERVICE_ACCOUNT}
serviceAccountName: ${SERVICE_ACCOUNT}
terminationGracePeriodSeconds: 0
volumes:
- name: ${APPLICATION_NAME}
persistentVolumeClaim:
claimName: ${APPLICATION_NAME}
volumeClaimTemplates:
- metadata:
labels:
application: ${APPLICATION_NAME}
name: ${APPLICATION_NAME}
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: ${PVC_SIZE}
updateStrategy:
type: OnDelete
- apiVersion: v1
kind: Endpoints
metadata:
name: ${APPLICATION_NAME}
labels:
application: ${APPLICATION_NAME}
cluster-name: ${PATRONI_CLUSTER_NAME}
subsets: []
- apiVersion: v1
kind: ServiceAccount
metadata:
name: ${SERVICE_ACCOUNT}
- apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: ${SERVICE_ACCOUNT}
rules:
- apiGroups:
- ""
resources:
- configmaps
verbs:
- create
- get
- list
- patch
- update
- watch
# delete is required only for 'patronictl remove'
- delete
- apiGroups:
- ""
resources:
- endpoints
verbs:
- get
- patch
- update
# the following three privileges are necessary only when using endpoints
- create
- list
- watch
# delete is required only for for 'patronictl remove'
- delete
- apiGroups:
- ""
resources:
- pods
verbs:
- get
- list
- patch
- update
- watch
- apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: ${SERVICE_ACCOUNT}
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: ${SERVICE_ACCOUNT}
subjects:
- kind: ServiceAccount
name: ${SERVICE_ACCOUNT}
parameters:
- description: The name of the application for labelling all artifacts.
displayName: Application Name
name: APPLICATION_NAME
value: patroni-persistent
- description: The name of the patroni-pgsql cluster.
displayName: Cluster Name
name: PATRONI_CLUSTER_NAME
value: patroni-persistent
- description: The name of the OpenShift Service exposed for the patroni-persistent-master container.
displayName: Master service name.
name: PATRONI_MASTER_SERVICE_NAME
value: patroni-persistent-master
- description: The name of the OpenShift Service exposed for the patroni-persistent-replica containers.
displayName: Replica service name.
name: PATRONI_REPLICA_SERVICE_NAME
value: patroni-persistent-replica
- description: Maximum amount of memory the container can use.
displayName: Memory Limit
name: MEMORY_LIMIT
value: 512Mi
- description: The OpenShift Namespace where the patroni and postgresql ImageStream resides.
displayName: ImageStream Namespace
name: NAMESPACE
value: openshift
- description: Username of the superuser account for initialization.
displayName: Superuser Username
name: PATRONI_SUPERUSER_USERNAME
value: postgres
- description: Password of the superuser account for initialization.
displayName: Superuser Passsword
name: PATRONI_SUPERUSER_PASSWORD
value: postgres
- description: Username of the replication account for initialization.
displayName: Replication Username
name: PATRONI_REPLICATION_USERNAME
value: postgres
- description: Password of the replication account for initialization.
displayName: Repication Passsword
name: PATRONI_REPLICATION_PASSWORD
value: postgres
- description: Service account name used for pods and rolebindings to form a cluster in the project.
displayName: Service Account
name: SERVICE_ACCOUNT
value: patroni-persistent
- description: The size of the persistent volume to create.
displayName: Persistent Volume Size
name: PVC_SIZE
value: 5Gi
-43
View File
@@ -1,43 +0,0 @@
pipeline {
agent any
stages {
stage ('Deploy test pod'){
when {
expression {
openshift.withCluster() {
openshift.withProject() {
return !openshift.selector( "dc", "pgbench" ).exists()
}
}
}
}
steps {
script {
openshift.withCluster() {
openshift.withProject() {
def pgbench = openshift.newApp( "https://github.com/stewartshea/docker-pgbench/", "--name=pgbench", "-e PGPASSWORD=postgres", "-e PGUSER=postgres", "-e PGHOST=patroni-persistent-master", "-e PGDATABASE=postgres", "-e TEST_CLIENT_COUNT=20", "-e TEST_DURATION=120" )
def pgbenchdc = openshift.selector( "dc", "pgbench" )
timeout(5) {
pgbenchdc.rollout().status()
}
}
}
}
}
}
stage ('Run benchmark Test'){
steps {
sh '''
oc exec $(oc get pods -l app=pgbench | grep Running | awk '{print $1}') ./test.sh
'''
}
}
stage ('Clean up pgtest pod'){
steps {
sh '''
oc delete all -l app=pgbench
'''
}
}
}
}
@@ -1,2 +0,0 @@
# Jenkins Test
This pipeline test will create a separate deployment config for a pgbench pod and execute a test against the patroni cluster. This is a sample and should be customized.
+27
View File
@@ -1,3 +1,5 @@
# This manifest (and Docker image) works only with emptyDir volume and are designed for demo purpose.
# For production usage please take a look at https://github.com/zalando/spilo/tree/master/postgres-appliance
apiVersion: apps/v1beta1
kind: StatefulSet
metadata:
@@ -19,6 +21,9 @@ spec:
- name: *cluster_name
image: patroni # docker build -t patroni .
imagePullPolicy: IfNotPresent
# resources:
# limits:
# cpu: 100m
ports:
- containerPort: 8008
protocol: TCP
@@ -88,6 +93,7 @@ spec:
# storage: 5Gi
---
# Endpoint for leader elections
apiVersion: v1
kind: Endpoints
metadata:
@@ -98,6 +104,7 @@ metadata:
subsets: []
---
# master Service
apiVersion: v1
kind: Service
metadata:
@@ -112,6 +119,26 @@ spec:
targetPort: 5432
---
# replica Service
apiVersion: v1
kind: Service
metadata:
name: patronidemo-repl
labels:
application: patroni
cluster-name: &cluster_name patronidemo
spec:
type: ClusterIP
selector:
application: patroni
cluster-name: *cluster_name
role: replica
ports:
- port: 5432
targetPort: 5432
---
# Secrets (superuser password and replication passwor)
apiVersion: v1
kind: Secret
metadata:
+9 -14
View File
@@ -14,7 +14,6 @@ class Patroni(object):
from patroni.config import Config
from patroni.dcs import get_dcs
from patroni.ha import Ha
from patroni.log import PatroniLogger
from patroni.postgresql import Postgresql
from patroni.version import __version__
from patroni.watchdog import Watchdog
@@ -22,9 +21,7 @@ class Patroni(object):
self.setup_signal_handlers()
self.version = __version__
self.logger = PatroniLogger()
self.config = Config()
self.logger.reload_config(self.config.get('log', {}))
self.dcs = get_dcs(self.config)
self.watchdog = Watchdog(self.config)
self.load_dynamic_configuration()
@@ -52,7 +49,6 @@ class Patroni(object):
break
except DCSError:
logger.warning('Can not get cluster from dcs')
time.sleep(5)
def get_tags(self):
return {tag: value for tag, value in self.config.get('tags', {}).items()
@@ -69,7 +65,6 @@ class Patroni(object):
def reload_config(self):
try:
self.tags = self.get_tags()
self.logger.reload_config(self.config.get('log', {}))
self.dcs.reload_config(self.config)
self.watchdog.reload_config(self.config)
self.api.reload_config(self.config['restapi'])
@@ -130,8 +125,7 @@ class Patroni(object):
def setup_signal_handlers(self):
self._received_sighup = False
self._received_sigterm = False
if os.name != 'nt':
signal.signal(signal.SIGHUP, self.sighup_handler)
signal.signal(signal.SIGHUP, self.sighup_handler)
signal.signal(signal.SIGTERM, self.sigterm_handler)
def shutdown(self):
@@ -143,6 +137,12 @@ class Patroni(object):
def patroni_main():
logformat = os.environ.get('PATRONI_LOGFORMAT', '%(asctime)s %(levelname)s: %(message)s')
loglevel = os.environ.get('PATRONI_LOGLEVEL', 'INFO')
requests_loglevel = os.environ.get('PATRONI_REQUESTS_LOGLEVEL', 'WARNING')
logging.basicConfig(format=logformat, level=loglevel)
logging.getLogger('requests').setLevel(requests_loglevel)
patroni = Patroni()
try:
patroni.run()
@@ -150,13 +150,10 @@ def patroni_main():
pass
finally:
patroni.shutdown()
logging.shutdown()
def pg_ctl_start(args):
import subprocess
if os.name != 'nt':
os.setsid()
postmaster = subprocess.Popen(args)
print(postmaster.pid)
@@ -200,13 +197,11 @@ def main():
os.kill(pid, signo)
signal.signal(signal.SIGCHLD, sigchld_handler)
if os.name != 'nt':
signal.signal(signal.SIGHUP, passtochild)
signal.signal(signal.SIGQUIT, passtochild)
signal.signal(signal.SIGHUP, passtochild)
signal.signal(signal.SIGINT, passtochild)
signal.signal(signal.SIGUSR1, passtochild)
signal.signal(signal.SIGUSR2, passtochild)
signal.signal(signal.SIGABRT, passtochild)
signal.signal(signal.SIGQUIT, passtochild)
signal.signal(signal.SIGTERM, passtochild)
patroni = call_self(sys.argv[1:])
+15 -33
View File
@@ -1,12 +1,11 @@
import base64
import fcntl
import json
import logging
import psycopg2
import time
import traceback
import dateutil.parser
import datetime
import os
from patroni.postgresql import PostgresConnectionException, PostgresException, Postgresql
from patroni.utils import deep_compare, parse_bool, patch_config, Retry, \
@@ -88,14 +87,8 @@ class RestApiHandler(BaseHTTPRequestHandler):
replica_status_code = 200 if not patroni.noloadbalance and response.get('role') == 'replica' else 503
status_code = 503
if patroni.ha.is_standby_cluster() and ('standby_leader' in path or 'standby-leader' in path):
if 'master' in path:
status_code = 200 if patroni.ha.is_leader() else 503
elif 'master' in path or 'leader' in path or 'primary' in path:
# Round-robing across all masters in pause mode if DCS is not accessible
if not cluster and patroni.ha.is_paused():
status_code = 200 if response['role'] == 'master' else 503
else:
status_code = 200 if patroni.ha.is_leader() else 503
elif 'replica' in path:
status_code = replica_status_code
elif cluster: # dcs is available
@@ -413,20 +406,17 @@ class RestApiHandler(BaseHTTPRequestHandler):
raise RetryFailedError('')
stmt = ("WITH replication_info AS ("
"SELECT usename, application_name, client_addr, state, sync_state, sync_priority"
" FROM pg_catalog.pg_stat_replication) SELECT"
" pg_catalog.to_char(pg_catalog.pg_postmaster_start_time(), 'YYYY-MM-DD HH24:MI:SS.MS TZ'),"
" CASE WHEN pg_catalog.pg_is_in_recovery() THEN 0"
" ELSE ('x' || pg_catalog.substr(pg_catalog.pg_{0}file_name("
"pg_catalog.pg_current_{0}_{1}()), 1, 8))::bit(32)::int END,"
" CASE WHEN pg_catalog.pg_is_in_recovery() THEN 0"
" ELSE pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_current_{0}_{1}(), '0/0')::bigint END,"
" pg_catalog.pg_{0}_{1}_diff(COALESCE(pg_catalog.pg_last_{0}_receive_{1}(),"
" pg_catalog.pg_last_{0}_replay_{1}()), '0/0')::bigint,"
" pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_last_{0}_replay_{1}(), '0/0')::bigint,"
" pg_catalog.to_char(pg_catalog.pg_last_xact_replay_timestamp(), 'YYYY-MM-DD HH24:MI:SS.MS TZ'),"
" pg_catalog.pg_is_in_recovery() AND pg_catalog.pg_is_{0}_replay_paused(), "
"(SELECT pg_catalog.array_to_json(pg_catalog.array_agg("
"pg_catalog.row_to_json(ri))) FROM replication_info ri)")
" FROM pg_stat_replication) SELECT"
" to_char(pg_postmaster_start_time(), 'YYYY-MM-DD HH24:MI:SS.MS TZ'),"
" CASE WHEN pg_is_in_recovery() THEN 0"
" ELSE ('x' || SUBSTR(pg_{0}file_name(pg_current_{0}_{1}()), 1, 8))::bit(32)::int END,"
" CASE WHEN pg_is_in_recovery() THEN 0"
" ELSE pg_{0}_{1}_diff(pg_current_{0}_{1}(), '0/0')::bigint END,"
" pg_{0}_{1}_diff(COALESCE(pg_last_{0}_receive_{1}(), pg_last_{0}_replay_{1}()), '0/0')::bigint,"
" pg_{0}_{1}_diff(pg_last_{0}_replay_{1}(), '0/0')::bigint,"
" to_char(pg_last_xact_replay_timestamp(), 'YYYY-MM-DD HH24:MI:SS.MS TZ'),"
" pg_is_in_recovery() AND pg_is_{0}_replay_paused(),"
" (SELECT array_to_json(array_agg(row_to_json(ri))) FROM replication_info ri)")
row = self.query(stmt.format(self.server.patroni.postgresql.wal_name,
self.server.patroni.postgresql.lsn_name), retry=retry)[0]
@@ -489,10 +479,8 @@ class RestApiServer(ThreadingMixIn, HTTPServer, Thread):
@staticmethod
def _set_fd_cloexec(fd):
if os.name != 'nt':
import fcntl
flags = fcntl.fcntl(fd, fcntl.F_GETFD)
fcntl.fcntl(fd, fcntl.F_SETFD, flags | fcntl.FD_CLOEXEC)
flags = fcntl.fcntl(fd, fcntl.F_GETFD)
fcntl.fcntl(fd, fcntl.F_SETFD, flags | fcntl.FD_CLOEXEC)
def check_basic_auth_key(self, key):
return self.__auth_key == key
@@ -548,9 +536,3 @@ class RestApiServer(ThreadingMixIn, HTTPServer, Thread):
and self.__initialize(config):
self.start()
self.__set_config_parameters(config)
@staticmethod
def handle_error(request, client_address):
address, port = client_address
logger.warning('Exception happened during processing of request from {}:{}'.format(address, port))
logger.warning(traceback.format_exc())
+24 -39
View File
@@ -1,14 +1,14 @@
import json
import logging
import os
import shutil
import six
import sys
import tempfile
import yaml
from collections import defaultdict
from copy import deepcopy
from patroni.dcs import ClusterConfig
from patroni.dcs import ClusterConfig, is_standby_cluster
from patroni.postgresql import Postgresql
from patroni.utils import deep_compare, parse_bool, parse_int, patch_config
from requests.structures import CaseInsensitiveDict
@@ -43,7 +43,6 @@ class Config(object):
__DEFAULT_CONFIG = {
'ttl': 30, 'loop_wait': 10, 'retry_timeout': 10,
'maximum_lag_on_failover': 1048576,
'check_timeline': False,
'master_start_timeout': 300,
'synchronous_mode': False,
'synchronous_mode_strict': False,
@@ -99,6 +98,10 @@ class Config(object):
def dynamic_configuration(self):
return deepcopy(self._dynamic_configuration)
@property
def is_standby_cluster(self):
return is_standby_cluster(self._dynamic_configuration.get('standby_cluster'))
def check_mode(self, mode):
return bool(parse_bool(self._dynamic_configuration.get(mode)))
@@ -125,7 +128,7 @@ class Config(object):
with os.fdopen(fd, 'w') as f:
fd = None
json.dump(self.dynamic_configuration, f)
tmpfile = shutil.move(tmpfile, self._cache_file)
tmpfile = os.rename(tmpfile, self._cache_file)
self._cache_needs_saving = False
except Exception:
logger.exception('Exception when saving file: %s', self._cache_file)
@@ -195,9 +198,12 @@ class Config(object):
elif name not in ('connect_address', 'listen', 'data_dir', 'pgpass', 'authentication'):
config['postgresql'][name] = deepcopy(value)
elif name == 'standby_cluster':
for name, value in (value or {}).items():
if name in self.__DEFAULT_CONFIG['standby_cluster']:
config['standby_cluster'][name] = deepcopy(value)
allowed_keys = self.__DEFAULT_CONFIG['standby_cluster'].keys()
expected = {
k: v for k, v in (value or {}).items()
if (k in allowed_keys and isinstance(v, six.string_types))
}
config['standby_cluster'].update(expected)
elif name in config: # only variables present in __DEFAULT_CONFIG allowed to be overriden from DCS
if name in ('synchronous_mode', 'synchronous_mode_strict'):
config[name] = value
@@ -217,15 +223,6 @@ class Config(object):
if value:
ret[param] = value
def _fix_log_env(name, oldname):
value = _popenv(oldname)
name = Config.PATRONI_ENV_PREFIX + 'LOG_' + name.upper()
if value and name not in os.environ:
os.environ[name] = value
for name, oldname in (('level', 'loglevel'), ('format', 'logformat'), ('dateformat', 'log_datefmt')):
_fix_log_env(name, oldname)
def _set_section_values(section, params):
for param in params:
value = _popenv(section + '_' + param)
@@ -234,22 +231,6 @@ class Config(object):
_set_section_values('restapi', ['listen', 'connect_address', 'certfile', 'keyfile'])
_set_section_values('postgresql', ['listen', 'connect_address', 'data_dir', 'pgpass', 'bin_dir'])
_set_section_values('log', ['level', 'format', 'dateformat', 'dir', 'file_size', 'file_num', 'loggers'])
def _parse_dict(value):
if not value.strip().startswith('{'):
value = '{{{0}}}'.format(value)
try:
return yaml.safe_load(value)
except Exception:
logger.exception('Exception when parsing dict %s', value)
return None
value = ret.get('log', {}).pop('loggers', None)
if value:
value = _parse_dict(value)
if value:
ret['log']['loggers'] = value
def _get_auth(name):
ret = {}
@@ -288,18 +269,22 @@ class Config(object):
name, suffix = (param[8:].split('_', 1) + [''])[:2]
if name and suffix:
# PATRONI_(ETCD|CONSUL|ZOOKEEPER|EXHIBITOR|...)_(HOSTS?|PORT|..)
if suffix in ('HOST', 'HOSTS', 'PORT', 'SRV', 'URL', 'PROXY', 'CACERT', 'CERT', 'KEY', 'VERIFY',
'TOKEN', 'CHECKS', 'DC', 'REGISTER_SERVICE', 'SERVICE_CHECK_INTERVAL', 'NAMESPACE',
'CONTEXT', 'USE_ENDPOINTS', 'SCOPE_LABEL', 'ROLE_LABEL', 'POD_IP', 'PORTS', 'LABELS'):
if suffix in ('HOST', 'HOSTS', 'PORT', 'SRV', 'URL', 'PROXY', 'CACERT', 'CERT',
'KEY', 'VERIFY', 'TOKEN', 'CHECKS', 'DC', 'NAMESPACE', 'CONTEXT',
'USE_ENDPOINTS', 'SCOPE_LABEL', 'ROLE_LABEL', 'POD_IP', 'PORTS', 'LABELS'):
value = os.environ.pop(param)
if suffix == 'PORT':
value = value and parse_int(value)
elif suffix in ('HOSTS', 'PORTS', 'CHECKS'):
value = value and _parse_list(value)
elif suffix == 'LABELS':
value = _parse_dict(value)
elif suffix == 'REGISTER_SERVICE':
value = parse_bool(value)
if not value.strip().startswith('{'):
value = '{{{0}}}'.format(value)
try:
value = yaml.safe_load(value)
except Exception:
logger.exception('Exception when parsing dict %s', value)
value = None
if value:
ret[name.lower()][suffix.lower()] = value
# PATRONI_<username>_PASSWORD=<password>, PATRONI_<username>_OPTIONS=<option1,option2,...>
@@ -358,7 +343,7 @@ class Config(object):
'scope',
'retry_timeout',
'synchronous_mode',
'synchronous_mode_strict',
'maximum_lag_on_failover'
)
pg_config.update({p: config[p] for p in updated_fields if p in config})
+2 -2
View File
@@ -252,7 +252,7 @@ def get_cursor(cluster, connect_parameters, role='master', member=None):
if role == 'any':
return cursor
cursor.execute('SELECT pg_catalog.pg_is_in_recovery()')
cursor.execute('SELECT pg_is_in_recovery()')
in_recovery = cursor.fetchone()[0]
if in_recovery and role == 'replica' or not in_recovery and role == 'master':
@@ -392,7 +392,7 @@ def query_member(cluster, cursor, member, role, command, connect_parameters):
logging.debug(message)
return [[timestamp(0), message]], None
cursor.execute('SELECT pg_catalog.pg_is_in_recovery()')
cursor.execute('SELECT pg_is_in_recovery()')
in_recovery = cursor.fetchone()[0]
if in_recovery and role == 'master' or not in_recovery and role == 'replica':
+26 -119
View File
@@ -6,38 +6,19 @@ import json
import logging
import os
import pkgutil
import re
import six
import sys
from collections import defaultdict, namedtuple
from copy import deepcopy
from collections import namedtuple
from patroni.exceptions import PatroniException
from patroni.utils import parse_bool
from random import randint
from six.moves.urllib_parse import urlparse, urlunparse, parse_qsl
from threading import Event, Lock
slot_name_re = re.compile('^[a-z0-9_]{1,63}$')
logger = logging.getLogger(__name__)
def slot_name_from_member_name(member_name):
"""Translate member name to valid PostgreSQL slot name.
PostgreSQL replication slot names must be valid PostgreSQL names. This function maps the wider space of
member names to valid PostgreSQL names. Names are lowercased, dashes and periods common in hostnames
are replaced with underscores, other characters are encoded as their unicode codepoint. Name is truncated
to 64 characters. Multiple different member names may map to a single slot name."""
def replace_char(match):
c = match.group(0)
return '_' if c in '-.' else "u{:04d}".format(ord(c))
slot_name = re.sub('[^a-z0-9_]', replace_char, member_name.lower())
return slot_name[0:63]
def parse_connection_string(value):
"""Original Governor stores connection strings for each cluster members if a following format:
postgres://{username}:{password}@{connect_address}/postgres
@@ -207,12 +188,13 @@ class RemoteMember(Member):
'create_replica_methods',
'restore_command',
'archive_cleanup_command',
'recovery_min_apply_delay',
'no_replication_slot')
'recovery_min_apply_delay')
def __getattr__(self, name):
if name in RemoteMember.allowed_keys():
return self.data.get(name)
if name not in RemoteMember.allowed_keys():
return
return self.data.get(name)
class Leader(namedtuple('Leader', 'index,session,member')):
@@ -292,24 +274,14 @@ class ClusterConfig(namedtuple('ClusterConfig', 'index,data,modify_index')):
def from_node(index, data, modify_index=None):
"""
>>> ClusterConfig.from_node(1, '{') is None
False
True
"""
try:
data = json.loads(data)
except (TypeError, ValueError):
data = None
modify_index = 0
if not isinstance(data, dict):
data = {}
return ClusterConfig(index, data, index if modify_index is None else modify_index)
@property
def permanent_slots(self):
return isinstance(self.data, dict) and (
self.data.get('permanent_replication_slots') or
self.data.get('permanent_slots') or self.data.get('slots')
) or {}
return None
return ClusterConfig(index, data, modify_index or index)
class SyncState(namedtuple('SyncState', 'index,leader,sync_standby')):
@@ -368,7 +340,7 @@ class SyncState(namedtuple('SyncState', 'index,leader,sync_standby')):
return name is not None and name in (self.leader, self.sync_standby)
class TimelineHistory(namedtuple('TimelineHistory', 'index,value,lines')):
class TimelineHistory(namedtuple('TimelineHistory', 'index,lines')):
"""Object representing timeline history file"""
@staticmethod
@@ -384,7 +356,7 @@ class TimelineHistory(namedtuple('TimelineHistory', 'index,value,lines')):
lines = None
if not isinstance(lines, list):
lines = []
return TimelineHistory(index, value, lines)
return TimelineHistory(index, lines)
class Cluster(namedtuple('Cluster', 'initialize,config,leader,last_leader_operation,members,failover,sync,history')):
@@ -425,84 +397,8 @@ class Cluster(namedtuple('Cluster', 'initialize,config,leader,last_leader_operat
def is_synchronous_mode(self):
return self.check_mode('synchronous_mode')
def get_replication_slots(self, name, role):
# if the replicatefrom tag is set on the member - we should not create the replication slot for it on
# the current master, because that member would replicate from elsewhere. We still create the slot if
# the replicatefrom destination member is currently not a member of the cluster (fallback to the
# master), or if replicatefrom destination member happens to be the current master
if role in ('master', 'standby_leader'):
slot_members = [m.name for m in self.members if m.name != name and
(m.replicatefrom is None or m.replicatefrom == name or
not self.has_member(m.replicatefrom))]
permanent_slots = (self.config and self.config.permanent_slots or {}).copy()
else:
# only manage slots for replicas that replicate from this one, except for the leader among them
slot_members = [m.name for m in self.members if m.replicatefrom == name and m.name != self.leader.name]
permanent_slots = {}
slots = {slot_name_from_member_name(name): {'type': 'physical'} for name in slot_members}
if len(slots) < len(slot_members):
# Find which names are conflicting for a nicer error message
slot_conflicts = defaultdict(list)
for name in slot_members:
slot_conflicts[slot_name_from_member_name(name)].append(name)
logger.error("Following cluster members share a replication slot name: %s",
"; ".join("{} map to {}".format(", ".join(v), k)
for k, v in slot_conflicts.items() if len(v) > 1))
# "merge" replication slots for members with permanent_replication_slots
for name, value in permanent_slots.items():
if not slot_name_re.match(name):
logger.error("Invalid permanent replication slot name '%s'", name)
logger.error("Slot name may only contain lower case letters, numbers, and the underscore chars")
continue
if name in slots:
logger.error("Permanent replication slot {'%s': %s} is conflicting with" +
" physical replication slot for cluster member", name, value)
continue
value = deepcopy(value)
if not value:
value = {'type': 'physical'}
if isinstance(value, dict):
if 'type' not in value:
value['type'] = 'logical' if value.get('database') and value.get('plugin') else 'physical'
if value['type'] == 'physical' or value['type'] == 'logical' \
and value.get('database') and value.get('plugin'):
slots[name] = value
continue
logger.error("Bad value for slot '%s' in permanent_slots: %s", name, permanent_slots[name])
return slots
def has_permanent_logical_slots(self, name):
slots = self.get_replication_slots(name, 'master').values()
return any(v for v in slots if v.get("type") == "logical")
@property
def timeline(self):
"""
>>> Cluster(0, 0, 0, 0, 0, 0, 0, 0).timeline
0
>>> Cluster(0, 0, 0, 0, 0, 0, 0, TimelineHistory.from_node(1, '[]')).timeline
1
>>> Cluster(0, 0, 0, 0, 0, 0, 0, TimelineHistory.from_node(1, '[["a"]]')).timeline
0
"""
if self.history:
if self.history.lines:
try:
return int(self.history.lines[-1][0]) + 1
except Exception:
logger.error('Failed to parse cluster history from DCS: %s', self.history.lines)
elif self.history.value == '[]':
return 1
return 0
def is_standby_cluster(self):
return is_standby_cluster(self.config and self.config.data.get('standby_cluster'))
@six.add_metaclass(abc.ABCMeta)
@@ -524,7 +420,7 @@ class AbstractDCS(object):
i.e.: `zookeeper` for zookeeper, `etcd` for etcd, etc...
"""
self._name = config['name']
self._base_path = re.sub('/+', '/', '/'.join(['', config.get('namespace', 'service'), config['scope']]))
self._base_path = os.path.join('/', config.get('namespace', '/service/').strip('/'), config['scope'])
self._set_loop_wait(config.get('loop_wait', 10))
self._ctl = bool(config.get('patronictl', False))
@@ -640,7 +536,7 @@ class AbstractDCS(object):
You have to use CAS (Compare And Swap) operation in order to update leader key,
for example for etcd `prevValue` parameter must be used."""
def update_leader(self, last_operation, access_is_restricted=False):
def update_leader(self, last_operation):
"""Update leader key (or session) ttl and optime/leader
:param last_operation: absolute xlog location in bytes
@@ -757,3 +653,14 @@ class AbstractDCS(object):
self.event.wait(timeout)
return self.event.isSet()
def is_standby_cluster(config):
""" Check whether or not provided configuration describes a standby cluster.
Config can be both patroni config or cluster.config.data
"""
return isinstance(config, dict) and (
config.get('host') or
config.get('port') or
config.get('restore_command')
)
+13 -29
View File
@@ -27,14 +27,10 @@ class ConsulInternalError(ConsulException):
"""An internal Consul server error occurred"""
class InvalidSessionTTL(ConsulException):
class InvalidSessionTTL(ConsulInternalError):
"""Session TTL is too small or too big"""
class InvalidSession(ConsulException):
"""invalid session"""
class HTTPClient(object):
def __init__(self, host='127.0.0.1', port=8500, token=None, scheme='http', verify=True, cert=None, ca_cert=None):
@@ -76,8 +72,6 @@ class HTTPClient(object):
msg = '{0} {1}'.format(response.status, data)
if data.startswith('Invalid Session TTL'):
raise InvalidSessionTTL(msg)
elif data.startswith('invalid session'):
raise InvalidSession(msg)
else:
raise ConsulInternalError(msg)
return base.Response(response.status, response.headers, data)
@@ -294,7 +288,7 @@ class Consul(AbstractDCS):
nodes = {}
for node in results:
node['Value'] = (node['Value'] or b'').decode('utf-8')
nodes[os.path.relpath(node['Key'], path).replace('\\', '/')] = node
nodes[os.path.relpath(node['Key'], path)] = node
# get initialize flag
initialize = nodes.get(self._INITIALIZE)
@@ -344,15 +338,17 @@ class Consul(AbstractDCS):
logger.exception('get_cluster')
raise ConsulError('Consul is not responding properly')
@catch_consul_errors
def touch_member(self, data, ttl=None, permanent=False):
cluster = self.cluster
member = cluster and cluster.get_member(self._name, fallback_to_leader=False)
create_member = not permanent and self.refresh_session()
if member and (create_member or member.session != self._session):
self._client.kv.delete(self.member_path)
create_member = True
try:
self._client.kv.delete(self.member_path)
create_member = True
except Exception:
return False
if not create_member and member and deep_compare(data, member.data):
return True
@@ -363,9 +359,6 @@ class Consul(AbstractDCS):
if self._register_service:
self.update_service(not create_member and member and member.data or {}, data)
return True
except InvalidSession:
self._session = None
logger.error('Our session disappeared from Consul, can not "touch_member"')
except Exception:
logger.exception('touch_member')
return False
@@ -384,13 +377,12 @@ class Consul(AbstractDCS):
def _update_service(self, data):
service_name = self._service_name
role = data['role'].replace('_', '-')
role = data['role']
state = data['state']
api_parts = urlparse(data['api_url'])
api_parts = api_parts._replace(path='/{0}'.format(role))
conn_parts = urlparse(data['conn_url'])
check = base.Check.http(api_parts.geturl(), self._service_check_interval,
deregister='{0}s'.format(self._client.http.ttl * 10))
check = base.Check.http(api_parts.geturl(), self._service_check_interval, deregister=self._client.http.ttl * 10)
params = {
'service_id': '{0}/{1}'.format(self._scope, self._name),
'address': conn_parts.hostname,
@@ -402,7 +394,7 @@ class Consul(AbstractDCS):
if state == 'stopped':
return self.deregister_service(params['service_id'])
if role in ['master', 'replica', 'standby-leader']:
if role in ['master', 'replica']:
if state != 'running':
return
return self.register_service(service_name, **params)
@@ -424,21 +416,14 @@ class Consul(AbstractDCS):
return self._update_service(new_data)
@catch_consul_errors
def _do_attempt_to_acquire_leader(self, permanent):
try:
kwargs = {} if permanent else {'acquire': self._session}
return self.retry(self._client.kv.put, self.leader_path, self._name, **kwargs)
except InvalidSession:
self._session = None
logger.error('Our session disappeared from Consul. Will try to get a new one and retry attempt')
self.refresh_session()
return self.retry(self._client.kv.put, self.leader_path, self._name, acquire=self._session)
def _do_attempt_to_acquire_leader(self, kwargs):
return self.retry(self._client.kv.put, self.leader_path, self._name, **kwargs)
def attempt_to_acquire_leader(self, permanent=False):
if not self._session and not permanent:
self.refresh_session()
ret = self._do_attempt_to_acquire_leader(permanent)
ret = self._do_attempt_to_acquire_leader({} if permanent else {'acquire': self._session})
if not ret:
logger.info('Could not take out TTL lock')
@@ -516,5 +501,4 @@ class Consul(AbstractDCS):
try:
return super(Consul, self).watch(None, timeout)
finally:
self._last_session_refresh = 0
self.event.clear()
+1 -1
View File
@@ -448,7 +448,7 @@ class Etcd(AbstractDCS):
def _load_cluster(self):
try:
result = self.retry(self._client.read, self.client_path(''), recursive=True)
nodes = {os.path.relpath(node.key, result.key).replace('\\', '/'): node for node in result.leaves}
nodes = {os.path.relpath(node.key, result.key): node for node in result.leaves}
# get initialize flag
initialize = nodes.get(self._INITIALIZE)
+4 -10
View File
@@ -152,8 +152,7 @@ class Kubernetes(AbstractDCS):
# get global dynamic configuration
config = ClusterConfig.from_node(metadata and metadata.resource_version,
annotations.get(self._CONFIG) or '{}',
metadata.resource_version if self._CONFIG in annotations else 0)
annotations.get(self._CONFIG) or '{}')
# get timeline history
history = TimelineHistory.from_node(metadata and metadata.resource_version,
@@ -280,7 +279,7 @@ class Kubernetes(AbstractDCS):
def _update_leader(self):
"""Unused"""
def update_leader(self, last_operation, access_is_restricted=False):
def update_leader(self, last_operation):
now = datetime.datetime.now(tzutc).isoformat()
annotations = {self._LEADER: self._name, 'ttl': str(self._ttl), 'renewTime': now,
'acquireTime': self._leader_observed_record.get('acquireTime') or now,
@@ -288,11 +287,7 @@ class Kubernetes(AbstractDCS):
if last_operation:
annotations[self._OPTIME] = last_operation
subsets = self.__subsets
if subsets is not None and access_is_restricted:
subsets = []
ret = self.patch_or_create(self.leader_path, annotations, self._leader_resource_version, subsets=subsets)
ret = self.patch_or_create(self.leader_path, annotations, self._leader_resource_version, subsets=self.__subsets)
if ret:
self._leader_resource_version = ret.metadata.resource_version
return ret
@@ -312,8 +307,7 @@ class Kubernetes(AbstractDCS):
else:
annotations['acquireTime'] = self._leader_observed_record.get('acquireTime') or now
annotations['transitions'] = str(transitions)
subsets = [] if self.__subsets else None
ret = self.patch_or_create(self.leader_path, annotations, self._leader_resource_version, subsets=subsets)
ret = self.patch_or_create(self.leader_path, annotations, self._leader_resource_version, subsets=self.__subsets)
if ret:
self._leader_resource_version = ret.metadata.resource_version
else:
+1 -11
View File
@@ -1,6 +1,5 @@
import json
import logging
import select
import time
from kazoo.client import KazooClient, KazooState, KazooRetry
@@ -38,21 +37,12 @@ class PatroniSequentialThreadingHandler(SequentialThreadingHandler):
`connect_timeout` (negotiated session timeout) as the second element."""
args = list(args)
if len(args) == 0: # kazoo 2.6.0 slightly changed the way how it calls create_connection method
kwargs['timeout'] = max(self._connect_timeout, kwargs.get('timeout', self._connect_timeout*10)/10.0)
elif len(args) == 1:
if len(args) == 1:
args.append(self._connect_timeout)
else:
args[1] = max(self._connect_timeout, args[1]/10.0)
return super(PatroniSequentialThreadingHandler, self).create_connection(*args, **kwargs)
def select(self, *args, **kwargs):
"""Python3 raises `ValueError` if socket is closed, because fd == -1"""
try:
return super(PatroniSequentialThreadingHandler, self).select(*args, **kwargs)
except ValueError as e:
raise select.error(9, str(e))
class ZooKeeper(AbstractDCS):
+31 -88
View File
@@ -12,7 +12,7 @@ from collections import namedtuple
from multiprocessing.pool import ThreadPool
from patroni.async_executor import AsyncExecutor, CriticalTask
from patroni.exceptions import DCSError, PostgresConnectionException, PatroniException
from patroni.postgresql import ACTION_ON_START, ACTION_ON_ROLE_CHANGE
from patroni.postgresql import ACTION_ON_START
from patroni.utils import polling_loop, tzutc
from patroni.dcs import RemoteMember
from threading import RLock
@@ -20,28 +20,25 @@ from threading import RLock
logger = logging.getLogger(__name__)
class _MemberStatus(namedtuple('_MemberStatus', ['member', 'reachable', 'in_recovery', 'timeline',
'wal_position', 'tags', 'watchdog_failed'])):
class _MemberStatus(namedtuple('_MemberStatus', 'member,reachable,in_recovery,wal_position,tags,watchdog_failed')):
"""Node status distilled from API response:
member - dcs.Member object of the node
reachable - `!False` if the node is not reachable or is not responding with correct JSON
in_recovery - `!True` if pg_is_in_recovery() == true
timeline - timeline value from JSON
wal_position - maximum value of `replayed_location` or `received_location` from JSON
wal_position - value of `replayed_location` or `location` from JSON, dependin on its role.
tags - dictionary with values of different tags (i.e. nofailover)
watchdog_failed - indicates that watchdog is required by configuration but not available or failed
"""
@classmethod
def from_api_response(cls, member, json):
is_master = json['role'] == 'master'
timeline = json.get('timeline', 0)
wal = not is_master and max(json['xlog'].get('received_location', 0), json['xlog'].get('replayed_location', 0))
return cls(member, True, not is_master, timeline, wal, json.get('tags', {}), json.get('watchdog_failed', False))
return cls(member, True, not is_master, wal, json.get('tags', {}), json.get('watchdog_failed', False))
@classmethod
def unknown(cls, member):
return cls(member, False, None, 0, 0, {}, False)
return cls(member, False, None, 0, {}, False)
def failover_limitation(self):
"""Returns reason why this node can't promote or None if everything is ok."""
@@ -64,7 +61,6 @@ class Ha(object):
self.old_cluster = None
self._is_leader = False
self._is_leader_lock = RLock()
self._leader_access_is_restricted = False
self._was_paused = False
self._leader_timeline = None
self.recovering = False
@@ -95,33 +91,14 @@ class Ha(object):
def is_paused(self):
return self.check_mode('pause')
def check_timeline(self):
return self.check_mode('check_timeline')
def get_standby_cluster_config(self):
if self.cluster and self.cluster.config and self.cluster.config.modify_index:
config = self.cluster.config.data
else:
config = self.patroni.config.dynamic_configuration
return config.get('standby_cluster')
def is_standby_cluster(self):
config = self.get_standby_cluster_config()
# Check whether or not provided configuration describes a standby cluster
return isinstance(config, dict) and (config.get('host') or config.get('port') or config.get('restore_command'))
def is_leader(self):
with self._is_leader_lock:
return self._is_leader and not self._leader_access_is_restricted
return self._is_leader
def set_is_leader(self, value):
with self._is_leader_lock:
self._is_leader = value
def set_leader_access_is_restricted(self, value):
with self._is_leader_lock:
self._leader_access_is_restricted = value
def load_cluster_from_dcs(self):
cluster = self.dcs.get_cluster()
@@ -136,7 +113,6 @@ class Ha(object):
self._leader_timeline = None if cluster.is_unlocked() else cluster.leader.timeline
def acquire_lock(self):
self.set_leader_access_is_restricted(self.cluster.has_permanent_logical_slots(self.state_handler.name))
ret = self.dcs.attempt_to_acquire_leader()
self.set_is_leader(ret)
return ret
@@ -148,7 +124,7 @@ class Ha(object):
last_operation = self.state_handler.last_operation()
except Exception:
logger.exception('Exception when called state_handler.last_operation()')
ret = self.dcs.update_leader(last_operation, self._leader_access_is_restricted)
ret = self.dcs.update_leader(last_operation)
self.set_is_leader(ret)
if ret:
self.watchdog.keepalive()
@@ -175,10 +151,6 @@ class Ha(object):
'state': self.state_handler.state,
'role': self.state_handler.role
}
# following two lines are mainly necessary for consul, to avoid creation of master service
if data['role'] == 'master' and not self.is_leader():
data['role'] = 'promoted'
tags = self.get_effective_tags()
if tags:
data['tags'] = tags
@@ -230,7 +202,7 @@ class Ha(object):
self.state_handler.bootstrapping = True
self._post_bootstrap_task = CriticalTask()
if self.is_standby_cluster():
if self.patroni.config.is_standby_cluster:
self._async_executor.schedule('bootstrap_standby_leader')
self._async_executor.run_async(self.bootstrap_standby_leader)
return 'trying to bootstrap a new standby leader'
@@ -256,7 +228,8 @@ class Ha(object):
not a real master, but a 'standby leader', that will take base backup
from a remote master and start follow it.
"""
clone_source = self.get_remote_master()
patroni_config = self.patroni.config.dynamic_configuration
clone_source = self.get_remote_master(patroni_config)
msg = 'clone from remote master {0}'.format(clone_source.conn_url)
result = self.clone(clone_source, msg)
self._post_bootstrap_task.complete(result)
@@ -266,10 +239,9 @@ class Ha(object):
return result
def _handle_rewind(self):
leader = self.get_remote_master() if self.is_standby_cluster() else self.cluster.leader
if self.state_handler.rewind_needed_and_possible(leader):
self._async_executor.schedule('running pg_rewind from ' + leader.name)
self._async_executor.run_async(self.state_handler.rewind, (leader,))
if self.state_handler.rewind_needed_and_possible(self.cluster.leader):
self._async_executor.schedule('running pg_rewind from ' + self.cluster.leader.name)
self._async_executor.run_async(self.state_handler.rewind, (self.cluster.leader,))
return True
def recover(self):
@@ -301,24 +273,16 @@ class Ha(object):
self.load_cluster_from_dcs()
if self.is_standby_cluster() or not self.has_lock():
if self.has_lock():
msg = "starting as readonly because i had the session lock"
node_to_follow = None
else:
if not self.state_handler.rewind_executed:
self.state_handler.trigger_check_diverged_lsn()
if self._handle_rewind():
return self._async_executor.scheduled_action
if self.has_lock(): # in standby cluster
msg = "starting as a standby leader because i had the session lock"
node_to_follow = self._get_node_to_follow(self.cluster)
elif self.is_standby_cluster() and self.cluster.is_unlocked():
msg = "trying to follow a remote master because standby cluster is unhealthy"
node_to_follow = self.get_remote_master()
else:
msg = "starting as a secondary"
node_to_follow = self._get_node_to_follow(self.cluster)
elif self.has_lock():
msg = "starting as readonly because i had the session lock"
node_to_follow = None
msg = "starting as a secondary"
node_to_follow = self._get_node_to_follow(self.cluster)
self.recovering = True
@@ -331,8 +295,8 @@ class Ha(object):
# try to follow the node mentioned there, otherwise, follow the leader.
is_leader = self.cluster.leader and self.state_handler.name == self.cluster.leader.name
if self.is_standby_cluster() and (is_leader or self.cluster.is_unlocked()):
node_to_follow = self.get_remote_master()
if self.cluster.is_standby_cluster() and is_leader:
node_to_follow = self.get_remote_master(cluster.config.data)
elif self.patroni.replicatefrom and self.patroni.replicatefrom != self.state_handler.name:
node_to_follow = cluster.get_member(self.patroni.replicatefrom)
else:
@@ -515,10 +479,8 @@ class Ha(object):
return 'Postponing promotion because synchronous replication state was updated by somebody else'
self.state_handler.set_synchronous_standby('*' if self.is_synchronous_mode_strict() else None)
if self.state_handler.role != 'master':
self.set_leader_access_is_restricted(self.cluster.has_permanent_logical_slots(self.state_handler.name))
self._async_executor.schedule('promote')
self._async_executor.run_async(self.state_handler.promote,
args=(self.dcs.loop_wait, self._leader_access_is_restricted))
self._async_executor.run_async(self.state_handler.promote, args=(self.dcs.loop_wait,))
return promote_message
@staticmethod
@@ -549,23 +511,15 @@ class Ha(object):
:returns True when node is lagging
"""
lag = (self.cluster.last_leader_operation or 0) - wal_position
return lag > self.patroni.config.get('maximum_lag_on_failover', 0)
return lag > self.state_handler.config.get('maximum_lag_on_failover', 0)
def _is_healthiest_node(self, members, check_replication_lag=True):
"""This method tries to determine whether I am healthy enough to became a new leader candidate or not."""
_, my_wal_position = self.state_handler.timeline_wal_position()
if check_replication_lag and self.is_lagging(my_wal_position):
logger.info('My wal position exceeds maximum replication lag')
return False # Too far behind last reported wal position on master
if not self.is_standby_cluster() and self.check_timeline():
cluster_timeline = self.cluster.timeline
my_timeline = self.state_handler.replica_cached_timeline(cluster_timeline)
if my_timeline < cluster_timeline:
logger.info('My timeline %s is behind last known cluster timeline %s', my_timeline, cluster_timeline)
return False
# Prepare list of nodes to run check against
members = [m for m in members if m.name != self.state_handler.name and not m.nofailover and m.api_url]
@@ -576,13 +530,11 @@ class Ha(object):
logger.warning('Master (%s) is still alive', st.member.name)
return False
if my_wal_position < st.wal_position:
logger.info('Wal position of %s is ahead of my wal position', st.member.name)
return False
return True
def is_failover_possible(self, members):
ret = False
cluster_timeline = self.cluster.timeline
members = [m for m in members if m.name != self.state_handler.name and not m.nofailover and m.api_url]
if members:
for st in self.fetch_nodes_statuses(members):
@@ -591,9 +543,6 @@ class Ha(object):
logger.info('Member %s is %s', st.member.name, not_allowed_reason)
elif self.is_lagging(st.wal_position):
logger.info('Member %s exceeds maximum replication lag', st.member.name)
elif self.check_timeline() and (not st.timeline or st.timeline < cluster_timeline):
logger.info('Timeline %s of member %s is behind the cluster timeline %s',
st.timeline, st.member.name, cluster_timeline)
else:
ret = True
else:
@@ -831,7 +780,7 @@ class Ha(object):
self.dcs.manual_failover('', '')
self.load_cluster_from_dcs()
if self.is_standby_cluster():
if self.cluster.is_standby_cluster():
# standby leader disappeared, and this is a healthiest
# replica, so it should become a new standby leader.
# This imply that we need to start following a remote master
@@ -868,17 +817,12 @@ class Ha(object):
self.dcs.reset_cluster()
return 'removed leader lock because postgres is not running as master'
if self.state_handler.is_leader() and self._leader_access_is_restricted:
self.state_handler.sync_replication_slots(self.cluster)
self.state_handler.call_nowait(ACTION_ON_ROLE_CHANGE)
self.set_leader_access_is_restricted(False)
if self.update_lock(True):
msg = self.process_manual_failover_from_leader()
if msg is not None:
return msg
if self.is_standby_cluster():
if self.cluster.is_standby_cluster():
# in case of standby cluster we don't really need to
# enforce anything, since the leader is not a master.
# So just remind the role.
@@ -1018,8 +962,7 @@ class Ha(object):
def _do_reinitialize(self, cluster):
self.state_handler.stop('immediate')
# Commented redundant data directory cleanup here
# self.state_handler.remove_data_directory()
self.state_handler.remove_data_directory()
clone_member = self.cluster.get_clone_member(self.state_handler.name)
member_role = 'leader' if clone_member == self.cluster.leader else 'replica'
@@ -1112,7 +1055,6 @@ class Ha(object):
if not self.watchdog.activate():
logger.error('Cancelling bootstrap because watchdog activation failed')
self.cancel_initialization()
self.state_handler.sync_replication_slots(self.cluster)
self.dcs.take_leader()
self.set_is_leader(True)
self.state_handler.call_nowait(ACTION_ON_START)
@@ -1271,11 +1213,11 @@ class Ha(object):
# stops PostgreSQL, therefore, we only reload replication slots if no
# asynchronous processes are running (should be always the case for the master)
if not self._async_executor.busy and not self.state_handler.is_starting():
self.state_handler.sync_replication_slots(self.cluster)
if not self.state_handler.cb_called:
if not self.state_handler.is_leader():
self.state_handler.trigger_check_diverged_lsn()
self.state_handler.call_nowait(ACTION_ON_START)
self.state_handler.sync_replication_slots(self.cluster)
except DCSError:
dcs_failed = True
logger.error('Error communicating with DCS')
@@ -1333,14 +1275,15 @@ class Ha(object):
This usually happens on the master or if the node is running async action"""
self.dcs.event.set()
def get_remote_master(self):
def get_remote_master(self, config):
""" In case of standby cluster this will tel us from which remote
master to stream. Config can be both patroni config or
cluster.config.data
"""
cluster_params = self.get_standby_cluster_config()
config = config or (self.config is not None and self.config.data)
if cluster_params:
if config and config.get('standby_cluster'):
cluster_params = config.get('standby_cluster')
unique_name = 'remote_master:{}'.format(uuid.uuid1())
data = {
'conn_kwargs': {
-66
View File
@@ -1,66 +0,0 @@
import logging
import os
from copy import deepcopy
from logging.handlers import RotatingFileHandler
from patroni.utils import deep_compare
class PatroniLogger(object):
DEFAULT_LEVEL = 'INFO'
DEFAULT_FORMAT = '%(asctime)s %(levelname)s: %(message)s'
def __init__(self):
self.root_logger = logging.getLogger()
self.config = None
self.handler = None
self.reload_config({'level': 'DEBUG'})
def update_loggers(self):
loggers = deepcopy(self.config.get('loggers') or {})
for name, logger in self.root_logger.manager.loggerDict.items():
if not isinstance(logger, logging.PlaceHolder):
level = loggers.pop(name, logging.NOTSET)
logger.setLevel(level)
for name, level in loggers.items():
logger = self.root_logger.manager.getLogger(name)
logger.setLevel(level)
def reload_config(self, config):
if self.config is None or not deep_compare(self.config, config):
self.root_logger.setLevel(config.get('level', PatroniLogger.DEFAULT_LEVEL))
add_handler = None
if 'dir' in config:
if not isinstance(self.handler, RotatingFileHandler):
add_handler = RotatingFileHandler(os.path.join(config['dir'], __name__))
handler = add_handler or self.handler
handler.maxBytes = int(config.get('file_size', 25000000))
handler.backupCount = int(config.get('file_num', 4))
else:
if self.handler is None or isinstance(self.handler, RotatingFileHandler):
add_handler = logging.StreamHandler()
handler = add_handler or self.handler
oldlogformat = (self.config or {}).get('format', PatroniLogger.DEFAULT_FORMAT)
logformat = config.get('format', PatroniLogger.DEFAULT_FORMAT)
olddateformat = (self.config or {}).get('dateformat') or None
dateformat = config.get('dateformat') or None # Convert empty string to `None`
if oldlogformat != logformat or olddateformat != dateformat or add_handler:
handler.setFormatter(logging.Formatter(logformat, dateformat))
if add_handler:
self.root_logger.addHandler(add_handler)
if self.handler is not None:
self.root_logger.removeHandler(self.handler)
self.handler.close()
self.handler = add_handler
self.config = config.copy()
self.update_loggers()
+102 -131
View File
@@ -5,7 +5,6 @@ import re
import shlex
import shutil
import socket
import stat
import subprocess
import tempfile
import time
@@ -16,7 +15,7 @@ from patroni.callback_executor import CallbackExecutor
from patroni.exceptions import PostgresConnectionException, PostgresException
from patroni.utils import compare_values, parse_bool, parse_int, Retry, RetryFailedError, polling_loop, split_host_port
from patroni.postmaster import PostmasterProcess
from patroni.dcs import slot_name_from_member_name, RemoteMember
from patroni.dcs import RemoteMember
from requests.structures import CaseInsensitiveDict
from six import string_types
from six.moves.urllib.parse import quote_plus
@@ -38,16 +37,14 @@ STATE_UNKNOWN = 'unknown'
STOP_POLLING_INTERVAL = 1
REWIND_STATUS = type('Enum', (), {'INITIAL': 0, 'CHECK': 1, 'NEED': 2, 'NOT_NEED': 3, 'SUCCESS': 4, 'FAILED': 5})
sync_standby_name_re = re.compile(r'^[A-Za-z_][A-Za-z_0-9\$]*$')
sync_standby_name_re = re.compile('^[A-Za-z_][A-Za-z_0-9\$]*$')
cluster_info_query = ("SELECT CASE WHEN pg_catalog.pg_is_in_recovery() THEN 0 "
"ELSE ('x' || pg_catalog.substr(pg_catalog.pg_{0}file_name("
"pg_catalog.pg_current_{0}_{1}()), 1, 8))::bit(32)::int END, "
"CASE WHEN pg_catalog.pg_is_in_recovery() THEN GREATEST("
" pg_catalog.pg_{0}_{1}_diff(COALESCE("
"pg_catalog.pg_last_{0}_receive_{1}(), '0/0'), '0/0')::bigint,"
" pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_last_{0}_replay_{1}(), '0/0')::bigint)"
"ELSE pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_current_{0}_{1}(), '0/0')::bigint END")
cluster_info_query = ("SELECT CASE WHEN pg_is_in_recovery() THEN 0 "
"ELSE ('x' || SUBSTR(pg_{0}file_name(pg_current_{0}_{1}()), 1, 8))::bit(32)::int END, "
"CASE WHEN pg_is_in_recovery() THEN GREATEST("
" pg_{0}_{1}_diff(COALESCE(pg_last_{0}_receive_{1}(), '0/0'), '0/0')::bigint,"
" pg_{0}_{1}_diff(pg_last_{0}_replay_{1}(), '0/0')::bigint)"
"ELSE pg_{0}_{1}_diff(pg_current_{0}_{1}(), '0/0')::bigint END")
def quote_ident(value):
@@ -55,6 +52,22 @@ def quote_ident(value):
return value if sync_standby_name_re.match(value) else '"' + value + '"'
def slot_name_from_member_name(member_name):
"""Translate member name to valid PostgreSQL slot name.
PostgreSQL replication slot names must be valid PostgreSQL names. This function maps the wider space of
member names to valid PostgreSQL names. Names are lowercased, dashes and periods common in hostnames
are replaced with underscores, other characters are encoded as their unicode codepoint. Name is truncated
to 64 characters. Multiple different member names may map to a single slot name."""
def replace_char(match):
c = match.group(0)
return '_' if c in '-.' else "u{:04d}".format(ord(c))
slot_name = re.sub('[^a-z0-9_]', replace_char, member_name.lower())
return slot_name[0:63]
@contextmanager
def null_context():
yield
@@ -141,7 +154,7 @@ class Postgresql(object):
self._connection = None
self._cursor_holder = None
self._sysid = None
self._replication_slots = {} # already existing replication slots
self._replication_slots = [] # list of already existing replication slots
self.retry = Retry(max_tries=-1, deadline=config['retry_timeout']/2.0, max_delay=1,
retry_exceptions=PostgresConnectionException)
@@ -311,8 +324,8 @@ class Postgresql(object):
changes['wal_segment_size'] = '16384kB'
# XXX: query can raise an exception
for r in self.query("""SELECT name, setting, unit, vartype, context
FROM pg_catalog.pg_settings
WHERE pg_catalog.lower(name) IN (""" + ', '.join(['%s'] * len(changes)) + """)
FROM pg_settings
WHERE LOWER(name) IN (""" + ', '.join(['%s'] * len(changes)) + """)
ORDER BY 1 DESC""", *(k.lower() for k in changes.keys())):
if r[4] == 'internal':
if r[0] == 'wal_segment_size':
@@ -392,7 +405,7 @@ class Postgresql(object):
we have either wal_log_hints or checksums turned on
"""
# low-hanging fruit: check if pg_rewind configuration is there
if not self.config.get('use_pg_rewind'):
if not (self.config.get('use_pg_rewind') and all(self._superuser.get(n) for n in ('username', 'password'))):
return False
cmd = [self._pgcommand('pg_rewind'), '--help']
@@ -627,8 +640,7 @@ class Postgresql(object):
return os.environ.copy()
with open(self._pgpass, 'w') as f:
if os.name != 'nt':
os.fchmod(f.fileno(), 0o600)
os.fchmod(f.fileno(), 0o600)
f.write('{host}:{port}:*:{user}:{password}\n'.format(**record))
env = os.environ.copy()
@@ -693,10 +705,8 @@ class Postgresql(object):
# if basebackup succeeds, exit with success
break
else:
if not self.data_directory_empty() and not self.config.get(replica_method, {}).get('keep_data', False):
if not self.data_directory_empty():
self.remove_data_directory()
else:
logger.info('Leaving data directory uncleaned')
cmd = replica_method
method_config = {}
@@ -709,14 +719,10 @@ class Postgresql(object):
cmd = method_config.pop('command', cmd)
# add the default parameters
if not method_config.get('no_params', False):
method_config.update({"scope": self.scope,
"role": "replica",
"datadir": self._data_dir,
"connstring": connstring})
else:
for param in ('no_params', 'no_master', 'keep_data'):
method_config.pop(param, None)
method_config.update({"scope": self.scope,
"role": "replica",
"datadir": self._data_dir,
"connstring": connstring})
params = ["--{0}={1}".format(arg, val) for arg, val in method_config.items()]
try:
# call script with the full set of parameters
@@ -944,7 +950,7 @@ class Postgresql(object):
with self._get_connection_cursor(**connect_kwargs) as cur:
cur.execute("SET statement_timeout = 0")
if check_not_is_in_recovery:
cur.execute('SELECT pg_catalog.pg_is_in_recovery()')
cur.execute('SELECT pg_is_in_recovery()')
if cur.fetchone()[0]:
return 'is_in_recovery=true'
return cur.execute('CHECKPOINT')
@@ -1111,21 +1117,16 @@ class Postgresql(object):
f.write(self._CONFIG_WARNING_HEADER)
f.write("include '{0}'\n\n".format(self.config.get('custom_conf') or self._postgresql_base_conf_name))
for name, value in sorted((configuration or self._server_parameters).items()):
if not self._running_custom_bootstrap or name not in ('hba_file', 'archive_mode'):
if not self._running_custom_bootstrap or name != 'hba_file':
f.write("{0} = '{1}'\n".format(name, value))
# we want to set archive_mode to 'off' during the custom bootstrap
# in order to avoid premature archiving of wals and history files
if self._running_custom_bootstrap:
f.write("archive_mode = 'off'\n")
# when we are doing custom bootstrap we assume that we don't know superuser password
# and in order to be able to change it, we are opening trust access from a certain address
# therefore we need to make sure that hba_file is not overriden
# after changing superuser password we will "revert" all these "changes"
if self._running_custom_bootstrap or 'hba_file' not in self._server_parameters:
f.write("hba_file = '{0}'\n".format(self._pg_hba_conf.replace('\\', '\\\\')))
f.write("hba_file = '{0}'\n".format(self._pg_hba_conf))
if 'ident_file' not in self._server_parameters:
s = "ident_file = '{0}'\n".format(os.path.join(self._config_dir, 'pg_ident.conf').replace('\\', '\\\\'))
f.write(s)
f.write("ident_file = '{0}'\n".format(os.path.join(self._config_dir, 'pg_ident.conf')))
def is_healthy(self):
if not self.is_running():
@@ -1159,10 +1160,8 @@ class Postgresql(object):
with open(self._pg_hba_conf, 'w') as f:
f.write(self._CONFIG_WARNING_HEADER)
for address, t in addresses.items():
f.write((
'{0}\treplication\t{1}\t{3}\ttrust\n'
'{0}\tall\t{2}\t{3}\ttrust\n'
).format(t, self._replication['username'], self._superuser.get('username') or 'all', address))
f.write('{0}\t{1}\t{2}\t{3}\ttrust\n'.format(t, 'all',
self._superuser.get('username') or 'all', address))
elif not self._server_parameters.get('hba_file') and self.config.get('pg_hba'):
with open(self._pg_hba_conf, 'w') as f:
f.write(self._CONFIG_WARNING_HEADER)
@@ -1193,7 +1192,6 @@ class Postgresql(object):
def write_recovery_conf(self, recovery_params):
with open(self._recovery_conf, 'w') as f:
os.chmod(self._recovery_conf, stat.S_IWRITE | stat.S_IREAD)
for name, value in recovery_params.items():
f.write("{0} = '{1}'\n".format(name, value))
@@ -1204,7 +1202,7 @@ class Postgresql(object):
('user', r.get('user')),
('host', r.get('host')),
('port', r.get('port')),
('dbname', r.get('database') or self._database),
('dbname', r.get('database')),
('sslmode', 'prefer'),
('sslcompression', '1'),
]
@@ -1223,10 +1221,8 @@ class Postgresql(object):
# Don't try to call pg_controldata during backup restore
if self._version_file_exists() and self.state != 'creating replica':
try:
env = {'LANG': 'C', 'LC_ALL': 'C', 'PATH': os.getenv('PATH')}
if os.getenv('SYSTEMROOT') is not None:
env['SYSTEMROOT'] = os.getenv('SYSTEMROOT')
data = subprocess.check_output([self._pgcommand('pg_controldata'), self._data_dir], env=env)
data = subprocess.check_output([self._pgcommand('pg_controldata'), self._data_dir],
env={'LANG': 'C', 'LC_ALL': 'C', 'PATH': os.environ['PATH']})
if data:
data = data.decode('utf-8').splitlines()
# pg_controldata output depends on major verion. Some of parameters are prefixed by 'Current '
@@ -1249,10 +1245,8 @@ class Postgresql(object):
yield cur
@contextmanager
def _get_replication_connection_cursor(self, host='localhost', port=5432, database=None, **kwargs):
database = database or self._database
replication = 'database' if self._major_version >= 90400 else 1
with self._get_connection_cursor(host=host, port=int(port), database=database, replication=replication,
def _get_replication_connection_cursor(self, host='localhost', port=5432, **kwargs):
with self._get_connection_cursor(host=host, port=int(port), database=self._database, replication=1,
user=self._replication['username'], password=self._replication['password'],
connect_timeout=3, options='-c statement_timeout=2000') as cur:
yield cur
@@ -1260,7 +1254,7 @@ class Postgresql(object):
def check_leader_is_not_in_recovery(self, **kwargs):
try:
with self._get_connection_cursor(connect_timeout=3, options='-c statement_timeout=2000', **kwargs) as cur:
cur.execute('SELECT pg_catalog.pg_is_in_recovery()')
cur.execute('SELECT pg_is_in_recovery()')
if not cur.fetchone()[0]:
return True
logger.info('Leader is still in_recovery and therefore can\'t be used for rewind')
@@ -1371,10 +1365,10 @@ class Postgresql(object):
history_path = 'pg_{0}/{1:08X}.history'.format(self.wal_name, timeline)
try:
cursor = self._cursor()
cursor.execute('SELECT isdir, modification FROM pg_catalog.pg_stat_file(%s)', (history_path,))
cursor.execute('SELECT isdir, modification FROM pg_stat_file(%s)', (history_path,))
isdir, modification = cursor.fetchone()
if not isdir:
cursor.execute('SELECT pg_catalog.pg_read_file(%s)', (history_path,))
cursor.execute('SELECT pg_read_file(%s)', (history_path,))
history = list(self.parse_history(cursor.fetchone()[0]))
if history[-1][0] == timeline - 1:
history[-1].append(modification.isoformat())
@@ -1503,7 +1497,7 @@ class Postgresql(object):
if data.get('Database cluster state') == 'in production':
return True
def promote(self, wait_seconds, access_is_restricted=False):
def promote(self, wait_seconds):
if self.role == 'master':
return True
ret = self.pg_ctl('promote', '-W')
@@ -1511,8 +1505,7 @@ class Postgresql(object):
self.set_role('master')
logger.info("cleared rewind state after becoming the leader")
self._rewind_state = REWIND_STATUS.INITIAL
if not access_is_restricted:
self.call_nowait(ACTION_ON_ROLE_CHANGE)
self.call_nowait(ACTION_ON_ROLE_CHANGE)
ret = self._wait_promote(wait_seconds)
return ret
@@ -1545,88 +1538,61 @@ $$""".format(name, ' '.join(options)), name, password, password)
def load_replication_slots(self):
if self.use_slots and self._schedule_load_slots:
replication_slots = {}
cursor = self._query('SELECT slot_name, slot_type, plugin, database FROM pg_catalog.pg_replication_slots')
for r in cursor:
value = {'type': r[1]}
if r[1] == 'logical':
value.update({'plugin': r[2], 'database': r[3]})
replication_slots[r[0]] = value
self._replication_slots = replication_slots
cursor = self._query("SELECT slot_name FROM pg_replication_slots WHERE slot_type='physical'")
self._replication_slots = [r[0] for r in cursor]
self._schedule_load_slots = False
def postmaster_start_time(self):
try:
cursor = self.query("SELECT pg_catalog.to_char(pg_catalog.pg_postmaster_start_time(),"
" 'YYYY-MM-DD HH24:MI:SS.MS TZ')")
cursor = self.query("""SELECT to_char(pg_postmaster_start_time(), 'YYYY-MM-DD HH24:MI:SS.MS TZ')""")
return cursor.fetchone()[0]
except psycopg2.Error:
return None
def drop_replication_slot(self, name):
cursor = self._query(('SELECT pg_catalog.pg_drop_replication_slot(%s) WHERE EXISTS (SELECT 1 ' +
'FROM pg_catalog.pg_replication_slots WHERE slot_name = %s AND NOT active)'), name, name)
# In normal situation rowcount should be 1, otherwise either slot doesn't exists or it is still active
return cursor.rowcount == 1
@staticmethod
def compare_slots(s1, s2):
return s1['type'] == s2['type'] and\
(s1['type'] == 'physical' or s1['database'] == s2['database'] and s1['plugin'] == s2['plugin'])
def sync_replication_slots(self, cluster):
if self.use_slots:
try:
self.load_replication_slots()
# if the replicatefrom tag is set on the member - we should not create the replication slot for it on
# the current master, because that member would replicate from elsewhere. We still create the slot if
# the replicatefrom destination member is currently not a member of the cluster (fallback to the
# master), or if replicatefrom destination member happens to be the current master
if self.role in ('master', 'standby_leader'):
slot_members = [m.name for m in cluster.members if m.name != self.name and
(m.replicatefrom is None or m.replicatefrom == self.name or
not cluster.has_member(m.replicatefrom))]
else:
# only manage slots for replicas that replicate from this one, except for the leader among them
slot_members = [m.name for m in cluster.members if m.replicatefrom == self.name and
m.name != cluster.leader.name]
slots = set(slot_name_from_member_name(name) for name in slot_members)
slots = cluster.get_replication_slots(self.name, self.role)
if len(slots) < len(slot_members):
# Find which names are conflicting for a nicer error message
slot_conflicts = defaultdict(list)
for name in slot_members:
slot_conflicts[slot_name_from_member_name(name)].append(name)
logger.error("Following cluster members share a replication slot name: %s",
"; ".join("{} map to {}".format(", ".join(v), k)
for k, v in slot_conflicts.items() if len(v) > 1))
# drop old replication slots which are not presented in desired slots
for name in set(self._replication_slots) - set(slots):
if not self.drop_replication_slot(name):
logger.error("Failed to drop replication slot '%s'", name)
self._schedule_load_slots = True
# drop unused slots
for slot in set(self._replication_slots) - slots:
cursor = self._query("""SELECT pg_drop_replication_slot(%s)
WHERE EXISTS(SELECT 1 FROM pg_replication_slots
WHERE slot_name = %s AND NOT active)""", slot, slot)
if cursor.rowcount != 1: # Either slot doesn't exists or it is still active
self._schedule_load_slots = True # schedule load_replication_slots on the next iteration
immediately_reserve = ', true' if self._major_version >= 90600 else ''
logical_slots = defaultdict(dict)
for name, value in slots.items():
if name in self._replication_slots and not self.compare_slots(value, self._replication_slots[name]):
logger.info("Trying to drop replication slot '%s' because value is changing from %s to %s",
name, self._replication_slots[name], value)
if not self.drop_replication_slot(name):
logger.error("Failed to drop replication slot '%s'", name)
self._schedule_load_slots = True
continue
self._replication_slots.pop(name)
if name not in self._replication_slots:
if value['type'] == 'physical':
try:
self._query(("SELECT pg_catalog.pg_create_physical_replication_slot(%s{0})" +
" WHERE NOT EXISTS (SELECT 1 FROM pg_catalog.pg_replication_slots" +
" WHERE slot_type = 'physical' AND slot_name = %s)").format(
immediately_reserve), name, name)
except Exception:
logger.exception("Failed to create physical replication slot '%s'", name)
self._schedule_load_slots = True
elif value['type'] == 'logical' and name not in self._replication_slots:
logical_slots[value['database']][name] = value
# create new slots
for slot in slots - set(self._replication_slots):
self._query("""SELECT pg_create_physical_replication_slot(%s{0})
WHERE NOT EXISTS (SELECT 1 FROM pg_replication_slots
WHERE slot_name = %s)""".format(immediately_reserve), slot, slot)
# create new logical slots
for database, values in logical_slots.items():
conn_kwargs = self._local_connect_kwargs
conn_kwargs['database'] = database
with self._get_connection_cursor(**conn_kwargs) as cur:
for name, value in values.items():
try:
cur.execute("SELECT pg_catalog.pg_create_logical_replication_slot(%s, %s)" +
" WHERE NOT EXISTS (SELECT 1 FROM pg_catalog.pg_replication_slots" +
" WHERE slot_type = 'logical' AND slot_name = %s)",
(name, value['plugin'], name))
except Exception:
logger.exception("Failed to create logical replication slot '%s' plugin='%s'",
name, value['plugin'])
self._schedule_load_slots = True
self._replication_slots = slots
except Exception:
logger.exception('Exception when changing replication slots')
@@ -1673,8 +1639,7 @@ $$""".format(name, ' '.join(options)), name, password, password)
def post_bootstrap(self, config, task):
try:
if 'username' in self._superuser and 'password' in self._superuser:
self.create_or_update_role(self._superuser['username'], self._superuser['password'], ['SUPERUSER'])
self.create_or_update_role(self._superuser['username'], self._superuser['password'], ['SUPERUSER'])
task.complete(self.run_bootstrap_post_init(config))
if task.result:
@@ -1693,11 +1658,17 @@ $$""".format(name, ' '.join(options)), name, password, password)
os.unlink(self._pg_hba_conf)
self.restore_configuration_files()
self._write_postgresql_conf()
self._replace_pg_hba()
# at this point there should be no recovery.conf
if os.path.isfile(self._recovery_conf) or os.path.islink(self._recovery_conf):
os.unlink(self._recovery_conf)
self.restart()
if self._server_parameters.get('hba_file') and \
self._server_parameters['hba_file'] != self._pg_hba_conf:
self.restart()
else:
self._replace_pg_hba()
if self.pending_restart:
self.restart()
else:
self.reload()
time.sleep(1) # give a time to postgres to "reload" configuration files
self.close_connection() # close connection to reconnect with a new password
except Exception:
logger.exception('post_bootstrap')
task.complete(False)
@@ -1778,9 +1749,9 @@ $$""".format(name, ' '.join(options)), name, password, password)
# Pick candidates based on who has flushed WAL farthest.
# TODO: for synchronous_commit = remote_write we actually want to order on write_location
for app_name, state, sync_state in self.query(
"SELECT pg_catalog.lower(application_name), state, sync_state"
" FROM pg_catalog.pg_stat_replication"
" ORDER BY flush_{0} DESC".format(self.lsn_name)):
"""SELECT LOWER(application_name), state, sync_state
FROM pg_stat_replication
ORDER BY flush_{0} DESC""".format(self.lsn_name)):
member = members.get(app_name)
if state != 'streaming' or not member or member.tags.get('nosync', False):
continue
+25 -29
View File
@@ -12,7 +12,7 @@ logger = logging.getLogger(__name__)
STOP_SIGNALS = {
'smart': signal.SIGTERM,
'fast': signal.SIGINT,
'immediate': signal.SIGQUIT if os.name != 'nt' else signal.SIGABRT,
'immediate': signal.SIGQUIT,
}
@@ -102,31 +102,28 @@ class PostmasterProcess(psutil.Process):
return None
def wait_for_user_backends_to_close(self):
# These regexps are cross checked against versions PostgreSQL 9.1 .. 11
aux_proc_re = re.compile("(?:postgres:)( .*:)? (?:(?:archiver|startup|autovacuum launcher|autovacuum worker|"
"checkpointer|logger|stats collector|wal receiver|wal writer|writer)(?: process )?|"
"walreceiver|wal sender process|walsender|walwriter|background writer|"
"logical replication launcher|logical replication worker for|bgworker:) ")
# These regexps are cross checked against versions PostgreSQL 9.1 .. 9.6
aux_proc_re = re.compile("(?:postgres:)( .*:)? (?:""(?:startup|logger|checkpointer|writer|wal writer|"
"autovacuum launcher|autovacuum worker|stats collector|wal receiver|archiver|"
"wal sender) process|bgworker: )")
try:
children = self.children()
user_backends = []
user_backends_cmdlines = []
for child in self.children():
try:
cmdline = child.cmdline()[0]
if not aux_proc_re.match(cmdline):
user_backends.append(child)
user_backends_cmdlines.append(cmdline)
except psutil.NoSuchProcess:
pass
if user_backends:
logger.debug('Waiting for user backends %s to close', ', '.join(user_backends_cmdlines))
psutil.wait_procs(user_backends)
logger.debug("Backends closed")
except psutil.Error:
return logger.debug('Failed to get list of postmaster children')
user_backends = []
user_backends_cmdlines = []
for child in children:
try:
cmdline = child.cmdline()[0]
if not aux_proc_re.match(cmdline):
user_backends.append(child)
user_backends_cmdlines.append(cmdline)
except psutil.NoSuchProcess:
pass
if user_backends:
logger.debug('Waiting for user backends %s to close', ', '.join(user_backends_cmdlines))
psutil.wait_procs(user_backends)
logger.debug("Backends closed")
logger.exception('wait_for_user_backends_to_close')
@staticmethod
def start(pgcommand, data_dir, conf, options):
@@ -141,8 +138,7 @@ class PostmasterProcess(psutil.Process):
# of init process to take care about postmaster.
# In order to make everything portable we can't use fork&exec approach here, so we will call
# ourselves and pass list of arguments which must be used to start postgres.
# On Windows, in order to run a side-by-side assembly the specified env must include a valid SYSTEMROOT.
env = {p: os.environ[p] for p in ('PATH', 'LD_LIBRARY_PATH', 'LC_ALL', 'LANG', 'SYSTEMROOT') if p in os.environ}
env = {p: os.environ[p] for p in ('PATH', 'LD_LIBRARY_PATH', 'LC_ALL', 'LANG') if p in os.environ}
try:
proc = PostmasterProcess._from_pidfile(data_dir)
if proc and not proc._is_postmaster_process():
@@ -157,10 +153,10 @@ class PostmasterProcess(psutil.Process):
env['PG_GRANDPARENT_PID'] = str(proc.pid)
except psutil.NoSuchProcess:
pass
cmdline = [pgcommand, '-D', data_dir, '--config-file={}'.format(conf)] + options
logger.debug("Starting postgres: %s", " ".join(cmdline))
proc = call_self(['pg_ctl_start'] + cmdline, close_fds=(os.name != 'nt'),
stdout=subprocess.PIPE, env=env)
proc = call_self(['pg_ctl_start', pgcommand, '-D', data_dir,
'--config-file={}'.format(conf)] + options, close_fds=True,
preexec_fn=os.setsid, stdout=subprocess.PIPE, env=env)
pid = int(proc.stdout.readline().strip())
proc.wait()
logger.info('postmaster pid=%s', pid)
+7 -6
View File
@@ -224,12 +224,13 @@ class WALERestore(object):
lsn_name = 'location'
con.autocommit = True
with con.cursor() as cur:
cur.execute(("SELECT CASE WHEN pg_catalog.pg_is_in_recovery()"
" THEN GREATEST(pg_catalog.pg_{0}_{1}_diff(COALESCE("
"pg_last_{0}_receive_{1}(), '0/0'), %s)::bigint, "
"pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_last_{0}_replay_{1}(), %s)::bigint)"
" ELSE pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_current_{0}_{1}(), %s)::bigint"
" END").format(wal_name, lsn_name),
cur.execute("""SELECT CASE WHEN pg_is_in_recovery()
THEN GREATEST(
pg_{0}_{1}_diff(COALESCE(
pg_last_{0}_receive_{1}(), '0/0'), %s)::bigint,
pg_{0}_{1}_diff(pg_last_{0}_replay_{1}(), %s)::bigint)
ELSE pg_{0}_{1}_diff(pg_current_{0}_{1}(), %s)::bigint
END""".format(wal_name, lsn_name),
(backup_start_lsn, backup_start_lsn, backup_start_lsn))
diff_in_bytes = int(cur.fetchone()[0])
+1 -1
View File
@@ -1 +1 @@
__version__ = '1.5.4'
__version__ = '1.4.6'
-3
View File
@@ -52,7 +52,6 @@ CLASSIFIERS = [
'Operating System :: MacOS',
'Operating System :: POSIX :: Linux',
'Operating System :: POSIX :: BSD :: FreeBSD',
'Operating System :: Microsoft :: Windows',
'Programming Language :: Python',
'Programming Language :: Python :: 2.7',
'Programming Language :: Python :: 3',
@@ -107,8 +106,6 @@ class PyTest(TestCommand):
silence = logging.WARNING
logging.basicConfig(format='%(asctime)s %(levelname)s: %(message)s', level=os.getenv('LOGLEVEL', silence))
params['args'] += ['-s' if logging.getLogger().getEffectiveLevel() < silence else '--capture=fd']
if not os.getenv('SYSTEMROOT'):
os.environ['SYSTEMROOT'] = '/'
errno = pytest.main(**params)
sys.exit(errno)
+1 -10
View File
@@ -73,7 +73,7 @@ class MockHa(object):
@staticmethod
def fetch_nodes_statuses(members):
return [_MemberStatus(None, True, None, 0, None, {}, False)]
return [_MemberStatus(None, True, None, None, {}, False)]
@staticmethod
def schedule_future_restart(data):
@@ -95,10 +95,6 @@ class MockHa(object):
def is_paused():
return True
@staticmethod
def is_standby_cluster():
return False
class MockPatroni(object):
@@ -171,8 +167,6 @@ class TestRestApiHandler(unittest.TestCase):
self.assertIsNotNone(MockRestApiServer(RestApiHandler, 'GET /master'))
with patch.object(RestApiServer, 'query', Mock(return_value=[('', 1, '', '', '', '', False, '')])):
self.assertIsNotNone(MockRestApiServer(RestApiHandler, 'GET /patroni'))
with patch.object(MockHa, 'is_standby_cluster', Mock(return_value=True)):
MockRestApiServer(RestApiHandler, 'GET /standby_leader')
def test_do_OPTIONS(self):
self.assertIsNotNone(MockRestApiServer(RestApiHandler, 'OPTIONS / HTTP/1.0'))
@@ -401,6 +395,3 @@ class TestRestApiServer(unittest.TestCase):
self.assertRaises(ValueError, srv.reload_config, bad_config)
self.assertRaises(ValueError, srv.reload_config, {})
srv.reload_config({'listen': '127.0.0.2:8008'})
def test_handle_error(self):
self.assertIsNone(MockRestApiServer.handle_error(None, ('127.0.0.1', 55555)))
+2 -17
View File
@@ -1,6 +1,6 @@
import os
import sys
import unittest
import sys
from mock import MagicMock, Mock, patch
from patroni.config import Config
@@ -30,8 +30,6 @@ class TestConfig(unittest.TestCase):
'PATRONI_NAME': 'postgres0',
'PATRONI_NAMESPACE': '/patroni/',
'PATRONI_SCOPE': 'batman2',
'PATRONI_LOGLEVEL': 'ERROR',
'PATRONI_LOG_LOGGERS': 'patroni.postmaster: WARNING, urllib3: DEBUG',
'PATRONI_RESTAPI_USERNAME': 'username',
'PATRONI_RESTAPI_PASSWORD': 'password',
'PATRONI_RESTAPI_LISTEN': '0.0.0.0:8008',
@@ -51,7 +49,6 @@ class TestConfig(unittest.TestCase):
'PATRONI_ETCD_CERT': '/cert',
'PATRONI_ETCD_KEY': '/key',
'PATRONI_CONSUL_HOST': '127.0.0.1:8500',
'PATRONI_CONSUL_REGISTER_SERVICE': 'on',
'PATRONI_KUBERNETES_LABELS': 'a:b:c',
'PATRONI_KUBERNETES_SCOPE_LABEL': 'a',
'PATRONI_KUBERNETES_PORTS': '[{"name": "postgresql"}]',
@@ -79,7 +76,7 @@ class TestConfig(unittest.TestCase):
@patch('os.path.exists', Mock(return_value=True))
@patch('os.remove', Mock(side_effect=IOError))
@patch('os.close', Mock(side_effect=IOError))
@patch('shutil.move', Mock(return_value=None))
@patch('os.rename', Mock(return_value=None))
@patch('json.dump', Mock())
def test_save_cache(self):
self.config.set_dynamic_configuration({'ttl': 30, 'postgresql': {'foo': 'bar'}})
@@ -87,15 +84,3 @@ class TestConfig(unittest.TestCase):
self.config.save_cache()
with patch('os.fdopen', MagicMock()):
self.config.save_cache()
def test_standby_cluster_parameters(self):
dynamic_configuration = {
'standby_cluster': {
'create_replica_methods': ['wal_e', 'basebackup'],
'host': 'localhost',
'port': 5432
}
}
self.config.set_dynamic_configuration(dynamic_configuration)
for name, value in dynamic_configuration['standby_cluster'].items():
self.assertEqual(self.config['standby_cluster'][name], value)
+9 -10
View File
@@ -4,7 +4,7 @@ import unittest
from consul import ConsulException, NotFound
from mock import Mock, patch
from patroni.dcs.consul import AbstractDCS, Cluster, Consul, ConsulInternalError, \
ConsulError, HTTPClient, InvalidSessionTTL, InvalidSession
ConsulError, HTTPClient, InvalidSessionTTL
from test_etcd import SleepException
@@ -52,8 +52,6 @@ class TestHTTPClient(unittest.TestCase):
self.assertRaises(ConsulInternalError, self.client.get, Mock(), '')
self.client.http.request.return_value.data = b"Invalid Session TTL '3000000000', must be between [10s=24h0m0s]"
self.assertRaises(InvalidSessionTTL, self.client.get, Mock(), '')
self.client.http.request.return_value.data = b"invalid session '16492f43-c2d6-5307-432f-e32d6f7bcbd0'"
self.assertRaises(InvalidSession, self.client.get, Mock(), '')
def test_unknown_method(self):
try:
@@ -112,18 +110,19 @@ class TestConsul(unittest.TestCase):
self.c._session = 'fd4f44fe-2cac-bba5-a60b-304b51ff39b8'
self.assertIsInstance(self.c.get_cluster(), Cluster)
@patch.object(consul.Consul.KV, 'delete', Mock(side_effect=[ConsulException, True, True, True]))
@patch.object(consul.Consul.KV, 'put', Mock(side_effect=[True, ConsulException, InvalidSession]))
@patch.object(consul.Consul.KV, 'delete', Mock(side_effect=[ConsulException, True, True]))
@patch.object(consul.Consul.KV, 'put', Mock(side_effect=[True, ConsulException]))
def test_touch_member(self):
self.c._register_service = True
self.c.refresh_session = Mock(return_value=True)
self.c.touch_member({'balbla': 'blabla'})
self.c.touch_member({'balbla': 'blabla'})
self.c.touch_member({'balbla': 'blabla'})
self.c.refresh_session = Mock(return_value=False)
self.c.touch_member({'conn_url': 'postgres://replicator:[email protected]:5433/postgres',
'api_url': 'http://127.0.0.1:8009/patroni'})
self.c._register_service = True
self.c.refresh_session = Mock(return_value=True)
for _ in range(0, 4):
self.c.touch_member({'balbla': 'blabla'})
@patch.object(consul.Consul.KV, 'put', Mock(side_effect=InvalidSession))
@patch.object(consul.Consul.KV, 'put', Mock(return_value=False))
def test_take_leader(self):
self.c.set_ttl(20)
self.c.refresh_session = Mock()
+22 -22
View File
@@ -192,24 +192,24 @@ class TestCtl(unittest.TestCase):
def test_query_member(self):
with patch('patroni.ctl.get_cursor', Mock(return_value=MockConnect().cursor())):
rows = query_member(None, None, None, 'master', 'SELECT pg_catalog.pg_is_in_recovery()', {})
rows = query_member(None, None, None, 'master', 'SELECT pg_is_in_recovery()', {})
self.assertTrue('False' in str(rows))
rows = query_member(None, None, None, 'replica', 'SELECT pg_catalog.pg_is_in_recovery()', {})
self.assertEqual(rows, (None, None))
rows = query_member(None, None, None, 'replica', 'SELECT pg_is_in_recovery()', {})
self.assertEquals(rows, (None, None))
with patch('test_postgresql.MockCursor.execute', Mock(side_effect=OperationalError('bla'))):
rows = query_member(None, None, None, 'replica', 'SELECT pg_catalog.pg_is_in_recovery()', {})
rows = query_member(None, None, None, 'replica', 'SELECT pg_is_in_recovery()', {})
with patch('patroni.ctl.get_cursor', Mock(return_value=None)):
rows = query_member(None, None, None, None, 'SELECT pg_catalog.pg_is_in_recovery()', {})
rows = query_member(None, None, None, None, 'SELECT pg_is_in_recovery()', {})
self.assertTrue('No connection to' in str(rows))
rows = query_member(None, None, None, 'replica', 'SELECT pg_catalog.pg_is_in_recovery()', {})
rows = query_member(None, None, None, 'replica', 'SELECT pg_is_in_recovery()', {})
self.assertTrue('No connection to' in str(rows))
with patch('patroni.ctl.get_cursor', Mock(side_effect=OperationalError('bla'))):
rows = query_member(None, None, None, 'replica', 'SELECT pg_catalog.pg_is_in_recovery()', {})
rows = query_member(None, None, None, 'replica', 'SELECT pg_is_in_recovery()', {})
@patch('patroni.ctl.get_dcs')
def test_dsn(self, mock_get_dcs):
@@ -364,20 +364,20 @@ class TestCtl(unittest.TestCase):
self.assertIsNone(get_any_member(get_cluster_initialized_without_leader(), role='master'))
m = get_any_member(get_cluster_initialized_with_leader(), role='master')
self.assertEqual(m.name, 'leader')
self.assertEquals(m.name, 'leader')
def test_get_all_members(self):
self.assertEqual(list(get_all_members(get_cluster_initialized_without_leader(), role='master')), [])
self.assertEquals(list(get_all_members(get_cluster_initialized_without_leader(), role='master')), [])
r = list(get_all_members(get_cluster_initialized_with_leader(), role='master'))
self.assertEqual(len(r), 1)
self.assertEqual(r[0].name, 'leader')
self.assertEquals(len(r), 1)
self.assertEquals(r[0].name, 'leader')
r = list(get_all_members(get_cluster_initialized_with_leader(), role='replica'))
self.assertEqual(len(r), 1)
self.assertEqual(r[0].name, 'other')
self.assertEquals(len(r), 1)
self.assertEquals(r[0].name, 'other')
self.assertEqual(len(list(get_all_members(get_cluster_initialized_without_leader(), role='replica'))), 2)
self.assertEquals(len(list(get_all_members(get_cluster_initialized_without_leader(), role='replica'))), 2)
@patch('patroni.ctl.get_dcs')
def test_members(self, mock_get_dcs):
@@ -499,23 +499,23 @@ class TestCtl(unittest.TestCase):
after_editing, changed_config = apply_config_changes(before_editing, config,
["postgresql.parameters.work_mem = 5MB",
"ttl=15", "postgresql.use_pg_rewind=off", 'a.b=c'])
self.assertEqual(changed_config, {"a": {"b": "c"}, "postgresql": {"parameters": {"work_mem": "5MB"},
"use_pg_rewind": False}, "ttl": 15})
self.assertEquals(changed_config, {"a": {"b": "c"}, "postgresql": {"parameters": {"work_mem": "5MB"},
"use_pg_rewind": False}, "ttl": 15})
# postgresql.parameters namespace is flattened
after_editing, changed_config = apply_config_changes(before_editing, config,
["postgresql.parameters.work_mem.sub = x"])
self.assertEqual(changed_config, {"postgresql": {"parameters": {"work_mem": "4MB", "work_mem.sub": "x"},
"use_pg_rewind": True}, "ttl": 30})
self.assertEquals(changed_config, {"postgresql": {"parameters": {"work_mem": "4MB", "work_mem.sub": "x"},
"use_pg_rewind": True}, "ttl": 30})
# Setting to null deletes
after_editing, changed_config = apply_config_changes(before_editing, config,
["postgresql.parameters.work_mem=null"])
self.assertEqual(changed_config, {"postgresql": {"use_pg_rewind": True}, "ttl": 30})
self.assertEquals(changed_config, {"postgresql": {"use_pg_rewind": True}, "ttl": 30})
after_editing, changed_config = apply_config_changes(before_editing, config,
["postgresql.use_pg_rewind=null",
"postgresql.parameters.work_mem=null"])
self.assertEqual(changed_config, {"ttl": 30})
self.assertEquals(changed_config, {"ttl": 30})
self.assertRaises(PatroniCtlException, apply_config_changes, before_editing, config, ['a'])
@@ -572,5 +572,5 @@ class TestCtl(unittest.TestCase):
assert 'failed to get version' in result.output
def test_format_pg_version(self):
self.assertEqual(format_pg_version(100001), '10.1')
self.assertEqual(format_pg_version(90605), '9.6.5')
self.assertEquals(format_pg_version(100001), '10.1')
self.assertEquals(format_pg_version(90605), '9.6.5')
+3 -3
View File
@@ -215,8 +215,8 @@ class TestClient(unittest.TestCase):
self.assertRaises(etcd.EtcdException, self.client.api_execute, '/', 'GET')
def test_get_srv_record(self):
self.assertEqual(self.client.get_srv_record('_etcd-server._tcp.blabla'), [])
self.assertEqual(self.client.get_srv_record('_etcd-server._tcp.exception'), [])
self.assertEquals(self.client.get_srv_record('_etcd-server._tcp.blabla'), [])
self.assertEquals(self.client.get_srv_record('_etcd-server._tcp.exception'), [])
def test__get_machines_cache_from_srv(self):
self.client._get_machines_cache_from_srv('foobar')
@@ -259,7 +259,7 @@ class TestEtcd(unittest.TestCase):
'host': 'localhost:2379', 'scope': 'test', 'name': 'foo'})
def test_base_path(self):
self.assertEqual(self.etcd._base_path, '/patroni/test')
self.assertEquals(self.etcd._base_path, '/patroni/test')
@patch('dns.resolver.query', dns_query)
def test_get_etcd_client(self):
+181 -153
View File
@@ -5,6 +5,7 @@ import unittest
import sys
from mock import Mock, MagicMock, PropertyMock, patch
from patroni.async_executor import CriticalTask
from patroni.config import Config
from patroni.dcs import Cluster, ClusterConfig, Failover, Leader, Member, get_dcs, SyncState, TimelineHistory
from patroni.dcs.etcd import Client
@@ -28,10 +29,8 @@ def false(*args, **kwargs):
def get_cluster(initialize, leader, members, failover, sync, cluster_config=None):
t = datetime.datetime.now().isoformat()
history = TimelineHistory(1, '[[1,67197376,"no recovery target specified","' + t + '"]]',
[(1, 67197376, 'no recovery target specified', t)])
cluster_config = cluster_config or ClusterConfig(1, {'check_timeline': True}, 1)
history = TimelineHistory(1, [(1, 67197376, 'no recovery target specified', datetime.datetime.now().isoformat())])
cluster_config = cluster_config or ClusterConfig(1, {1: 2}, 1)
return Cluster(initialize, cluster_config, leader, 10, members, failover, sync, history)
@@ -63,6 +62,17 @@ def get_cluster_initialized_with_only_leader(failover=None, cluster_config=None)
return get_cluster(True, leader, [leader], failover, None, cluster_config)
def get_cluster_not_initialized_standby(failover=None, sync=None):
return get_cluster_not_initialized_without_leader(
cluster_config=ClusterConfig(1, {
"standby_cluster": {
"host": "localhost",
"port": 5432,
"primary_slot_name": "",
}}, 1)
)
def get_standby_cluster_initialized_with_only_leader(failover=None, sync=None):
return get_cluster_initialized_with_only_leader(
cluster_config=ClusterConfig(1, {
@@ -74,13 +84,12 @@ def get_standby_cluster_initialized_with_only_leader(failover=None, sync=None):
)
def get_node_status(reachable=True, in_recovery=True, timeline=2,
wal_position=10, nofailover=False, watchdog_failed=False):
def get_node_status(reachable=True, in_recovery=True, wal_position=10, nofailover=False, watchdog_failed=False):
def fetch_node_status(e):
tags = {}
if nofailover:
tags['nofailover'] = True
return _MemberStatus(e, reachable, in_recovery, timeline, wal_position, tags, watchdog_failed)
return _MemberStatus(e, reachable, in_recovery, wal_position, tags, watchdog_failed)
return fetch_node_status
@@ -118,7 +127,6 @@ zookeeper:
sys.argv = sys.argv[:1]
self.config = Config()
self.config.set_dynamic_configuration({'maximum_lag_on_failover': 5})
self.postgresql = p
self.dcs = d
self.api = Mock()
@@ -150,8 +158,8 @@ def run_async(self, func, args=()):
@patch.object(Postgresql, 'write_recovery_conf', Mock())
@patch.object(Postgresql, 'query', Mock())
@patch.object(Postgresql, 'checkpoint', Mock())
@patch.object(Postgresql, 'call_nowait', Mock())
@patch.object(Postgresql, 'cancellable_subprocess_call', Mock(return_value=0))
@patch.object(Postgresql, '_get_local_timeline_lsn_from_replication_connection', Mock(return_value=[2, 10]))
@patch.object(etcd.Client, 'write', etcd_write)
@patch.object(etcd.Client, 'read', etcd_read)
@patch.object(etcd.Client, 'delete', Mock(side_effect=etcd.EtcdException))
@@ -171,6 +179,7 @@ class TestHa(unittest.TestCase):
mock_machines.__get__ = Mock(return_value=['http://remotehost:2379'])
self.p = Postgresql({'name': 'postgresql0', 'scope': 'dummy', 'listen': '127.0.0.1:5432',
'data_dir': 'data/postgresql0', 'retry_timeout': 10,
'maximum_lag_on_failover': 5,
'authentication': {'superuser': {'username': 'foo', 'password': 'bar'},
'replication': {'username': '', 'password': ''}},
'parameters': {'wal_level': 'hot_standby', 'max_replication_slots': 5, 'foo': 'bar',
@@ -200,15 +209,22 @@ class TestHa(unittest.TestCase):
def test_start_as_replica(self):
self.p.is_healthy = false
self.assertEqual(self.ha.run_cycle(), 'starting as a secondary')
self.assertEquals(self.ha.run_cycle(), 'starting as a secondary')
@patch('patroni.dcs.etcd.Etcd.initialize', return_value=True)
def test_start_as_standby_leader(self, initialize):
self.p.data_directory_empty = true
self.ha.cluster = get_cluster_not_initialized_without_leader(cluster_config=ClusterConfig(0, {}, 0))
self.ha.cluster = get_cluster_not_initialized_standby()
self.ha.cluster.is_unlocked = true
self.ha.patroni.config._dynamic_configuration = {"standby_cluster": {"port": 5432}}
self.assertEqual(self.ha.run_cycle(), 'trying to bootstrap a new standby leader')
self.ha.patroni.config._dynamic_configuration = {"standby_cluster": {
"host": "localhost",
"port": 5432,
"primary_slot_name": "",
}}
self.assertEquals(
self.ha.run_cycle(),
'trying to bootstrap a new standby leader'
)
@patch.object(Cluster, 'get_clone_member',
Mock(return_value=Member(0, 'test', 1, {'api_url': 'http://127.0.0.1:8011/patroni'})))
@@ -217,14 +233,34 @@ class TestHa(unittest.TestCase):
self.p.data_directory_empty = true
self.ha.cluster = get_standby_cluster_initialized_with_only_leader()
self.ha.cluster.is_unlocked = false
self.assertEqual(self.ha.run_cycle(), "trying to bootstrap from replica 'test'")
self.ha.patroni.config._dynamic_configuration = {"standby_cluster": {
"host": "localhost",
"port": 5432,
"primary_slot_name": "",
}}
self.assertEquals(
self.ha.run_cycle(),
"trying to bootstrap from replica 'test'"
)
@patch.object(Postgresql, 'create_replica', Mock(return_value=0))
def test_bootstrap_standby_leader(self):
self.ha.cluster = get_cluster_not_initialized_standby()
self.ha.cluster.is_unlocked = true
self.ha.patroni.config._dynamic_configuration = {"standby_cluster": {
"host": "localhost",
"port": 5432,
"primary_slot_name": "",
}}
self.ha._post_bootstrap_task = CriticalTask()
self.assertEquals(self.ha.bootstrap_standby_leader(), True)
def test_recover_replica_failed(self):
self.p.controldata = lambda: {'Database cluster state': 'in recovery', 'Database system identifier': SYSID}
self.p.is_running = false
self.p.follow = false
self.assertEqual(self.ha.run_cycle(), 'starting as a secondary')
self.assertEqual(self.ha.run_cycle(), 'failed to start postgres')
self.assertEquals(self.ha.run_cycle(), 'starting as a secondary')
self.assertEquals(self.ha.run_cycle(), 'failed to start postgres')
def test_recover_former_master(self):
self.p.follow = false
@@ -233,19 +269,19 @@ class TestHa(unittest.TestCase):
self.p.set_role('master')
self.p.controldata = lambda: {'Database cluster state': 'shut down', 'Database system identifier': SYSID}
self.ha.cluster = get_cluster_initialized_with_leader()
self.assertEqual(self.ha.run_cycle(), 'starting as readonly because i had the session lock')
self.assertEquals(self.ha.run_cycle(), 'starting as readonly because i had the session lock')
@patch.object(Postgresql, 'fix_cluster_state', Mock())
def test_crash_recovery(self):
self.p.is_running = false
self.p.controldata = lambda: {'Database cluster state': 'in production', 'Database system identifier': SYSID}
self.assertEqual(self.ha.run_cycle(), 'doing crash recovery in a single user mode')
self.assertEquals(self.ha.run_cycle(), 'doing crash recovery in a single user mode')
@patch.object(Postgresql, 'rewind_needed_and_possible', Mock(return_value=True))
def test_recover_with_rewind(self):
self.p.is_running = false
self.ha.cluster = get_cluster_initialized_with_leader()
self.assertEqual(self.ha.run_cycle(), 'running pg_rewind from leader')
self.assertEquals(self.ha.run_cycle(), 'running pg_rewind from leader')
@patch('sys.exit', return_value=1)
@patch('patroni.ha.Ha.sysid_valid', MagicMock(return_value=True))
@@ -260,129 +296,129 @@ class TestHa(unittest.TestCase):
self.p.is_healthy = true
self.ha.has_lock = true
self.p.controldata = lambda: {'Database cluster state': 'in production', 'Database system identifier': SYSID}
self.assertEqual(self.ha.run_cycle(), 'promoted self to leader because i had the session lock')
self.assertEquals(self.ha.run_cycle(), 'promoted self to leader because i had the session lock')
@patch('psycopg2.connect', psycopg2_connect)
def test_acquire_lock_as_master(self):
self.assertEqual(self.ha.run_cycle(), 'acquired session lock as a leader')
self.assertEquals(self.ha.run_cycle(), 'acquired session lock as a leader')
def test_promoted_by_acquiring_lock(self):
self.ha.is_healthiest_node = true
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
self.assertEquals(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
def test_long_promote(self):
self.ha.cluster.is_unlocked = false
self.ha.has_lock = true
self.p.is_leader = false
self.p.set_role('master')
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
def test_demote_after_failing_to_obtain_lock(self):
self.ha.acquire_lock = false
self.assertEqual(self.ha.run_cycle(), 'demoted self after trying and failing to obtain lock')
self.assertEquals(self.ha.run_cycle(), 'demoted self after trying and failing to obtain lock')
def test_follow_new_leader_after_failing_to_obtain_lock(self):
self.ha.is_healthiest_node = true
self.ha.acquire_lock = false
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'following new leader after trying and failing to obtain lock')
self.assertEquals(self.ha.run_cycle(), 'following new leader after trying and failing to obtain lock')
def test_demote_because_not_healthiest(self):
self.ha.is_healthiest_node = false
self.assertEqual(self.ha.run_cycle(), 'demoting self because i am not the healthiest node')
self.assertEquals(self.ha.run_cycle(), 'demoting self because i am not the healthiest node')
def test_follow_new_leader_because_not_healthiest(self):
self.ha.is_healthiest_node = false
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'following a different leader because i am not the healthiest node')
self.assertEquals(self.ha.run_cycle(), 'following a different leader because i am not the healthiest node')
def test_promote_because_have_lock(self):
self.ha.cluster.is_unlocked = false
self.ha.has_lock = true
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'promoted self to leader because i had the session lock')
self.assertEquals(self.ha.run_cycle(), 'promoted self to leader because i had the session lock')
def test_promote_without_watchdog(self):
self.ha.cluster.is_unlocked = false
self.ha.has_lock = true
self.p.is_leader = true
with patch.object(Watchdog, 'activate', Mock(return_value=False)):
self.assertEqual(self.ha.run_cycle(), 'Demoting self because watchdog could not be activated')
self.assertEquals(self.ha.run_cycle(), 'Demoting self because watchdog could not be activated')
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'Not promoting self because watchdog could not be activated')
self.assertEquals(self.ha.run_cycle(), 'Not promoting self because watchdog could not be activated')
def test_leader_with_lock(self):
self.ha.cluster = get_cluster_not_initialized_without_leader()
self.ha.cluster.is_unlocked = false
self.ha.has_lock = true
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
def test_demote_because_not_having_lock(self):
self.ha.cluster.is_unlocked = false
with patch.object(Watchdog, 'is_running', PropertyMock(return_value=True)):
self.assertEqual(self.ha.run_cycle(), 'demoting self because i do not have the lock and i was a leader')
self.assertEquals(self.ha.run_cycle(), 'demoting self because i do not have the lock and i was a leader')
def test_demote_because_update_lock_failed(self):
self.ha.cluster.is_unlocked = false
self.ha.has_lock = true
self.ha.update_lock = false
self.assertEqual(self.ha.run_cycle(), 'demoted self because failed to update leader lock in DCS')
self.assertEquals(self.ha.run_cycle(), 'demoted self because failed to update leader lock in DCS')
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'not promoting because failed to update leader lock in DCS')
self.assertEquals(self.ha.run_cycle(), 'not promoting because failed to update leader lock in DCS')
def test_follow(self):
self.ha.cluster.is_unlocked = false
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'no action. i am a secondary and i am following a leader')
self.assertEquals(self.ha.run_cycle(), 'no action. i am a secondary and i am following a leader')
self.ha.patroni.replicatefrom = "foo"
self.assertEqual(self.ha.run_cycle(), 'no action. i am a secondary and i am following a leader')
self.assertEquals(self.ha.run_cycle(), 'no action. i am a secondary and i am following a leader')
def test_follow_in_pause(self):
self.ha.cluster.is_unlocked = false
self.ha.is_paused = true
self.assertEqual(self.ha.run_cycle(), 'PAUSE: continue to run as master without lock')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: continue to run as master without lock')
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'PAUSE: no action')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: no action')
@patch.object(Postgresql, 'rewind_needed_and_possible', Mock(return_value=True))
def test_follow_triggers_rewind(self):
self.p.is_leader = false
self.p.trigger_check_diverged_lsn()
self.ha.cluster = get_cluster_initialized_with_leader()
self.assertEqual(self.ha.run_cycle(), 'running pg_rewind from leader')
self.assertEquals(self.ha.run_cycle(), 'running pg_rewind from leader')
def test_no_etcd_connection_master_demote(self):
self.ha.load_cluster_from_dcs = Mock(side_effect=DCSError('Etcd is not responding properly'))
self.assertEqual(self.ha.run_cycle(), 'demoted self because DCS is not accessible and i was a leader')
self.assertEquals(self.ha.run_cycle(), 'demoted self because DCS is not accessible and i was a leader')
@patch('time.sleep', Mock())
def test_bootstrap_from_another_member(self):
self.ha.cluster = get_cluster_initialized_with_leader()
self.assertEqual(self.ha.bootstrap(), 'trying to bootstrap from replica \'other\'')
self.assertEquals(self.ha.bootstrap(), 'trying to bootstrap from replica \'other\'')
def test_bootstrap_waiting_for_leader(self):
self.ha.cluster = get_cluster_initialized_without_leader()
self.assertEqual(self.ha.bootstrap(), 'waiting for leader to bootstrap')
self.assertEquals(self.ha.bootstrap(), 'waiting for leader to bootstrap')
def test_bootstrap_without_leader(self):
self.ha.cluster = get_cluster_initialized_without_leader()
self.p.can_create_replica_without_replication_connection = MagicMock(return_value=True)
self.assertEqual(self.ha.bootstrap(), 'trying to bootstrap (without leader)')
self.assertEquals(self.ha.bootstrap(), 'trying to bootstrap (without leader)')
def test_bootstrap_initialize_lock_failed(self):
self.ha.cluster = get_cluster_not_initialized_without_leader()
self.assertEqual(self.ha.bootstrap(), 'failed to acquire initialize lock')
self.assertEquals(self.ha.bootstrap(), 'failed to acquire initialize lock')
def test_bootstrap_initialized_new_cluster(self):
self.ha.cluster = get_cluster_not_initialized_without_leader()
self.e.initialize = true
self.assertEqual(self.ha.bootstrap(), 'trying to bootstrap a new cluster')
self.assertEquals(self.ha.bootstrap(), 'trying to bootstrap a new cluster')
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'waiting for end of recovery after bootstrap')
self.assertEquals(self.ha.run_cycle(), 'waiting for end of recovery after bootstrap')
self.p.is_leader = true
self.assertEqual(self.ha.run_cycle(), 'running post_bootstrap')
self.assertEqual(self.ha.run_cycle(), 'initialized a new cluster')
self.assertEquals(self.ha.run_cycle(), 'running post_bootstrap')
self.assertEquals(self.ha.run_cycle(), 'initialized a new cluster')
def test_bootstrap_release_initialize_key_on_failure(self):
self.ha.cluster = get_cluster_not_initialized_without_leader()
@@ -398,7 +434,7 @@ class TestHa(unittest.TestCase):
self.p.is_running.return_value = MockPostmaster()
self.p.is_leader = true
with patch.object(Watchdog, 'activate', Mock(return_value=False)):
self.assertEqual(self.ha.post_bootstrap(), 'running post_bootstrap')
self.assertEquals(self.ha.post_bootstrap(), 'running post_bootstrap')
self.assertRaises(PatroniException, self.ha.post_bootstrap)
@patch('psycopg2.connect', psycopg2_connect)
@@ -415,35 +451,35 @@ class TestHa(unittest.TestCase):
@patch('time.sleep', Mock())
def test_restart(self):
self.assertEqual(self.ha.restart({}), (True, 'restarted successfully'))
self.assertEquals(self.ha.restart({}), (True, 'restarted successfully'))
self.p.restart = Mock(return_value=None)
self.assertEqual(self.ha.restart({}), (False, 'postgres is still starting'))
self.assertEquals(self.ha.restart({}), (False, 'postgres is still starting'))
self.p.restart = false
self.assertEqual(self.ha.restart({}), (False, 'restart failed'))
self.assertEquals(self.ha.restart({}), (False, 'restart failed'))
self.ha.cluster = get_cluster_initialized_with_leader()
self.ha.reinitialize()
self.assertEqual(self.ha.restart({}), (False, 'reinitialize already in progress'))
self.assertEquals(self.ha.restart({}), (False, 'reinitialize already in progress'))
with patch.object(self.ha, "restart_matches", return_value=False):
self.assertEqual(self.ha.restart({'foo': 'bar'}), (False, "restart conditions are not satisfied"))
self.assertEquals(self.ha.restart({'foo': 'bar'}), (False, "restart conditions are not satisfied"))
@patch('os.kill', Mock())
def test_restart_in_progress(self):
with patch('patroni.async_executor.AsyncExecutor.busy', PropertyMock(return_value=True)):
self.ha.restart({}, run_async=True)
self.assertTrue(self.ha.restart_scheduled())
self.assertEqual(self.ha.run_cycle(), 'restart in progress')
self.assertEquals(self.ha.run_cycle(), 'restart in progress')
self.ha.cluster = get_cluster_initialized_with_leader()
self.assertEqual(self.ha.run_cycle(), 'restart in progress')
self.assertEquals(self.ha.run_cycle(), 'restart in progress')
self.ha.has_lock = true
self.assertEqual(self.ha.run_cycle(), 'updated leader lock during restart')
self.assertEquals(self.ha.run_cycle(), 'updated leader lock during restart')
self.ha.update_lock = false
self.p.set_role('master')
with patch('patroni.async_executor.CriticalTask.cancel', Mock(return_value=False)):
with patch('patroni.postgresql.Postgresql.terminate_starting_postmaster') as mock_terminate:
self.assertEqual(self.ha.run_cycle(), 'lost leader lock during restart')
self.assertEquals(self.ha.run_cycle(), 'lost leader lock during restart')
mock_terminate.assert_called()
@patch('requests.get', requests_get)
@@ -451,27 +487,25 @@ class TestHa(unittest.TestCase):
self.ha.fetch_node_status = get_node_status()
self.ha.has_lock = true
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, 'blabla', '', None))
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, '', self.p.name, None))
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, '', 'blabla', None))
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
f = Failover(0, self.p.name, '', None)
self.ha.cluster = get_cluster_initialized_with_leader(f)
self.assertEqual(self.ha.run_cycle(), 'manual failover: demoting myself')
self.assertEquals(self.ha.run_cycle(), 'manual failover: demoting myself')
self.p.rewind_needed_and_possible = true
self.assertEqual(self.ha.run_cycle(), 'manual failover: demoting myself')
self.assertEquals(self.ha.run_cycle(), 'manual failover: demoting myself')
self.ha.fetch_node_status = get_node_status(nofailover=True)
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.ha.fetch_node_status = get_node_status(watchdog_failed=True)
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.ha.fetch_node_status = get_node_status(timeline=1)
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.ha.fetch_node_status = get_node_status(wal_position=1)
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
# manual failover from the previous leader to us won't happen if we hold the nofailover flag
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, 'blabla', self.p.name, None))
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
# Failover scheduled time must include timezone
scheduled = datetime.datetime.now()
@@ -480,19 +514,19 @@ class TestHa(unittest.TestCase):
scheduled = datetime.datetime.utcnow().replace(tzinfo=tzutc)
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, 'blabla', self.p.name, scheduled))
self.assertEqual('no action. i am the leader with the lock', self.ha.run_cycle())
self.assertEquals('no action. i am the leader with the lock', self.ha.run_cycle())
scheduled = scheduled + datetime.timedelta(seconds=30)
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, 'blabla', self.p.name, scheduled))
self.assertEqual('no action. i am the leader with the lock', self.ha.run_cycle())
self.assertEquals('no action. i am the leader with the lock', self.ha.run_cycle())
scheduled = scheduled + datetime.timedelta(seconds=-600)
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, 'blabla', self.p.name, scheduled))
self.assertEqual('no action. i am the leader with the lock', self.ha.run_cycle())
self.assertEquals('no action. i am the leader with the lock', self.ha.run_cycle())
scheduled = None
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, 'blabla', self.p.name, scheduled))
self.assertEqual('no action. i am the leader with the lock', self.ha.run_cycle())
self.assertEquals('no action. i am the leader with the lock', self.ha.run_cycle())
@patch('requests.get', requests_get)
def test_manual_failover_from_leader_in_pause(self):
@@ -500,9 +534,9 @@ class TestHa(unittest.TestCase):
self.ha.is_paused = true
scheduled = datetime.datetime.now()
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, 'blabla', self.p.name, scheduled))
self.assertEqual('PAUSE: no action. i am the leader with the lock', self.ha.run_cycle())
self.assertEquals('PAUSE: no action. i am the leader with the lock', self.ha.run_cycle())
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, self.p.name, '', None))
self.assertEqual('PAUSE: no action. i am the leader with the lock', self.ha.run_cycle())
self.assertEquals('PAUSE: no action. i am the leader with the lock', self.ha.run_cycle())
@patch('requests.get', requests_get)
def test_manual_failover_from_leader_in_synchronous_mode(self):
@@ -512,48 +546,48 @@ class TestHa(unittest.TestCase):
self.ha.is_failover_possible = false
self.ha.process_sync_replication = Mock()
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, self.p.name, 'a', None), (self.p.name, None))
self.assertEqual('no action. i am the leader with the lock', self.ha.run_cycle())
self.assertEquals('no action. i am the leader with the lock', self.ha.run_cycle())
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, self.p.name, 'a', None), (self.p.name, 'a'))
self.ha.is_failover_possible = true
self.assertEqual('manual failover: demoting myself', self.ha.run_cycle())
self.assertEquals('manual failover: demoting myself', self.ha.run_cycle())
@patch('requests.get', requests_get)
def test_manual_failover_process_no_leader(self):
self.p.is_leader = false
self.ha.cluster = get_cluster_initialized_without_leader(failover=Failover(0, '', self.p.name, None))
self.assertEqual(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
self.assertEquals(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
self.ha.cluster = get_cluster_initialized_without_leader(failover=Failover(0, '', 'leader', None))
self.p.set_role('replica')
self.assertEqual(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
self.assertEquals(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
self.ha.fetch_node_status = get_node_status() # accessible, in_recovery
self.assertEqual(self.ha.run_cycle(), 'following a different leader because i am not the healthiest node')
self.assertEquals(self.ha.run_cycle(), 'following a different leader because i am not the healthiest node')
self.ha.cluster = get_cluster_initialized_without_leader(failover=Failover(0, self.p.name, '', None))
self.assertEqual(self.ha.run_cycle(), 'following a different leader because i am not the healthiest node')
self.assertEquals(self.ha.run_cycle(), 'following a different leader because i am not the healthiest node')
self.ha.fetch_node_status = get_node_status(reachable=False) # inaccessible, in_recovery
self.p.set_role('replica')
self.assertEqual(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
self.assertEquals(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
# set failover flag to True for all members of the cluster
# this should elect the current member, as we are not going to call the API for it.
self.ha.cluster = get_cluster_initialized_without_leader(failover=Failover(0, '', 'other', None))
self.ha.fetch_node_status = get_node_status(nofailover=True) # accessible, in_recovery
self.p.set_role('replica')
self.assertEqual(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
self.assertEquals(self.ha.run_cycle(), 'promoted self to leader by acquiring session lock')
# same as previous, but set the current member to nofailover. In no case it should be elected as a leader
self.ha.patroni.nofailover = True
self.assertEqual(self.ha.run_cycle(), 'following a different leader because I am not allowed to promote')
self.assertEquals(self.ha.run_cycle(), 'following a different leader because I am not allowed to promote')
def test_manual_failover_process_no_leader_in_pause(self):
self.ha.is_paused = true
self.ha.cluster = get_cluster_initialized_without_leader(failover=Failover(0, '', 'other', None))
self.assertEqual(self.ha.run_cycle(), 'PAUSE: continue to run as master without lock')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: continue to run as master without lock')
self.ha.cluster = get_cluster_initialized_without_leader(failover=Failover(0, 'leader', '', None))
self.assertEqual(self.ha.run_cycle(), 'PAUSE: continue to run as master without lock')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: continue to run as master without lock')
self.ha.cluster = get_cluster_initialized_without_leader(failover=Failover(0, 'leader', 'blabla', None))
self.assertEqual('PAUSE: acquired session lock as a leader', self.ha.run_cycle())
self.assertEquals('PAUSE: acquired session lock as a leader', self.ha.run_cycle())
self.p.is_leader = false
self.p.set_role('replica')
self.ha.cluster = get_cluster_initialized_without_leader(failover=Failover(0, 'leader', self.p.name, None))
self.assertEqual(self.ha.run_cycle(), 'PAUSE: promoted self to leader by acquiring session lock')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: promoted self to leader by acquiring session lock')
def test_is_healthiest_node(self):
self.ha.state_handler.is_leader = false
@@ -578,8 +612,6 @@ class TestHa(unittest.TestCase):
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
with patch('patroni.postgresql.Postgresql.timeline_wal_position', return_value=(1, 1)):
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
with patch('patroni.postgresql.Postgresql.replica_cached_timeline', return_value=1):
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
self.ha.patroni.nofailover = True
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
self.ha.patroni.nofailover = False
@@ -636,7 +668,7 @@ class TestHa(unittest.TestCase):
def test_scheduled_restart(self):
self.ha.cluster = get_cluster_initialized_with_leader()
with patch.object(self.ha, "evaluate_scheduled_restart", Mock(return_value="restart scheduled")):
self.assertEqual(self.ha.run_cycle(), "restart scheduled")
self.assertEquals(self.ha.run_cycle(), "restart scheduled")
def test_restart_matches(self):
self.p._role = 'replica'
@@ -653,79 +685,84 @@ class TestHa(unittest.TestCase):
self.ha.is_paused = true
self.p.name = 'leader'
self.ha.cluster = get_cluster_initialized_with_leader()
self.assertEqual(self.ha.run_cycle(), 'PAUSE: removed leader lock because postgres is not running as master')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: removed leader lock because postgres is not running as master')
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, '', self.p.name, None))
self.assertEqual(self.ha.run_cycle(), 'PAUSE: waiting to become master after promote...')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: waiting to become master after promote...')
def test_process_healthy_standby_cluster_as_standby_leader(self):
self.p.is_leader = false
self.p.name = 'leader'
self.ha.patroni.config._dynamic_configuration = {"standby_cluster": {
"host": "localhost",
"port": 5432,
"primary_slot_name": "",
}}
self.ha.cluster = get_standby_cluster_initialized_with_only_leader()
msg = 'no action. i am the standby leader with the lock'
self.assertEqual(self.ha.run_cycle(), msg)
self.assertEquals(self.ha.run_cycle(), msg)
def test_process_healthy_standby_cluster_as_cascade_replica(self):
self.p.is_leader = false
self.p.name = 'replica'
self.ha.patroni.config._dynamic_configuration = {"standby_cluster": {
"host": "localhost",
"port": 5432,
"primary_slot_name": "",
}}
self.ha.cluster = get_standby_cluster_initialized_with_only_leader()
msg = 'no action. i am a secondary and i am following a leader'
self.assertEqual(self.ha.run_cycle(), msg)
self.assertEquals(self.ha.run_cycle(), msg)
def test_process_unhealthy_standby_cluster_as_standby_leader(self):
@patch('patroni.dcs.etcd.Etcd.initialize', return_value=True)
def test_process_unhealthy_standby_cluster_as_standby_leader(self, initialize):
self.p.is_leader = false
self.p.name = 'leader'
self.ha.patroni.config._dynamic_configuration = {"standby_cluster": {
"host": "localhost",
"port": 5432,
"primary_slot_name": "",
}}
self.ha.cluster = get_standby_cluster_initialized_with_only_leader()
self.ha.cluster.is_unlocked = true
self.ha.sysid_valid = true
self.p._sysid = True
msg = 'promoted self to a standby leader because i had the session lock'
self.assertEqual(self.ha.run_cycle(), msg)
self.assertEquals(self.ha.run_cycle(), msg)
@patch.object(Postgresql, 'rewind_needed_and_possible', Mock(return_value=True))
def test_process_unhealthy_standby_cluster_as_cascade_replica(self):
@patch('patroni.dcs.etcd.Etcd.initialize', return_value=True)
def test_process_unhealthy_standby_cluster_as_cascade_replica(self, initialize):
self.p.is_leader = false
self.p.name = 'replica'
self.ha.patroni.config._dynamic_configuration = {"standby_cluster": {
"host": "localhost",
"port": 5432,
"primary_slot_name": "",
}}
self.ha.cluster = get_standby_cluster_initialized_with_only_leader()
self.ha.is_unlocked = true
self.assertTrue(self.ha.run_cycle().startswith('running pg_rewind from remote_master:'))
def test_recover_unhealthy_leader_in_standby_cluster(self):
self.p.is_leader = false
self.p.name = 'leader'
self.p.is_running = false
self.p.follow = false
self.ha.cluster = get_standby_cluster_initialized_with_only_leader()
self.assertEqual(self.ha.run_cycle(), 'starting as a standby leader because i had the session lock')
def test_recover_unhealthy_unlocked_standby_cluster(self):
self.p.is_leader = false
self.p.name = 'leader'
self.p.is_running = false
self.p.follow = false
self.ha.cluster = get_standby_cluster_initialized_with_only_leader()
self.ha.cluster.is_unlocked = true
self.ha.has_lock = false
self.assertEqual(self.ha.run_cycle(), 'trying to follow a remote master because standby cluster is unhealthy')
msg = 'running pg_rewind from leader'
self.assertEquals(self.ha.run_cycle(), msg)
def test_failed_to_update_lock_in_pause(self):
self.ha.update_lock = false
self.ha.is_paused = true
self.p.name = 'leader'
self.ha.cluster = get_cluster_initialized_with_leader()
self.assertEqual(self.ha.run_cycle(),
'PAUSE: continue to run as master after failing to update leader lock in DCS')
self.assertEquals(self.ha.run_cycle(),
'PAUSE: continue to run as master after failing to update leader lock in DCS')
def test_postgres_unhealthy_in_pause(self):
self.ha.is_paused = true
self.p.is_healthy = false
self.assertEqual(self.ha.run_cycle(), 'PAUSE: postgres is not running')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: postgres is not running')
self.ha.has_lock = true
self.assertEqual(self.ha.run_cycle(), 'PAUSE: removed leader lock because postgres is not running')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: removed leader lock because postgres is not running')
def test_no_etcd_connection_in_pause(self):
self.ha.is_paused = true
self.ha.load_cluster_from_dcs = Mock(side_effect=DCSError('Etcd is not responding properly'))
self.assertEqual(self.ha.run_cycle(), 'PAUSE: DCS is not accessible')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: DCS is not accessible')
@patch('patroni.ha.Ha.update_lock', return_value=True)
@patch('patroni.ha.Ha.demote')
@@ -741,26 +778,26 @@ class TestHa(unittest.TestCase):
self.ha.cluster = get_cluster_initialized_with_leader()
self.p.check_for_startup = true
self.p.time_in_state = lambda: 30
self.assertEqual(self.ha.run_cycle(), 'PostgreSQL is still starting up, 270 seconds until timeout')
self.assertEquals(self.ha.run_cycle(), 'PostgreSQL is still starting up, 270 seconds until timeout')
check_calls([(update_lock, True), (demote, False)])
self.p.time_in_state = lambda: 350
self.ha.fetch_node_status = get_node_status(reachable=False) # inaccessible, in_recovery
self.assertEqual(self.ha.run_cycle(),
'master start has timed out, but continuing to wait because failover is not possible')
self.assertEquals(self.ha.run_cycle(),
'master start has timed out, but continuing to wait because failover is not possible')
check_calls([(update_lock, True), (demote, False)])
self.ha.fetch_node_status = get_node_status() # accessible, in_recovery
self.assertEqual(self.ha.run_cycle(), 'stopped PostgreSQL because of startup timeout')
self.assertEquals(self.ha.run_cycle(), 'stopped PostgreSQL because of startup timeout')
check_calls([(update_lock, True), (demote, True)])
update_lock.return_value = False
self.assertEqual(self.ha.run_cycle(), 'stopped PostgreSQL while starting up because leader key was lost')
self.assertEquals(self.ha.run_cycle(), 'stopped PostgreSQL while starting up because leader key was lost')
check_calls([(update_lock, True), (demote, True)])
self.ha.has_lock = false
self.p.is_leader = false
self.assertEqual(self.ha.run_cycle(), 'no action. i am a secondary and i am following a leader')
self.assertEquals(self.ha.run_cycle(), 'no action. i am a secondary and i am following a leader')
check_calls([(update_lock, False), (demote, False)])
def test_manual_failover_while_starting(self):
@@ -769,7 +806,7 @@ class TestHa(unittest.TestCase):
f = Failover(0, self.p.name, '', None)
self.ha.cluster = get_cluster_initialized_with_leader(f)
self.ha.fetch_node_status = get_node_status() # accessible, in_recovery
self.assertEqual(self.ha.run_cycle(), 'manual failover: demoting myself')
self.assertEquals(self.ha.run_cycle(), 'manual failover: demoting myself')
@patch('patroni.ha.Ha.demote')
def test_failover_immediately_on_zero_master_start_timeout(self, demote):
@@ -780,7 +817,7 @@ class TestHa(unittest.TestCase):
self.ha.has_lock = true
self.ha.update_lock = true
self.ha.fetch_node_status = get_node_status() # accessible, in_recovery
self.assertEqual(self.ha.run_cycle(), 'stopped PostgreSQL to fail over after a crash')
self.assertEquals(self.ha.run_cycle(), 'stopped PostgreSQL to fail over after a crash')
demote.assert_called_once()
@patch('patroni.postgresql.Postgresql.follow')
@@ -842,18 +879,18 @@ class TestHa(unittest.TestCase):
self.p.pick_synchronous_standby = Mock(return_value=('other2', True))
self.ha.run_cycle()
self.ha.dcs.get_cluster.assert_called_once()
self.assertEqual(self.ha.dcs.write_sync_state.call_count, 2)
self.assertEquals(self.ha.dcs.write_sync_state.call_count, 2)
# Test updating sync standby key failed due to race
self.ha.dcs.write_sync_state = Mock(side_effect=[True, False])
self.ha.run_cycle()
self.assertEqual(self.ha.dcs.write_sync_state.call_count, 2)
self.assertEquals(self.ha.dcs.write_sync_state.call_count, 2)
# Test changing sync standby failed due to race
self.ha.dcs.write_sync_state = Mock(return_value=True)
self.ha.dcs.get_cluster = Mock(return_value=get_cluster_initialized_with_leader(sync=('somebodyelse', None)))
self.ha.run_cycle()
self.assertEqual(self.ha.dcs.write_sync_state.call_count, 1)
self.assertEquals(self.ha.dcs.write_sync_state.call_count, 1)
# Test sync set to '*' when synchronous_mode_strict is enabled
mock_set_sync.reset_mock()
@@ -874,7 +911,7 @@ class TestHa(unittest.TestCase):
self.ha.cluster = get_cluster_initialized_with_leader(sync=('other', None))
# When we just became master nobody is sync
self.assertEqual(self.ha.enforce_master_role('msg', 'promote msg'), 'promote msg')
self.assertEquals(self.ha.enforce_master_role('msg', 'promote msg'), 'promote msg')
mock_set_sync.assert_called_once_with(None)
mock_write_sync.assert_called_once_with('leader', None, index=0)
@@ -902,7 +939,7 @@ class TestHa(unittest.TestCase):
self.ha.run_cycle()
mock_acquire.assert_not_called()
mock_follow.assert_called_once()
self.assertEqual(mock_follow.call_args[0][0], None)
self.assertEquals(mock_follow.call_args[0][0], None)
mock_write_sync.assert_not_called()
mock_follow.reset_mock()
@@ -953,15 +990,15 @@ class TestHa(unittest.TestCase):
def test_effective_tags(self):
self.ha._disable_sync = True
self.assertEqual(self.ha.get_effective_tags(), {'foo': 'bar', 'nosync': True})
self.assertEquals(self.ha.get_effective_tags(), {'foo': 'bar', 'nosync': True})
self.ha._disable_sync = False
self.assertEqual(self.ha.get_effective_tags(), {'foo': 'bar'})
self.assertEquals(self.ha.get_effective_tags(), {'foo': 'bar'})
def test_restore_cluster_config(self):
self.ha.cluster.config.data.clear()
self.ha.has_lock = true
self.ha.cluster.is_unlocked = false
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
def test_watch(self):
self.ha.cluster = get_cluster_initialized_with_leader()
@@ -980,19 +1017,19 @@ class TestHa(unittest.TestCase):
self.ha.cluster = get_cluster_initialized_with_leader()
self.ha.has_lock = true
self.p.data_directory_empty = true
self.assertEqual(self.ha.run_cycle(), 'released leader key voluntarily as data dir empty and currently leader')
self.assertEqual(self.p.role, 'uninitialized')
self.assertEquals(self.ha.run_cycle(), 'released leader key voluntarily as data dir empty and currently leader')
self.assertEquals(self.p.role, 'uninitialized')
# as has_lock is mocked out, we need to fake the leader key release
self.ha.has_lock = false
# will not say bootstrap from leader as replica can't self elect
self.assertEqual(self.ha.run_cycle(), "trying to bootstrap from replica 'other'")
self.assertEquals(self.ha.run_cycle(), "trying to bootstrap from replica 'other'")
def test_update_cluster_history(self):
self.p.get_master_timeline = Mock(return_value=1)
self.ha.has_lock = true
self.ha.cluster.is_unlocked = false
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
@patch('sys.exit', return_value=1)
def test_abort_join(self, exit_mock):
@@ -1005,15 +1042,6 @@ class TestHa(unittest.TestCase):
self.ha.has_lock = true
self.ha.cluster.is_unlocked = false
self.ha.is_paused = true
self.assertEqual(self.ha.run_cycle(), 'PAUSE: no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'PAUSE: no action. i am the leader with the lock')
self.ha.is_paused = false
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
@patch('psycopg2.connect', psycopg2_connect)
def test_permanent_logical_slots_after_promote(self):
config = ClusterConfig(1, {'slots': {'l': {'database': 'postgres', 'plugin': 'test_decoding'}}}, 1)
self.ha.cluster = get_cluster_initialized_without_leader(cluster_config=config)
self.assertEqual(self.ha.run_cycle(), 'acquired session lock as a leader')
self.ha.cluster = get_cluster_initialized_without_leader(leader=True, cluster_config=config)
self.ha.has_lock = true
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.assertEquals(self.ha.run_cycle(), 'no action. i am the leader with the lock')
-7
View File
@@ -50,13 +50,6 @@ class TestKubernetes(unittest.TestCase):
'labels': {'f': 'b'}, 'use_endpoints': True, 'pod_ip': '10.0.0.0'})
self.assertIsNotNone(k.update_leader('123'))
@patch('kubernetes.config.load_kube_config', Mock())
@patch.object(k8s_client.CoreV1Api, 'create_namespaced_endpoints', Mock())
def test_update_leader_with_restricted_access(self):
k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'retry_timeout': 10,
'labels': {'f': 'b'}, 'use_endpoints': True, 'pod_ip': '10.0.0.0'})
self.assertIsNotNone(k.update_leader('123', True))
def test_take_leader(self):
self.k.take_leader()
self.k._leader_observed_record['leader'] = 'test'
-38
View File
@@ -1,38 +0,0 @@
import os
import sys
import unittest
import yaml
from mock import Mock, patch
from patroni.config import Config
from patroni.log import PatroniLogger
class TestPatroniLogger(unittest.TestCase):
@patch('logging.FileHandler._open', Mock())
def setUp(self):
self.config = {
'log': {
'dir': 'foo',
'file_size': 4096,
'file_num': 5,
'loggers': {
'foo.bar': 'INFO'
}
},
'restapi': {}, 'postgresql': {'data_dir': 'foo'}
}
sys.argv = ['patroni.py']
os.environ[Config.PATRONI_CONFIG_VARIABLE] = yaml.dump(self.config, default_flow_style=False)
self.logger = PatroniLogger()
config = Config()
self.logger.reload_config(config['log'])
def test_rotating_handler(self):
self.assertEqual(self.logger.handler.maxBytes, self.config['log']['file_size'])
self.assertEqual(self.logger.handler.backupCount, self.config['log']['file_num'])
def test_reload_config(self):
self.config['log'].pop('dir')
self.logger.reload_config(self.config['log'])
+1 -1
View File
@@ -73,7 +73,7 @@ class TestPatroni(unittest.TestCase):
mock_getpid.return_value = 2
_main()
with patch('sys.frozen', Mock(return_value=True), create=True), patch('os.setsid', Mock()):
with patch('sys.frozen', Mock(return_value=True), create=True):
sys.argv = ['/patroni', 'pg_ctl_start', 'postgres', '-D', '/data', '--max_connections=100']
_main()
+70 -78
View File
@@ -8,7 +8,7 @@ import unittest
from mock import Mock, MagicMock, PropertyMock, patch, mock_open
from patroni.async_executor import CriticalTask
from patroni.dcs import Cluster, ClusterConfig, Leader, Member, RemoteMember, SyncState
from patroni.dcs import Cluster, Leader, Member, RemoteMember, SyncState
from patroni.exceptions import PostgresConnectionException, PostgresException
from patroni.postgresql import Postgresql, STATE_REJECT, STATE_NO_RESPONSE
from patroni.postmaster import PostmasterProcess
@@ -28,15 +28,15 @@ class MockCursor(object):
def execute(self, sql, *params):
if sql.startswith('blabla'):
raise psycopg2.ProgrammingError()
elif sql == 'CHECKPOINT' or sql.startswith('SELECT pg_catalog.pg_create_'):
elif sql == 'CHECKPOINT':
raise psycopg2.OperationalError()
elif sql.startswith('RetryFailedError'):
raise RetryFailedError('retry')
elif sql.startswith('SELECT slot_name'):
self.results = [('blabla', 'physical'), ('foobar', 'physical'), ('ls', 'logical', 'a', 'b')]
elif sql.startswith('SELECT CASE WHEN pg_catalog.pg_is_in_recovery()'):
self.results = [('blabla',), ('foobar',)]
elif sql.startswith('SELECT CASE WHEN pg_is_in_recovery()'):
self.results = [(1, 2)]
elif sql.startswith('SELECT pg_catalog.pg_is_in_recovery()'):
elif sql.startswith('SELECT pg_is_in_recovery()'):
self.results = [(False, 2)]
elif sql.startswith('WITH replication_info AS ('):
replication_info = '[{"application_name":"walreceiver","client_addr":"1.2.3.4",' +\
@@ -53,7 +53,7 @@ class MockCursor(object):
self.results = [('1', 2, '0/402EEC0', '')]
elif sql.startswith('SELECT isdir, modification'):
self.results = [(False, datetime.datetime.now())]
elif sql.startswith('SELECT pg_catalog.pg_read_file'):
elif sql.startswith('SELECT pg_read_file'):
self.results = [('1\t0/40159C0\tno recovery target specified\n\n' +
'2\t1/40159C0\tno recovery target specified\n',)]
elif sql.startswith('TIMELINE_HISTORY '):
@@ -308,7 +308,7 @@ class TestPostgresql(unittest.TestCase):
def test_restart(self):
self.p.start = Mock(return_value=False)
self.assertFalse(self.p.restart())
self.assertEqual(self.p.state, 'restart failed (restarting)')
self.assertEquals(self.p.state, 'restart failed (restarting)')
@patch.object(builtins, 'open', MagicMock())
def test_write_pgpass(self):
@@ -317,10 +317,10 @@ class TestPostgresql(unittest.TestCase):
def test_checkpoint(self):
with patch.object(MockCursor, 'fetchone', Mock(return_value=(True, ))):
self.assertEqual(self.p.checkpoint({'user': 'postgres'}), 'is_in_recovery=true')
self.assertEquals(self.p.checkpoint({'user': 'postgres'}), 'is_in_recovery=true')
with patch.object(MockCursor, 'execute', Mock(return_value=None)):
self.assertIsNone(self.p.checkpoint())
self.assertEqual(self.p.checkpoint(), 'not accessible or not healty')
self.assertEquals(self.p.checkpoint(), 'not accessible or not healty')
@patch.object(Postgresql, 'cancellable_subprocess_call')
@patch('patroni.postgresql.Postgresql.write_pgpass', MagicMock(return_value=dict()))
@@ -401,8 +401,7 @@ class TestPostgresql(unittest.TestCase):
@patch.object(Postgresql, 'is_running', Mock(return_value=False))
@patch.object(Postgresql, 'start', Mock())
def test_follow(self):
m = RemoteMember('1', {'restore_command': '2', 'recovery_min_apply_delay': 3, 'archive_cleanup_command': '4'})
self.p.follow(m)
self.p.follow(RemoteMember('123', {'recovery_command': 'foo'}))
@patch('subprocess.check_output', Mock(return_value=0, side_effect=pg_controldata_string))
def test_can_rewind(self):
@@ -421,20 +420,16 @@ class TestPostgresql(unittest.TestCase):
def test_create_replica(self, mock_cancellable_subprocess_call):
self.p.delete_trigger_file = Mock(side_effect=OSError)
self.p.config['create_replica_methods'] = ['pgBackRest']
self.p.config['pgBackRest'] = {'command': 'pgBackRest', 'keep_data': True, 'no_params': True}
mock_cancellable_subprocess_call.return_value = 0
self.assertEqual(self.p.create_replica(self.leader), 0)
self.p.config['create_replica_methods'] = ['wale', 'basebackup']
self.p.config['wale'] = {'command': 'foo'}
self.assertEqual(self.p.create_replica(self.leader), 0)
mock_cancellable_subprocess_call.return_value = 0
self.assertEquals(self.p.create_replica(self.leader), 0)
del self.p.config['wale']
self.assertEqual(self.p.create_replica(self.leader), 0)
self.assertEquals(self.p.create_replica(self.leader), 0)
self.p.config['create_replica_methods'] = ['basebackup']
self.p.config['basebackup'] = [{'max_rate': '100M'}, 'no-sync']
self.assertEqual(self.p.create_replica(self.leader), 0)
self.assertEquals(self.p.create_replica(self.leader), 0)
self.p.config['basebackup'] = [{'max_rate': '100M', 'compress': '9'}]
with mock.patch('patroni.postgresql.logger.error', new_callable=Mock()) as mock_logger:
@@ -451,24 +446,24 @@ class TestPostgresql(unittest.TestCase):
"not matching {0}".format(mock_logger.call_args[0][0]))
self.p.config['basebackup'] = {"foo": "bar"}
self.assertEqual(self.p.create_replica(self.leader), 0)
self.assertEquals(self.p.create_replica(self.leader), 0)
self.p.config['create_replica_methods'] = ['wale', 'basebackup']
del self.p.config['basebackup']
mock_cancellable_subprocess_call.return_value = 1
self.assertEqual(self.p.create_replica(self.leader), 1)
self.assertEquals(self.p.create_replica(self.leader), 1)
mock_cancellable_subprocess_call.side_effect = Exception('foo')
self.assertEqual(self.p.create_replica(self.leader), 1)
self.assertEquals(self.p.create_replica(self.leader), 1)
mock_cancellable_subprocess_call.side_effect = [1, 0]
self.assertEqual(self.p.create_replica(self.leader), 0)
self.assertEquals(self.p.create_replica(self.leader), 0)
mock_cancellable_subprocess_call.side_effect = [Exception(), 0]
self.assertEqual(self.p.create_replica(self.leader), 0)
self.assertEquals(self.p.create_replica(self.leader), 0)
self.p.cancel()
self.assertEqual(self.p.create_replica(self.leader), 1)
self.assertEquals(self.p.create_replica(self.leader), 1)
@patch('time.sleep', Mock())
@patch.object(Postgresql, 'cancellable_subprocess_call')
@@ -482,18 +477,18 @@ class TestPostgresql(unittest.TestCase):
self.p.config['create_replica_method'] = ['wale', 'basebackup']
self.p.config['wale'] = {'command': 'foo'}
mock_cancellable_subprocess_call.return_value = 0
self.assertEqual(self.p.create_replica(self.leader), 0)
self.assertEquals(self.p.create_replica(self.leader), 0)
del self.p.config['wale']
self.assertEqual(self.p.create_replica(self.leader), 0)
self.assertEquals(self.p.create_replica(self.leader), 0)
self.p.config['create_replica_method'] = ['basebackup']
self.p.config['basebackup'] = [{'max_rate': '100M'}, 'no-sync']
self.assertEqual(self.p.create_replica(self.leader), 0)
self.assertEquals(self.p.create_replica(self.leader), 0)
self.p.config['create_replica_method'] = ['wale', 'basebackup']
del self.p.config['basebackup']
mock_cancellable_subprocess_call.return_value = 1
self.assertEqual(self.p.create_replica(self.leader), 1)
self.assertEquals(self.p.create_replica(self.leader), 1)
def test_basebackup(self):
self.p.cancel()
@@ -502,25 +497,23 @@ class TestPostgresql(unittest.TestCase):
@patch.object(Postgresql, 'is_running', Mock(return_value=True))
def test_sync_replication_slots(self):
self.p.start()
config = ClusterConfig(1, {'slots': {'ls': {'database': 'a', 'plugin': 'b'},
'A': 0, 'test_3': 0, 'b': {'type': 'logical', 'plugin': '1'}}}, 1)
cluster = Cluster(True, config, self.leader, 0, [self.me, self.other, self.leadermem], None, None, None)
cluster = Cluster(True, None, self.leader, 0, [self.me, self.other, self.leadermem], None, None, None)
with mock.patch('patroni.postgresql.Postgresql._query', Mock(side_effect=psycopg2.OperationalError)):
self.p.sync_replication_slots(cluster)
self.p.sync_replication_slots(cluster)
with mock.patch('patroni.postgresql.Postgresql.role', new_callable=PropertyMock(return_value='replica')):
self.p.sync_replication_slots(cluster)
with patch.object(Postgresql, 'drop_replication_slot', Mock(return_value=True)),\
patch('patroni.dcs.logger.error', new_callable=Mock()) as errorlog_mock:
with mock.patch('patroni.postgresql.logger.error', new_callable=Mock()) as errorlog_mock:
self.p.query = Mock()
alias1 = Member(0, 'test-3', 28, {'conn_url': 'postgres://replicator:[email protected]:5436/postgres'})
alias2 = Member(0, 'test.3', 28, {'conn_url': 'postgres://replicator:[email protected]:5436/postgres'})
cluster.members.extend([alias1, alias2])
self.p.sync_replication_slots(cluster)
self.assertEqual(errorlog_mock.call_count, 5)
ca = errorlog_mock.call_args_list[0][0][1]
self.assertTrue("test-3" in ca, "non matching {0}".format(ca))
self.assertTrue("test.3" in ca, "non matching {0}".format(ca))
errorlog_mock.assert_called_once()
self.assertTrue("test-3" in errorlog_mock.call_args[0][1],
"non matching {0}".format(errorlog_mock.call_args[0][1]))
self.assertTrue("test.3" in errorlog_mock.call_args[0][1],
"non matching {0}".format(errorlog_mock.call_args[0][1]))
@patch.object(MockCursor, 'execute', Mock(side_effect=psycopg2.OperationalError))
def test__query(self):
@@ -556,25 +549,25 @@ class TestPostgresql(unittest.TestCase):
self.assertTrue(self.p.promote(0))
def test_timeline_wal_position(self):
self.assertEqual(self.p.timeline_wal_position(), (1, 2))
self.assertEquals(self.p.timeline_wal_position(), (1, 2))
Thread(target=self.p.timeline_wal_position).start()
@patch.object(PostmasterProcess, 'from_pidfile')
def test_is_running(self, mock_frompidfile):
# Cached postmaster running
mock_postmaster = self.p._postmaster_proc = MockPostmaster()
self.assertEqual(self.p.is_running(), mock_postmaster)
self.assertEquals(self.p.is_running(), mock_postmaster)
# Cached postmaster not running, no postmaster running
mock_postmaster.is_running.return_value = False
mock_frompidfile.return_value = None
self.assertEqual(self.p.is_running(), None)
self.assertEqual(self.p._postmaster_proc, None)
self.assertEquals(self.p.is_running(), None)
self.assertEquals(self.p._postmaster_proc, None)
# No cached postmaster, postmaster running
mock_frompidfile.return_value = mock_postmaster2 = MockPostmaster()
self.assertEqual(self.p.is_running(), mock_postmaster2)
self.assertEqual(self.p._postmaster_proc, mock_postmaster2)
self.assertEquals(self.p.is_running(), mock_postmaster2)
self.assertEquals(self.p._postmaster_proc, mock_postmaster2)
@patch('shlex.split', Mock(side_effect=OSError))
def test_call_nowait(self):
@@ -651,7 +644,6 @@ class TestPostgresql(unittest.TestCase):
@patch('time.sleep', Mock())
@patch('os.unlink', Mock())
@patch('os.path.isfile', Mock(return_value=True))
@patch.object(Postgresql, 'run_bootstrap_post_init', Mock(return_value=True))
@patch.object(Postgresql, '_custom_bootstrap', Mock(return_value=True))
@patch.object(Postgresql, 'start', Mock(return_value=True))
@@ -694,13 +686,13 @@ class TestPostgresql(unittest.TestCase):
mock_cancellable_subprocess_call.assert_called()
args, kwargs = mock_cancellable_subprocess_call.call_args
self.assertTrue('PGPASSFILE' in kwargs['env'])
self.assertEqual(args[0], ['/bin/false', 'postgres://127.0.0.2:5432/postgres'])
self.assertEquals(args[0], ['/bin/false', 'postgres://127.0.0.2:5432/postgres'])
mock_cancellable_subprocess_call.reset_mock()
self.p._local_address.pop('host')
self.assertTrue(self.p.run_bootstrap_post_init({'post_init': '/bin/false'}))
mock_cancellable_subprocess_call.assert_called()
self.assertEqual(mock_cancellable_subprocess_call.call_args[0][0], ['/bin/false', 'postgres://:5432/postgres'])
self.assertEquals(mock_cancellable_subprocess_call.call_args[0][0], ['/bin/false', 'postgres://:5432/postgres'])
mock_cancellable_subprocess_call.side_effect = OSError
self.assertFalse(self.p.run_bootstrap_post_init({'post_init': '/bin/false'}))
@@ -712,7 +704,7 @@ class TestPostgresql(unittest.TestCase):
@patch('os.listdir', Mock(return_value=['recovery.conf']))
@patch('os.path.exists', Mock(return_value=True))
def test_get_postgres_role_from_data_directory(self):
self.assertEqual(self.p.get_postgres_role_from_data_directory(), 'replica')
self.assertEquals(self.p.get_postgres_role_from_data_directory(), 'replica')
def test_remove_data_directory(self):
self.p.remove_data_directory()
@@ -727,13 +719,13 @@ class TestPostgresql(unittest.TestCase):
def test_controldata(self):
with patch('subprocess.check_output', Mock(return_value=0, side_effect=pg_controldata_string)):
data = self.p.controldata()
self.assertEqual(len(data), 50)
self.assertEqual(data['Database cluster state'], 'shut down in recovery')
self.assertEqual(data['wal_log_hints setting'], 'on')
self.assertEqual(int(data['Database block size']), 8192)
self.assertEquals(len(data), 50)
self.assertEquals(data['Database cluster state'], 'shut down in recovery')
self.assertEquals(data['wal_log_hints setting'], 'on')
self.assertEquals(int(data['Database block size']), 8192)
with patch('subprocess.check_output', Mock(side_effect=subprocess.CalledProcessError(1, ''))):
self.assertEqual(self.p.controldata(), {})
self.assertEquals(self.p.controldata(), {})
@patch('patroni.postgresql.Postgresql._version_file_exists', Mock(return_value=True))
@patch('subprocess.check_output', MagicMock(return_value=0, side_effect=pg_controldata_string))
@@ -787,9 +779,9 @@ class TestPostgresql(unittest.TestCase):
@patch.object(Postgresql, '_version_file_exists', Mock(return_value=True))
def test_get_major_version(self):
with patch.object(builtins, 'open', mock_open(read_data='9.4')):
self.assertEqual(self.p.get_major_version(), 90400)
self.assertEquals(self.p.get_major_version(), 90400)
with patch.object(builtins, 'open', Mock(side_effect=Exception)):
self.assertEqual(self.p.get_major_version(), 0)
self.assertEquals(self.p.get_major_version(), 0)
def test_postmaster_start_time(self):
with patch.object(MockCursor, "fetchone", Mock(return_value=('foo', True, '', '', '', '', False))):
@@ -801,31 +793,31 @@ class TestPostgresql(unittest.TestCase):
with patch('subprocess.call', return_value=0):
self.p._state = 'starting'
self.assertFalse(self.p.check_for_startup())
self.assertEqual(self.p.state, 'running')
self.assertEquals(self.p.state, 'running')
with patch('subprocess.call', return_value=1):
self.p._state = 'starting'
self.assertTrue(self.p.check_for_startup())
self.assertEqual(self.p.state, 'starting')
self.assertEquals(self.p.state, 'starting')
with patch('subprocess.call', return_value=2):
self.p._state = 'starting'
self.assertFalse(self.p.check_for_startup())
self.assertEqual(self.p.state, 'start failed')
self.assertEquals(self.p.state, 'start failed')
with patch('subprocess.call', return_value=0):
self.p._state = 'running'
self.assertFalse(self.p.check_for_startup())
self.assertEqual(self.p.state, 'running')
self.assertEquals(self.p.state, 'running')
with patch('subprocess.call', return_value=127):
self.p._state = 'running'
self.assertFalse(self.p.check_for_startup())
self.assertEqual(self.p.state, 'running')
self.assertEquals(self.p.state, 'running')
self.p._state = 'starting'
self.assertFalse(self.p.check_for_startup())
self.assertEqual(self.p.state, 'running')
self.assertEquals(self.p.state, 'running')
def test_wait_for_startup(self):
state = {'sleeps': 0, 'num_rejects': 0, 'final_return': 0}
@@ -850,12 +842,12 @@ class TestPostgresql(unittest.TestCase):
self.p._state = 'stopped'
self.assertTrue(self.p.wait_for_startup())
self.assertEqual(state['sleeps'], 0)
self.assertEquals(state['sleeps'], 0)
self.p._state = 'starting'
state['num_rejects'] = 5
self.assertTrue(self.p.wait_for_startup())
self.assertEqual(state['sleeps'], 5)
self.assertEquals(state['sleeps'], 5)
self.p._state = 'starting'
state['sleeps'] = 0
@@ -866,7 +858,7 @@ class TestPostgresql(unittest.TestCase):
state['sleeps'] = 0
state['final_return'] = 0
self.assertFalse(self.p.wait_for_startup(timeout=2))
self.assertEqual(state['sleeps'], 3)
self.assertEquals(state['sleeps'], 3)
with patch.object(Postgresql, 'check_startup_state_changed', Mock(return_value=False)):
self.p.cancel()
@@ -882,30 +874,30 @@ class TestPostgresql(unittest.TestCase):
(self.me.name, 'streaming', 'async'),
(self.other.name, 'streaming', 'async'),
]):
self.assertEqual(self.p.pick_synchronous_standby(cluster), (self.leadermem.name, True))
self.assertEquals(self.p.pick_synchronous_standby(cluster), (self.leadermem.name, True))
with patch.object(Postgresql, "query", return_value=[
(self.me.name, 'streaming', 'async'),
(self.leadermem.name, 'streaming', 'potential'),
(self.other.name, 'streaming', 'async'),
]):
self.assertEqual(self.p.pick_synchronous_standby(cluster), (self.leadermem.name, False))
self.assertEquals(self.p.pick_synchronous_standby(cluster), (self.leadermem.name, False))
with patch.object(Postgresql, "query", return_value=[
(self.me.name, 'streaming', 'async'),
(self.other.name, 'streaming', 'async'),
]):
self.assertEqual(self.p.pick_synchronous_standby(cluster), (self.me.name, False))
self.assertEquals(self.p.pick_synchronous_standby(cluster), (self.me.name, False))
with patch.object(Postgresql, "query", return_value=[
('missing', 'streaming', 'sync'),
(self.me.name, 'streaming', 'async'),
(self.other.name, 'streaming', 'async'),
]):
self.assertEqual(self.p.pick_synchronous_standby(cluster), (self.me.name, False))
self.assertEquals(self.p.pick_synchronous_standby(cluster), (self.me.name, False))
with patch.object(Postgresql, "query", return_value=[]):
self.assertEqual(self.p.pick_synchronous_standby(cluster), (None, False))
self.assertEquals(self.p.pick_synchronous_standby(cluster), (None, False))
def test_set_sync_standby(self):
def value_in_conf():
@@ -916,22 +908,22 @@ class TestPostgresql(unittest.TestCase):
mock_reload = self.p.reload = Mock()
self.p.set_synchronous_standby('n1')
self.assertEqual(value_in_conf(), "synchronous_standby_names = 'n1'")
self.assertEquals(value_in_conf(), "synchronous_standby_names = 'n1'")
mock_reload.assert_called()
mock_reload.reset_mock()
self.p.set_synchronous_standby('n1')
mock_reload.assert_not_called()
self.assertEqual(value_in_conf(), "synchronous_standby_names = 'n1'")
self.assertEquals(value_in_conf(), "synchronous_standby_names = 'n1'")
self.p.set_synchronous_standby('n2')
mock_reload.assert_called()
self.assertEqual(value_in_conf(), "synchronous_standby_names = 'n2'")
self.assertEquals(value_in_conf(), "synchronous_standby_names = 'n2'")
mock_reload.reset_mock()
self.p.set_synchronous_standby(None)
mock_reload.assert_called()
self.assertEqual(value_in_conf(), None)
self.assertEquals(value_in_conf(), None)
def test_get_server_parameters(self):
config = {'synchronous_mode': True, 'parameters': {'wal_level': 'hot_standby'}, 'listen': '0'}
@@ -965,8 +957,8 @@ class TestPostgresql(unittest.TestCase):
"--wal_log_hints=on" "--max_wal_senders=5" "--max_replication_slots=5"\n')
with patch.object(builtins, 'open', m):
data = self.p.read_postmaster_opts()
self.assertEqual(data['wal_level'], 'hot_standby')
self.assertEqual(int(data['max_replication_slots']), 5)
self.assertEquals(data['wal_level'], 'hot_standby')
self.assertEquals(int(data['max_replication_slots']), 5)
self.assertEqual(data.get('D'), None)
m.side_effect = IOError
@@ -976,7 +968,7 @@ class TestPostgresql(unittest.TestCase):
@patch('subprocess.Popen')
def test_single_user_mode(self, subprocess_popen_mock):
subprocess_popen_mock.return_value.wait.return_value = 0
self.assertEqual(self.p.single_user_mode('CHECKPOINT', {'archive_mode': 'on'}), 0)
self.assertEquals(self.p.single_user_mode('CHECKPOINT', {'archive_mode': 'on'}), 0)
@patch('os.listdir', Mock(side_effect=[OSError, ['a', 'b']]))
@patch('os.unlink', Mock(side_effect=OSError))
@@ -994,10 +986,10 @@ class TestPostgresql(unittest.TestCase):
self.assertTrue(self.p.fix_cluster_state())
def test_replica_cached_timeline(self):
self.assertEqual(self.p.replica_cached_timeline(1), 2)
self.assertEquals(self.p.replica_cached_timeline(1), 2)
def test_get_master_timeline(self):
self.assertEqual(self.p.get_master_timeline(), 1)
self.assertEquals(self.p.get_master_timeline(), 1)
def test_cancellable_subprocess_call(self):
self.p.cancel()
+10 -9
View File
@@ -46,7 +46,7 @@ class TestPostmasterProcess(unittest.TestCase):
@patch('psutil.Process.__init__')
def test_from_pid(self, mock_init):
mock_init.side_effect = psutil.NoSuchProcess(123)
self.assertEqual(PostmasterProcess.from_pid(123), None)
self.assertEquals(PostmasterProcess.from_pid(123), None)
mock_init.side_effect = None
self.assertNotEquals(PostmasterProcess.from_pid(123), None)
@@ -55,19 +55,19 @@ class TestPostmasterProcess(unittest.TestCase):
@patch('psutil.Process.pid', Mock(return_value=123))
def test_signal_stop(self, mock_send_signal):
proc = PostmasterProcess(-123)
self.assertEqual(proc.signal_stop('immediate'), False)
self.assertEquals(proc.signal_stop('immediate'), False)
mock_send_signal.side_effect = [None, psutil.NoSuchProcess(123), psutil.AccessDenied()]
proc = PostmasterProcess(123)
self.assertEqual(proc.signal_stop('immediate'), None)
self.assertEqual(proc.signal_stop('immediate'), True)
self.assertEqual(proc.signal_stop('immediate'), False)
self.assertEquals(proc.signal_stop('immediate'), None)
self.assertEquals(proc.signal_stop('immediate'), True)
self.assertEquals(proc.signal_stop('immediate'), False)
@patch('psutil.Process.__init__', Mock())
@patch('psutil.wait_procs')
def test_wait_for_user_backends_to_close(self, mock_wait):
c1 = Mock()
c1.cmdline = Mock(return_value=["postgres: startup process "])
c1.cmdline = Mock(return_value=["postgres: startup process"])
c2 = Mock()
c2.cmdline = Mock(return_value=["postgres: postgres postgres [local] idle"])
c3 = Mock()
@@ -77,7 +77,8 @@ class TestPostmasterProcess(unittest.TestCase):
self.assertIsNone(proc.wait_for_user_backends_to_close())
mock_wait.assert_called_with([c2])
with patch('psutil.Process.children', Mock(side_effect=psutil.NoSuchProcess(123))):
c3.cmdline = Mock(side_effect=psutil.AccessDenied(123))
with patch('psutil.Process.children', Mock(return_value=[c3])):
proc = PostmasterProcess(123)
self.assertIsNone(proc.wait_for_user_backends_to_close())
@@ -88,11 +89,11 @@ class TestPostmasterProcess(unittest.TestCase):
mock_frompidfile.return_value._is_postmaster_process.return_value = False
mock_frompid.return_value = "proc 123"
mock_popen.return_value.stdout.readline.return_value = '123'
self.assertEqual(PostmasterProcess.start('true', '/tmp', '/tmp/test.conf', []), "proc 123")
self.assertEquals(PostmasterProcess.start('true', '/tmp', '/tmp/test.conf', []), "proc 123")
mock_frompid.assert_called_with(123)
mock_frompidfile.side_effect = psutil.NoSuchProcess(123)
self.assertEqual(PostmasterProcess.start('true', '/tmp', '/tmp/test.conf', []), "proc 123")
self.assertEquals(PostmasterProcess.start('true', '/tmp', '/tmp/test.conf', []), "proc 123")
@patch('psutil.Process.__init__', Mock(side_effect=psutil.NoSuchProcess(123)))
def test_read_postmaster_pidfile(self):
+5 -5
View File
@@ -8,7 +8,7 @@ from patroni.utils import Retry, RetryFailedError, polling_loop
class TestUtils(unittest.TestCase):
def test_polling_loop(self):
self.assertEqual(list(polling_loop(0.001, interval=0.001)), [0])
self.assertEquals(list(polling_loop(0.001, interval=0.001)), [0])
@patch('time.sleep', Mock())
@@ -29,21 +29,21 @@ class TestRetrySleeper(unittest.TestCase):
def test_reset(self):
retry = Retry(delay=0, max_tries=2)
retry(self._fail())
self.assertEqual(retry._attempts, 1)
self.assertEquals(retry._attempts, 1)
retry.reset()
self.assertEqual(retry._attempts, 0)
self.assertEquals(retry._attempts, 0)
def test_too_many_tries(self):
retry = Retry(delay=0)
self.assertRaises(RetryFailedError, retry, self._fail(times=999))
self.assertEqual(retry._attempts, 1)
self.assertEquals(retry._attempts, 1)
def test_maximum_delay(self):
retry = Retry(delay=10, max_tries=100)
retry(self._fail(times=10))
self.assertTrue(retry._cur_delay < 4000, retry._cur_delay)
# gevent's sleep function is picky about the type
self.assertEqual(type(retry._cur_delay), float)
self.assertEquals(type(retry._cur_delay), float)
def test_deadline(self):
retry = Retry(deadline=0.0001)
+11 -11
View File
@@ -73,14 +73,14 @@ class TestWatchdog(unittest.TestCase):
@patch.object(LinuxWatchdogDevice, 'can_be_disabled', PropertyMock(return_value=True))
def test_unsafe_timeout_disable_watchdog_and_exit(self):
watchdog = Watchdog({'ttl': 30, 'loop_wait': 15, 'watchdog': {'mode': 'required', 'safety_margin': -1}})
self.assertEqual(watchdog.activate(), False)
self.assertEqual(watchdog.is_running, False)
self.assertEquals(watchdog.activate(), False)
self.assertEquals(watchdog.is_running, False)
@patch('platform.system', Mock(return_value='Linux'))
@patch.object(LinuxWatchdogDevice, 'get_timeout', Mock(return_value=16))
def test_timeout_does_not_ensure_safe_termination(self):
Watchdog({'ttl': 30, 'loop_wait': 15, 'watchdog': {'mode': 'auto', 'safety_margin': -1}}).activate()
self.assertEqual(len(mock_devices), 2)
self.assertEquals(len(mock_devices), 2)
@patch('platform.system', Mock(return_value='Linux'))
@patch.object(Watchdog, 'is_running', PropertyMock(return_value=False))
@@ -99,29 +99,29 @@ class TestWatchdog(unittest.TestCase):
watchdog = Watchdog({'ttl': 30, 'loop_wait': 10, 'watchdog': {'mode': 'required'}})
watchdog.activate()
self.assertEqual(len(mock_devices), 2)
self.assertEquals(len(mock_devices), 2)
device = mock_devices[-1]
self.assertTrue(device.open)
self.assertEqual(device.timeout, 24)
self.assertEquals(device.timeout, 24)
watchdog.keepalive()
self.assertEqual(len(device.writes), 1)
self.assertEquals(len(device.writes), 1)
watchdog.disable()
self.assertFalse(device.open)
self.assertEqual(device.writes[-1], b'V')
self.assertEquals(device.writes[-1], b'V')
def test_invalid_timings(self):
watchdog = Watchdog({'ttl': 30, 'loop_wait': 20, 'watchdog': {'mode': 'automatic', 'safety_margin': -1}})
watchdog.activate()
self.assertEqual(len(mock_devices), 1)
self.assertEquals(len(mock_devices), 1)
self.assertFalse(watchdog.is_running)
def test_parse_mode(self):
with patch('patroni.watchdog.base.logger.warning', new_callable=Mock()) as warning_mock:
watchdog = Watchdog({'ttl': 30, 'loop_wait': 10, 'watchdog': {'mode': 'bad'}})
self.assertEqual(watchdog.config.mode, 'off')
self.assertEquals(watchdog.config.mode, 'off')
warning_mock.assert_called_once()
@patch('platform.system', Mock(return_value='Unknown'))
@@ -170,7 +170,7 @@ class TestNullWatchdog(unittest.TestCase):
watchdog = NullWatchdog()
self.assertTrue(watchdog.can_be_disabled)
self.assertRaises(WatchdogError, watchdog.set_timeout, 1)
self.assertEqual(watchdog.describe(), 'NullWatchdog')
self.assertEquals(watchdog.describe(), 'NullWatchdog')
self.assertIsInstance(NullWatchdog.from_config({}), NullWatchdog)
@@ -210,7 +210,7 @@ class TestLinuxWatchdogDevice(unittest.TestCase):
self.assertRaises(WatchdogError, self.impl.get_timeout)
self.assertRaises(WatchdogError, self.impl.set_timeout, 10)
# We still try to output a reasonable string even if getting info errors
self.assertEqual(self.impl.describe(), "Linux watchdog device")
self.assertEquals(self.impl.describe(), "Linux watchdog device")
@patch('os.open', Mock(side_effect=OSError))
def test_open(self):
-6
View File
@@ -1,4 +1,3 @@
import select
import six
import unittest
@@ -114,11 +113,6 @@ class TestPatroniSequentialThreadingHandler(unittest.TestCase):
def test_create_connection(self):
self.assertIsNotNone(self.handler.create_connection(()))
self.assertIsNotNone(self.handler.create_connection((), 40))
self.assertIsNotNone(self.handler.create_connection(timeout=40))
@patch.object(SequentialThreadingHandler, 'select', Mock(side_effect=ValueError))
def test_select(self):
self.assertRaises(select.error, self.handler.select)
class TestZooKeeper(unittest.TestCase):