mirror of
https://github.com/outbackdingo/patroni.git
synced 2026-08-26 07:30:14 +00:00
Compare commits
112
Commits
+1
-1
@@ -7,7 +7,7 @@ addons:
|
||||
- expect-dev # for unbuffer
|
||||
env:
|
||||
global:
|
||||
- ETCDVERSION=3.0.17 ZKVERSION=3.4.11 CONSULVERSION=0.7.4
|
||||
- ETCDVERSION=3.0.17 ZKVERSION=3.4.14 CONSULVERSION=0.7.4
|
||||
- PYVERSIONS="2.7 3.5 3.6"
|
||||
- EXCLUDE_BEHAVE="3.5"
|
||||
- BOTO_CONFIG=/doesnotexist
|
||||
|
||||
+3
-1
@@ -25,11 +25,12 @@ RUN set -ex \
|
||||
| grep -Ev '^python3-(sphinx|etcd|consul|kazoo|kubernetes)' \
|
||||
| xargs apt-get install -y vim curl less jq locales haproxy sudo \
|
||||
python3-etcd python3-kazoo python3-pip busybox \
|
||||
net-tools iputils-ping --fix-missing \
|
||||
&& pip3 install dumb-init \
|
||||
\
|
||||
# Cleanup all locales but en_US.UTF-8
|
||||
&& find /usr/share/i18n/charmaps/ -type f ! -name UTF-8.gz -delete \
|
||||
&& find /usr/share/i18n/locales/ -type f ! -name en_US ! -name en_GB ! -name i18n ! -name iso14651_t1 ! -name iso14651_t1_common ! -name 'translit_*' -delete \
|
||||
&& find /usr/share/i18n/locales/ -type f ! -name en_US ! -name en_GB ! -name i18n* ! -name iso14651_t1 ! -name iso14651_t1_common ! -name 'translit_*' -delete \
|
||||
&& echo 'en_US.UTF-8 UTF-8' > /usr/share/i18n/SUPPORTED \
|
||||
\
|
||||
# Make sure we have a en_US.UTF-8 locale available
|
||||
@@ -149,6 +150,7 @@ RUN sed -i 's/env python/&3/' /patroni*.py \
|
||||
&& sed -i 's/^ \(replication\|superuser\|rewind\|unix_socket_directories\|\(\( \)\{0,1\}\(username\|password\)\)\):/#&/' postgres?.yml \
|
||||
&& sed -i 's/^ parameters:/ pg_hba:\n - local all all trust\n - host replication all all md5\n - host all all all md5\n&\n max_connections: 100/' postgres?.yml \
|
||||
&& if [ "$COMPRESS" = "true" ]; then chmod u+s /usr/bin/sudo; fi \
|
||||
&& chmod +s /bin/ping \
|
||||
&& chown -R postgres:postgres $PGHOME /run /etc/haproxy
|
||||
|
||||
USER postgres
|
||||
|
||||
+36
-9
@@ -11,7 +11,11 @@ Global/Universal
|
||||
- **PATRONI\_NAME**: name of the node where the current instance of Patroni is running. Must be unique for the cluster.
|
||||
- **PATRONI\_NAMESPACE**: path within the configuration store where Patroni will keep information about the cluster. Default value: "/service"
|
||||
- **PATRONI\_SCOPE**: cluster name
|
||||
|
||||
Log
|
||||
---
|
||||
- **PATRONI\_LOG\_LEVEL**: sets the general logging level. Default value is **INFO** (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
|
||||
- **PATRONI\_LOG\_TRACEBACK\_LEVEL**: sets the level where tracebacks will be visible. Default value is **ERROR**. Set it to **DEBUG** if you want to see tracebacks only if you enable **PATRONI\_LOG\_LEVEL=DEBUG**.
|
||||
- **PATRONI\_LOG\_FORMAT**: sets the log formatting string. Default value is **%(asctime)s %(levelname)s: %(message)s** (see `the LogRecord attributes <https://docs.python.org/3.6/library/logging.html#logrecord-attributes>`_)
|
||||
- **PATRONI\_LOG\_DATEFORMAT**: sets the datetime formatting string. (see the `formatTime() documentation <https://docs.python.org/3.6/library/logging.html#logging.Formatter.formatTime>`_)
|
||||
- **PATRONI\_LOG\_MAX\_QUEUE\_SIZE**: Patroni is using two-step logging. Log records are written into the in-memory queue and there is a separate thread which pulls them from the queue and writes to stderr or file. The maximum size of the internal queue is limited by default by **1000** records, which is enough to keep logs for the past 1h20m.
|
||||
@@ -19,8 +23,6 @@ Global/Universal
|
||||
- **PATRONI\_LOG\_FILE\_NUM**: The number of application logs to retain.
|
||||
- **PATRONI\_LOG\_FILE\_SIZE**: Size of patroni.log file (in bytes) that triggers a log rolling.
|
||||
- **PATRONI\_LOG\_LOGGERS**: Redefine logging level per python module. Example ``PATRONI_LOG_LOGGERS="{patroni.postmaster: WARNING, urllib3: DEBUG}"``
|
||||
- **PATRONI\_DEBUG\_MODE**: When set to a non-empty value, makes Patroni run in the special debug mode that enables stepping with a debugger through some execution paths within Patroni `
|
||||
Example ``PATRONI_DEBUG_MODE="on"``
|
||||
|
||||
Bootstrap configuration
|
||||
-----------------------
|
||||
@@ -33,8 +35,8 @@ Example: defining ``PATRONI_admin_PASSWORD=strongpasswd`` and ``PATRONI_admin_OP
|
||||
|
||||
Consul
|
||||
------
|
||||
- **PATRONI\_CONSUL\_HOST**: the host:port for the Consul endpoint.
|
||||
- **PATRONI\_CONSUL\_URL**: url for the Consul, in format: http(s)://host:port
|
||||
- **PATRONI\_CONSUL\_HOST**: the host:port for the Consul local agent.
|
||||
- **PATRONI\_CONSUL\_URL**: url for the Consul local agent, in format: http(s)://host:port
|
||||
- **PATRONI\_CONSUL\_PORT**: (optional) Consul port
|
||||
- **PATRONI\_CONSUL\_SCHEME**: (optional) **http** or **https**, defaults to **http**
|
||||
- **PATRONI\_CONSUL\_TOKEN**: (optional) ACL token
|
||||
@@ -44,7 +46,7 @@ Consul
|
||||
- **PATRONI\_CONSUL\_KEY**: (optional) File with the client key. Can be empty if the key is part of certificate.
|
||||
- **PATRONI\_CONSUL\_DC**: (optional) Datacenter to communicate with. By default the datacenter of the host is used.
|
||||
- **PATRONI\_CONSUL\_CONSISTENCY**: (optional) Select consul consistency mode. Possible values are ``default``, ``consistent``, or ``stale`` (more details in `consul API reference <https://www.consul.io/api/features/consistency.html/>`__)
|
||||
- **PATRONI\_CONSUL\_CHECKS**: (optional) list of Consul health checks used for the session. If not specified Consul will use "serfHealth" in additional to the TTL based check created by Patroni. Additional checks, in particular the "serfHealth", may cause the leader lock to expire faster than in `ttl` seconds when the leader instance becomes unavailable.
|
||||
- **PATRONI\_CONSUL\_CHECKS**: (optional) list of Consul health checks used for the session. By default an empty list is used.
|
||||
- **PATRONI\_CONSUL\_REGISTER\_SERVICE**: (optional) whether or not to register a service with the name defined by the scope parameter and the tag master, replica or standby-leader depending on the node's role. Defaults to **false**
|
||||
- **PATRONI\_CONSUL\_SERVICE\_CHECK\_INTERVAL**: (optional) how often to perform health check against registered url
|
||||
|
||||
@@ -64,6 +66,10 @@ Etcd
|
||||
- **PATRONI\_ETCD\_CERT**: File with the client certificate.
|
||||
- **PATRONI\_ETCD\_KEY**: File with the client key. Can be empty if the key is part of certificate.
|
||||
|
||||
ZooKeeper
|
||||
---------
|
||||
- **PATRONI\_ZOOKEEPER\_HOSTS**: comma separated list of ZooKeeper cluster members: "'host1:port1','host2:port2','etc...'". It is important to quote every single entity!
|
||||
|
||||
Exhibitor
|
||||
---------
|
||||
- **PATRONI\_EXHIBITOR\_HOSTS**: initial list of Exhibitor (ZooKeeper) nodes in format: 'host1,host2,etc...'. This list updates automatically whenever the Exhibitor (ZooKeeper) cluster topology changes.
|
||||
@@ -79,7 +85,7 @@ Kubernetes
|
||||
- **PATRONI\_KUBERNETES\_ROLE\_LABEL**: (optional) name of the label containing Postgres role (`master` or `replica`). Patroni will set this label on the pod it is running in. Default value is `role`.
|
||||
- **PATRONI\_KUBERNETES\_USE\_ENDPOINTS**: (optional) if set to true, Patroni will use Endpoints instead of ConfigMaps to run leader elections and keep cluster state.
|
||||
- **PATRONI\_KUBERNETES\_POD\_IP**: (optional) IP address of the pod Patroni is running in. This value is required when `PATRONI_KUBERNETES_USE_ENDPOINTS` is enabled and is used to populate the leader endpoint subsets when the pod's PostgreSQL is promoted.
|
||||
- **PATRONI\_KUBERNETES\_PORTS**: (optional) if the Service object has the name for the port, the same name must appear in the Endpoint object, otherwise service won't work. For example, if your service is defined as ``{Kind: Service, spec: {ports: [{name: postgresql, port: 5432, targetPort: 5432}]}}``, then you have to set ``PATRONI_KUBERNETES_PORTS='{[{"name": "postgresql", "port": 5432}]}'`` and Patroni will use it for updating subsets of the leader Endpoint. This parameter is used only if `PATRONI_KUBERNETES_USE_ENDPOINTS` is set.
|
||||
- **PATRONI\_KUBERNETES\_PORTS**: (optional) if the Service object has the name for the port, the same name must appear in the Endpoint object, otherwise service won't work. For example, if your service is defined as ``{Kind: Service, spec: {ports: [{name: postgresql, port: 5432, targetPort: 5432}]}}``, then you have to set ``PATRONI_KUBERNETES_PORTS='[{"name": "postgresql", "port": 5432}]'`` and Patroni will use it for updating subsets of the leader Endpoint. This parameter is used only if `PATRONI_KUBERNETES_USE_ENDPOINTS` is set.
|
||||
|
||||
PostgreSQL
|
||||
----------
|
||||
@@ -91,10 +97,25 @@ PostgreSQL
|
||||
- **PATRONI\_POSTGRESQL\_PGPASS**: path to the `.pgpass <https://www.postgresql.org/docs/current/static/libpq-pgpass.html>`__ password file. Patroni creates this file before executing pg\_basebackup and under some other circumstances. The location must be writable by Patroni.
|
||||
- **PATRONI\_REPLICATION\_USERNAME**: replication username; the user will be created during initialization. Replicas will use this user to access master via streaming replication
|
||||
- **PATRONI\_REPLICATION\_PASSWORD**: replication password; the user will be created during initialization.
|
||||
- **PATRONI\_REPLICATION\_SSLMODE**: (optional) maps to the `sslmode <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLMODE>`__ connection parameter, which allows a client to specify the type of TLS negotiation mode with the server. For more information on how each mode works, please visit the `PostgreSQL documentation <https://www.postgresql.org/docs/current/libpq-ssl.html#LIBPQ-SSL-SSLMODE-STATEMENTS>`__. The default mode is ``prefer``.
|
||||
- **PATRONI\_REPLICATION\_SSLKEY**: (optional) maps to the `sslkey <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLKEY>`__ connection parameter, which specifies the location of the secret key used with the client's certificate.
|
||||
- **PATRONI\_REPLICATION\_SSLCERT**: (optional) maps to the `sslcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCERT>`__ connection parameter, which specifies the location of the client certificate.
|
||||
- **PATRONI\_REPLICATION\_SSLROOTCERT**: (optional) maps to the `sslrootcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLROOTCERT>`__ connection parameter, which specifies the location of a file containing one ore more certificate authorities (CA) certificates that the client will use to verify a server's certificate.
|
||||
- **PATRONI\_REPLICATION\_SSLCRL**: (optional) maps to the `sslcrl <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCRL>`__ connection parameter, which specifies the location of a file containing a certificate revocation list. A client will reject connecting to any server that has a certificate present in this list.
|
||||
- **PATRONI\_SUPERUSER\_USERNAME**: name for the superuser, set during initialization (initdb) and later used by Patroni to connect to the postgres. Also this user is used by pg_rewind.
|
||||
- **PATRONI\_SUPERUSER\_PASSWORD**: password for the superuser, set during initialization (initdb).
|
||||
- **PATRONI\_SUPERUSER\_SSLMODE**: (optional) maps to the `sslmode <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLMODE>`__ connection parameter, which allows a client to specify the type of TLS negotiation mode with the server. For more information on how each mode works, please visit the `PostgreSQL documentation <https://www.postgresql.org/docs/current/libpq-ssl.html#LIBPQ-SSL-SSLMODE-STATEMENTS>`__. The default mode is ``prefer``.
|
||||
- **PATRONI\_SUPERUSER\_SSLKEY**: (optional) maps to the `sslkey <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLKEY>`__ connection parameter, which specifies the location of the secret key used with the client's certificate.
|
||||
- **PATRONI\_SUPERUSER\_SSLCERT**: (optional) maps to the `sslcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCERT>`__ connection parameter, which specifies the location of the client certificate.
|
||||
- **PATRONI\_SUPERUSER\_SSLROOTCERT**: (optional) maps to the `sslrootcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLROOTCERT>`__ connection parameter, which specifies the location of a file containing one ore more certificate authorities (CA) certificates that the client will use to verify a server's certificate.
|
||||
- **PATRONI\_SUPERUSER\_SSLCRL**: (optional) maps to the `sslcrl <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCRL>`__ connection parameter, which specifies the location of a file containing a certificate revocation list. A client will reject connecting to any server that has a certificate present in this list.
|
||||
- **PATRONI\_REWIND\_USERNAME**: name for the user for ``pg_rewind``; the user will be created during initialization of postgres 11+ and all necessary `permissions <https://www.postgresql.org/docs/11/app-pgrewind.html#id-1.9.5.8.8>`__ will be granted.
|
||||
- **PATRONI\_REWIND\_PASSWORD**: password for the user for ``pg_rewind``; the user will be created during initialization.
|
||||
- **PATRONI\_REWIND\_SSLMODE**: (optional) maps to the `sslmode <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLMODE>`__ connection parameter, which allows a client to specify the type of TLS negotiation mode with the server. For more information on how each mode works, please visit the `PostgreSQL documentation <https://www.postgresql.org/docs/current/libpq-ssl.html#LIBPQ-SSL-SSLMODE-STATEMENTS>`__. The default mode is ``prefer``.
|
||||
- **PATRONI\_REWIND\_SSLKEY**: (optional) maps to the `sslkey <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLKEY>`__ connection parameter, which specifies the location of the secret key used with the client's certificate.
|
||||
- **PATRONI\_REWIND\_SSLCERT**: (optional) maps to the `sslcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCERT>`__ connection parameter, which specifies the location of the client certificate.
|
||||
- **PATRONI\_REWIND\_SSLROOTCERT**: (optional) maps to the `sslrootcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLROOTCERT>`__ connection parameter, which specifies the location of a file containing one ore more certificate authorities (CA) certificates that the client will use to verify a server's certificate.
|
||||
- **PATRONI\_REWIND\_SSLCRL**: (optional) maps to the `sslcrl <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCRL>`__ connection parameter, which specifies the location of a file containing a certificate revocation list. A client will reject connecting to any server that has a certificate present in this list.
|
||||
|
||||
REST API
|
||||
--------
|
||||
@@ -104,7 +125,13 @@ REST API
|
||||
- **PATRONI\_RESTAPI\_PASSWORD**: Basic-auth password to protect unsafe REST API endpoints.
|
||||
- **PATRONI\_RESTAPI\_CERTFILE**: Specifies the file with the certificate in the PEM format. If the certfile is not specified or is left empty, the API server will work without SSL.
|
||||
- **PATRONI\_RESTAPI\_KEYFILE**: Specifies the file with the secret key in the PEM format.
|
||||
- **PATRONI\_RESTAPI\_CAFILE**: Specifies the file with the CA_BUNDLE with certificates of trusted CAs to use while verifying client certs.
|
||||
- **PATRONI\_RESTAPI\_VERIFY\_CLIENT**: ``none``, ``optional`` or ``required``. When ``none`` REST API will not check client certificates. When ``required`` client certificates are required for all REST API calls. When ``optional`` client certificates are required for all unsafe REST API endpoints. If ``verify_client`` is set to ``optional`` or ``required`` basic-auth is not checked.
|
||||
|
||||
ZooKeeper
|
||||
---------
|
||||
- **PATRONI\_ZOOKEEPER\_HOSTS**: comma separated list of ZooKeeper cluster members: "'host1:port1','host2:port2','etc...'". It is important to quote every single entity!
|
||||
CTL
|
||||
---
|
||||
- **PATRONICTL\_CONFIG\_FILE**: location of the configuration file.
|
||||
- **PATRONI\_CTL\_INSECURE**: Allow connections to REST API without verifying SSL certs.
|
||||
- **PATRONI\_CTL\_CACERT**: Specifies the file with the CA_BUNDLE file or directory with certificates of trusted CAs to use while verifying REST API SSL certs. If not provided patronictl will use the value provided for REST API "cafile" parameter.
|
||||
- **PATRONI\_CTL\_CERTFILE**: Specifies the file with the client certificate in the PEM format. If not provided patronictl will use the value provided for REST API "certfile" parameter.
|
||||
- **PATRONI\_CTL\_KEYFILE**: Specifies the file with the client secret key in the PEM format. If not provided patronictl will use the value provided for REST API "keyfile" parameter.
|
||||
|
||||
+1
-1
@@ -107,7 +107,7 @@ obtain those files from the git repository and replace `./patroni.py` below with
|
||||
To get started, do the following from different terminals:
|
||||
::
|
||||
|
||||
> etcd --data-dir=data/etcd
|
||||
> etcd --data-dir=data/etcd --enable-v2=true
|
||||
> ./patroni.py postgres0.yml
|
||||
> ./patroni.py postgres1.yml
|
||||
|
||||
|
||||
+84
-47
@@ -4,6 +4,41 @@
|
||||
YAML Configuration Settings
|
||||
===========================
|
||||
|
||||
.. _dynamic_configuration_settings:
|
||||
|
||||
Dynamic configuration settings
|
||||
------------------------------
|
||||
|
||||
Dynamic configuration is stored in the DCS (Distributed Configuration Store) and applied on all cluster nodes. Some parameters, like **loop_wait**, **ttl**, **postgresql.parameters.max_connections**, **postgresql.parameters.max_worker_processes** and so on could be set only in the dynamic configuration. Some other parameters like **postgresql.listen**, **postgresql.data_dir** could be set only locally, i.e. in the Patroni config file or via :ref:`configuration <environment>` variable. In most cases the local configuration will override the dynamic configuration. In order to change the dynamic configuration you can use either ``patronictl edit-config`` tool or Patroni :ref:`REST API <rest_api>`.
|
||||
|
||||
- **loop\_wait**: the number of seconds the loop will sleep. Default value: 10
|
||||
- **ttl**: the TTL to acquire the leader lock (in seconds). Think of it as the length of time before initiation of the automatic failover process. Default value: 30
|
||||
- **retry\_timeout**: timeout for DCS and PostgreSQL operation retries (in seconds). DCS or network issues shorter than this will not cause Patroni to demote the leader. Default value: 10
|
||||
- **maximum\_lag\_on\_failover**: the maximum bytes a follower may lag to be able to participate in leader election.
|
||||
- **max\_timelines\_history**: maximum number of timeline history items kept in DCS. Default value: 0. When set to 0, it keeps the full history in DCS.
|
||||
- **master\_start\_timeout**: the amount of time a master is allowed to recover from failures before failover is triggered (in seconds). Default is 300 seconds. When set to 0 failover is done immediately after a crash is detected if possible. When using asynchronous replication a failover can cause lost transactions. Worst case failover time for master failure is: loop\_wait + master\_start\_timeout + loop\_wait, unless master\_start\_timeout is zero, in which case it's just loop\_wait. Set the value according to your durability/availability tradeoff.
|
||||
- **master\_stop\_timeout**: The number of seconds Patroni is allowed to wait when stopping Postgres and effective only when synchronous_mode is enabled. When set to > 0 and the synchronous_mode is enabled, Patroni sends SIGKILL to the postmaster if the stop operation is running for more than the value set by master_stop_timeout. Set the value according to your durability/availability tradeoff. If the parameter is not set or set <= 0, master_stop_timeout does not apply.
|
||||
- **synchronous\_mode**: turns on synchronous replication mode. In this mode a replica will be chosen as synchronous and only the latest leader and synchronous replica are able to participate in leader election. Synchronous mode makes sure that successfully committed transactions will not be lost at failover, at the cost of losing availability for writes when Patroni cannot ensure transaction durability. See :ref:`replication modes documentation <replication_modes>` for details.
|
||||
- **synchronous\_mode\_strict**: prevents disabling synchronous replication if no synchronous replicas are available, blocking all client writes to the master. See :ref:`replication modes documentation <replication_modes>` for details.
|
||||
- **postgresql**:
|
||||
- **use\_pg\_rewind**: whether or not to use pg_rewind. Defaults to `false`.
|
||||
- **use\_slots**: whether or not to use replication slots. Defaults to `true` on PostgreSQL 9.4+.
|
||||
- **recovery\_conf**: additional configuration settings written to recovery.conf when configuring follower. There is no recovery.conf anymore in PostgreSQL 12, but you may continue using this section, because Patroni handles it transparently.
|
||||
- **parameters**: list of configuration settings for Postgres.
|
||||
- **standby\_cluster**: if this section is defined, we want to bootstrap a standby cluster.
|
||||
- **host**: an address of remote master
|
||||
- **port**: a port of remote master
|
||||
- **primary\_slot\_name**: which slot on the remote master to use for replication. This parameter is optional, the default value is derived from the instance name (see function `slot_name_from_member_name`).
|
||||
- **create\_replica\_methods**: an ordered list of methods that can be used to bootstrap standby leader from the remote master, can be different from the list defined in :ref:`postgresql_settings`
|
||||
- **restore\_command**: command to restore WAL records from the remote master to standby leader, can be different from the list defined in :ref:`postgresql_settings`
|
||||
- **archive\_cleanup\_command**: cleanup command for standby leader
|
||||
- **recovery\_min\_apply\_delay**: how long to wait before actually apply WAL records on a standby leader
|
||||
- **slots**: define permanent replication slots. These slots will be preserved during switchover/failover. Patroni will try to create slots before opening connections to the cluster.
|
||||
- **my_slot_name**: the name of replication slot. If the permanent slot name matches with the name of the current primary it will not be created. Everything else is the responsibility of the operator to make sure that there are no clashes in names between replication slots automatically created by Patroni for members and permanent replication slots.
|
||||
- **type**: slot type. Could be ``physical`` or ``logical``. If the slot is logical, you have to additionally define ``database`` and ``plugin``.
|
||||
- **database**: the database name where logical slots should be created.
|
||||
- **plugin**: the plugin name for the logical slot.
|
||||
|
||||
Global/Universal
|
||||
----------------
|
||||
- **name**: the name of the host. Must be unique for the cluster.
|
||||
@@ -13,6 +48,7 @@ Global/Universal
|
||||
Log
|
||||
---
|
||||
- **level**: sets the general logging level. Default value is **INFO** (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
|
||||
- **traceback\_level**: sets the level where tracebacks will be visible. Default value is **ERROR**. Set it to **DEBUG** if you want to see tracebacks only if you enable **log.level=DEBUG**.
|
||||
- **format**: sets the log formatting string. Default value is **%(asctime)s %(levelname)s: %(message)s** (see `the LogRecord attributes <https://docs.python.org/3.6/library/logging.html#logrecord-attributes>`_)
|
||||
- **dateformat**: sets the datetime formatting string. (see the `formatTime() documentation <https://docs.python.org/3.6/library/logging.html#logging.Formatter.formatTime>`_)
|
||||
- **max\_queue\_size**: Patroni is using two-step logging. Log records are written into the in-memory queue and there is a separate thread which pulls them from the queue and writes to stderr or file. The maximum size of the internal queue is limited by default by **1000** records, which is enough to keep logs for the past 1h20m.
|
||||
@@ -27,32 +63,7 @@ Log
|
||||
|
||||
Bootstrap configuration
|
||||
-----------------------
|
||||
- **dcs**: This section will be written into `/<namespace>/<scope>/config` of a given configuration store after initializing of new cluster. This is the global configuration for the cluster. If you want to change some parameters for all cluster nodes - just do it in DCS (or via Patroni API) and all nodes will apply this configuration.
|
||||
- **loop\_wait**: the number of seconds the loop will sleep. Default value: 10
|
||||
- **ttl**: the TTL to acquire the leader lock. Think of it as the length of time before initiation of the automatic failover process. Default value: 30
|
||||
- **retry\_timeout**: timeout for DCS and PostgreSQL operation retries. DCS or network issues shorter than this will not cause Patroni to demote the leader. Default value: 10
|
||||
- **maximum\_lag\_on\_failover**: the maximum bytes a follower may lag to be able to participate in leader election.
|
||||
- **master\_start\_timeout**: the amount of time a master is allowed to recover from failures before failover is triggered. Default is 300 seconds. When set to 0 failover is done immediately after a crash is detected if possible. When using asynchronous replication a failover can cause lost transactions. Best worst case failover time for master failure is: loop\_wait + master\_start\_timeout + loop\_wait, unless master\_start\_timeout is zero, in which case it's just loop\_wait. Set the value according to your durability/availability tradeoff.
|
||||
- **synchronous\_mode**: turns on synchronous replication mode. In this mode a replica will be chosen as synchronous and only the latest leader and synchronous replica are able to participate in leader election. Synchronous mode makes sure that successfully committed transactions will not be lost at failover, at the cost of losing availability for writes when Patroni cannot ensure transaction durability. See :ref:`replication modes documentation <replication_modes>` for details.
|
||||
- **synchronous\_mode\_strict**: prevents disabling synchronous replication if no synchronous replicas are available, blocking all client writes to the master. See :ref:`replication modes documentation <replication_modes>` for details.
|
||||
- **postgresql**:
|
||||
- **use\_pg\_rewind**: whether or not to use pg_rewind
|
||||
- **use\_slots**: whether or not to use replication_slots. Must be False for PostgreSQL 9.3. You should comment out max_replication_slots before it becomes ineligible for leader status.
|
||||
- **recovery\_conf**: additional configuration settings written to recovery.conf when configuring follower.
|
||||
- **parameters**: list of configuration settings for Postgres. Many of these are required for replication to work.
|
||||
- **standby\_cluster**: if this section is defined, we want to bootstrap a standby cluster.
|
||||
- **host**: an address of remote master
|
||||
- **port**: a port of remote master
|
||||
- **primary\_slot\_name**: which slot on the remote master to use for replication. This parameter is optional, the default value is derived from the instance name (see function `slot_name_from_member_name`).
|
||||
- **create\_replica\_methods**: an ordered list of methods that can be used to bootstrap standby leader from the remote master, can be different from the list defined in :ref:`postgresql_settings`
|
||||
- **restore\_command**: command to restore WAL records from the remote master to standby leader, can be different from the list defined in :ref:`postgresql_settings`
|
||||
- **archive\_cleanup\_command**: cleanup command for standby leader
|
||||
- **recovery\_min\_apply\_delay**: how long to wait before actually apply WAL records on a standby leader
|
||||
- **slots**: define permanent replication slots. These slots will be preserved during switchover/failover. Patroni will try to create slots before opening connections to the cluster.
|
||||
- **my_slot_name**: the name of replication slot. It is the responsibility of the operator to make sure that there are no clashes in names between replication slots automatically created by Patroni for members and permanent replication slots.
|
||||
- **type**: slot type. Could be ``physical`` or ``logical``. If the slot is logical, you have to additionally define ``database`` and ``plugin``.
|
||||
- **database**: the database name where logical slots should be created.
|
||||
- **plugin**: the plugin name for the logical slot.
|
||||
- **dcs**: This section will be written into `/<namespace>/<scope>/config` of the given configuration store after initializing of new cluster. The global dynamic configuration for the cluster. Under the ``bootstrap.dcs`` you can put any of the parameters described in the :ref:`Dynamic Configuration settings <dynamic_configuration_settings>` and after Patroni initialized (bootstrapped) the new cluster, it will write this section into `/<namespace>/<scope>/config` of the configuration store. All later changes of ``bootstrap.dcs`` will not take any effect! If you want to change them please use either ``patronictl edit-config`` or Patroni :ref:`REST API <rest_api>`.
|
||||
- **method**: custom script to use for bootstrapping this cluster.
|
||||
See :ref:`custom bootstrap methods documentation <custom_bootstrap>` for details.
|
||||
When ``initdb`` is specified revert to the default ``initdb`` command. ``initdb`` is also triggered when no ``method``
|
||||
@@ -78,20 +89,20 @@ Consul
|
||||
------
|
||||
Most of the parameters are optional, but you have to specify one of the **host** or **url**
|
||||
|
||||
- **host**: the host:port for the Consul endpoint, in format: http(s)://host:port
|
||||
- **url**: url for the Consul endpoint
|
||||
- **port**: (optional) Consul port
|
||||
- **scheme**: (optional) **http** or **https**, defaults to **http**
|
||||
- **token**: (optional) ACL token
|
||||
- **verify**: (optional) whether to verify the SSL certificate for HTTPS requests
|
||||
- **host**: the host:port for the Consul local agent.
|
||||
- **url**: url for the Consul local agent, in format: http(s)://host:port.
|
||||
- **port**: (optional) Consul port.
|
||||
- **scheme**: (optional) **http** or **https**, defaults to **http**.
|
||||
- **token**: (optional) ACL token.
|
||||
- **verify**: (optional) whether to verify the SSL certificate for HTTPS requests.
|
||||
- **cacert**: (optional) The ca certificate. If present it will enable validation.
|
||||
- **cert**: (optional) file with the client certificate
|
||||
- **cert**: (optional) file with the client certificate.
|
||||
- **key**: (optional) file with the client key. Can be empty if the key is part of **cert**.
|
||||
- **dc**: (optional) Datacenter to communicate with. By default the datacenter of the host is used.
|
||||
- **consistency**: (optional) Select consul consistency mode. Possible values are ``default``, ``consistent``, or ``stale`` (more details in `consul API reference <https://www.consul.io/api/features/consistency.html/>`__)
|
||||
- **checks**: (optional) list of Consul health checks used for the session. If not specified Consul will use "serfHealth" in additional to the TTL based check created by Patroni. Additional checks, in particular the "serfHealth", may cause the leader lock to expire faster than in `ttl` seconds when the leader instance becomes unavailable
|
||||
- **register\_service**: (optional) whether or not to register a service with the name defined by the scope parameter and the tag master, replica or standby-leader depending on the node's role. Defaults to **false**
|
||||
- **service\_check\_interval**: (optional) how often to perform health check against registered url
|
||||
- **checks**: (optional) list of Consul health checks used for the session. By default an empty list is used.
|
||||
- **register\_service**: (optional) whether or not to register a service with the name defined by the scope parameter and the tag master, replica or standby-leader depending on the node's role. Defaults to **false**.
|
||||
- **service\_check\_interval**: (optional) how often to perform health check against registered url.
|
||||
|
||||
Etcd
|
||||
----
|
||||
@@ -100,8 +111,8 @@ Most of the parameters are optional, but you have to specify one of the **host**
|
||||
- **host**: the host:port for the etcd endpoint.
|
||||
- **hosts**: list of etcd endpoint in format host1:port1,host2:port2,etc... Could be a comma separated string or an actual yaml list.
|
||||
- **use\_proxies**: If this parameter is set to true, Patroni will consider **hosts** as a list of proxies and will not perform a topology discovery of etcd cluster.
|
||||
- **url**: url for the etcd
|
||||
- **proxy**: proxy url for the etcd. If you are connecting to the etcd using proxy, use this parameter instead of **url**
|
||||
- **url**: url for the etcd.
|
||||
- **proxy**: proxy url for the etcd. If you are connecting to the etcd using proxy, use this parameter instead of **url**.
|
||||
- **srv**: Domain to search the SRV record(s) for cluster autodiscovery.
|
||||
- **protocol**: (optional) http or https, if not specified http is used. If the **url** or **proxy** is specified - will take protocol from them.
|
||||
- **username**: (optional) username for etcd authentication.
|
||||
@@ -110,10 +121,14 @@ Most of the parameters are optional, but you have to specify one of the **host**
|
||||
- **cert**: (optional) file with the client certificate.
|
||||
- **key**: (optional) file with the client key. Can be empty if the key is part of **cert**.
|
||||
|
||||
ZooKeeper
|
||||
----------
|
||||
- **hosts**: list of ZooKeeper cluster members in format: ['host1:port1', 'host2:port2', 'etc...'].
|
||||
|
||||
Exhibitor
|
||||
---------
|
||||
- **hosts**: initial list of Exhibitor (ZooKeeper) nodes in format: 'host1,host2,etc...'. This list updates automatically whenever the Exhibitor (ZooKeeper) cluster topology changes.
|
||||
- **poll\_interval**: how often the list of ZooKeeper and Exhibitor nodes should be updated from Exhibitor
|
||||
- **poll\_interval**: how often the list of ZooKeeper and Exhibitor nodes should be updated from Exhibitor.
|
||||
- **port**: Exhibitor port.
|
||||
|
||||
.. _kubernetes_settings:
|
||||
@@ -126,7 +141,7 @@ Kubernetes
|
||||
- **role\_label**: (optional) name of the label containing role (master or replica). Patroni will set this label on the pod it runs in. Default value is ``role``.
|
||||
- **use\_endpoints**: (optional) if set to true, Patroni will use Endpoints instead of ConfigMaps to run leader elections and keep cluster state.
|
||||
- **pod\_ip**: (optional) IP address of the pod Patroni is running in. This value is required when `use_endpoints` is enabled and is used to populate the leader endpoint subsets when the pod's PostgreSQL is promoted.
|
||||
- **ports**: (optional) if the Service object has the name for the port, the same name must appear in the Endpoint object, otherwise service won't work. For example, if your service is defined as ``{Kind: Service, spec: {ports: [{name: postgresql, port: 5432, targetPort: 5432}]}}``, then you have to set ``kubernetes.ports: {[{"name": "postgresql", "port": 5432}]}`` and Patroni will use it for updating subsets of the leader Endpoint. This parameter is used only if `kubernetes.use_endpoints` is set.
|
||||
- **ports**: (optional) if the Service object has the name for the port, the same name must appear in the Endpoint object, otherwise service won't work. For example, if your service is defined as ``{Kind: Service, spec: {ports: [{name: postgresql, port: 5432, targetPort: 5432}]}}``, then you have to set ``kubernetes.ports: [{"name": "postgresql", "port": 5432}]`` and Patroni will use it for updating subsets of the leader Endpoint. This parameter is used only if `kubernetes.use_endpoints` is set.
|
||||
|
||||
.. _postgresql_settings:
|
||||
|
||||
@@ -136,12 +151,27 @@ PostgreSQL
|
||||
- **superuser**:
|
||||
- **username**: name for the superuser, set during initialization (initdb) and later used by Patroni to connect to the postgres.
|
||||
- **password**: password for the superuser, set during initialization (initdb).
|
||||
- **sslmode**: (optional) maps to the `sslmode <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLMODE>`__ connection parameter, which allows a client to specify the type of TLS negotiation mode with the server. For more information on how each mode works, please visit the `PostgreSQL documentation <https://www.postgresql.org/docs/current/libpq-ssl.html#LIBPQ-SSL-SSLMODE-STATEMENTS>`__. The default mode is ``prefer``.
|
||||
- **sslkey**: (optional) maps to the `sslkey <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLKEY>`__ connection parameter, which specifies the location of the secret key used with the client's certificate.
|
||||
- **sslcert**: (optional) maps to the `sslcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCERT>`__ connection parameter, which specifies the location of the client certificate.
|
||||
- **sslrootcert**: (optional) maps to the `sslrootcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLROOTCERT>`__ connection parameter, which specifies the location of a file containing one ore more certificate authorities (CA) certificates that the client will use to verify a server's certificate.
|
||||
- **sslcrl**: (optional) maps to the `sslcrl <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCRL>`__ connection parameter, which specifies the location of a file containing a certificate revocation list. A client will reject connecting to any server that has a certificate present in this list.
|
||||
- **replication**:
|
||||
- **username**: replication username; the user will be created during initialization. Replicas will use this user to access master via streaming replication
|
||||
- **password**: replication password; the user will be created during initialization.
|
||||
- **sslmode**: (optional) maps to the `sslmode <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLMODE>`__ connection parameter, which allows a client to specify the type of TLS negotiation mode with the server. For more information on how each mode works, please visit the `PostgreSQL documentation <https://www.postgresql.org/docs/current/libpq-ssl.html#LIBPQ-SSL-SSLMODE-STATEMENTS>`__. The default mode is ``prefer``.
|
||||
- **sslkey**: (optional) maps to the `sslkey <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLKEY>`__ connection parameter, which specifies the location of the secret key used with the client's certificate.
|
||||
- **sslcert**: (optional) maps to the `sslcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCERT>`__ connection parameter, which specifies the location of the client certificate.
|
||||
- **sslrootcert**: (optional) maps to the `sslrootcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLROOTCERT>`__ connection parameter, which specifies the location of a file containing one ore more certificate authorities (CA) certificates that the client will use to verify a server's certificate.
|
||||
- **sslcrl**: (optional) maps to the `sslcrl <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCRL>`__ connection parameter, which specifies the location of a file containing a certificate revocation list. A client will reject connecting to any server that has a certificate present in this list.
|
||||
- **rewind**:
|
||||
- **username**: name for the user for ``pg_rewind``; the user will be created during initialization of postgres 11+ and all necessary `permissions <https://www.postgresql.org/docs/11/app-pgrewind.html#id-1.9.5.8.8>`__ will be granted.
|
||||
- **password**: password for the user for ``pg_rewind``; the user will be created during initialization.
|
||||
- **sslmode**: (optional) maps to the `sslmode <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLMODE>`__ connection parameter, which allows a client to specify the type of TLS negotiation mode with the server. For more information on how each mode works, please visit the `PostgreSQL documentation <https://www.postgresql.org/docs/current/libpq-ssl.html#LIBPQ-SSL-SSLMODE-STATEMENTS>`__. The default mode is ``prefer``.
|
||||
- **sslkey**: (optional) maps to the `sslkey <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLKEY>`__ connection parameter, which specifies the location of the secret key used with the client's certificate.
|
||||
- **sslcert**: (optional) maps to the `sslcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCERT>`__ connection parameter, which specifies the location of the client certificate.
|
||||
- **sslrootcert**: (optional) maps to the `sslrootcert <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLROOTCERT>`__ connection parameter, which specifies the location of a file containing one ore more certificate authorities (CA) certificates that the client will use to verify a server's certificate.
|
||||
- **sslcrl**: (optional) maps to the `sslcrl <https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-CONNECT-SSLCRL>`__ connection parameter, which specifies the location of a file containing a certificate revocation list. A client will reject connecting to any server that has a certificate present in this list.
|
||||
- **callbacks**: callback scripts to run on certain actions. Patroni will pass the action, role and cluster name. (See scripts/aws.py as an example of how to write them.)
|
||||
- **on\_reload**: run this script when configuration reload is triggered.
|
||||
- **on\_restart**: run this script when the postgres restarts (without changing role).
|
||||
@@ -175,7 +205,7 @@ PostgreSQL
|
||||
|
||||
REST API
|
||||
--------
|
||||
- **connect\_address**: IP address (or hostname) and port, to access the Patroni's REST API. All the members of the cluster must be able to connect to this address, so unless the Patroni setup is intended for a demo inside the localhost, this address must be a non "localhost" or loopback addres (ie: "localhost" or "127.0.0.1"). It can serve as a endpoint for HTTP health checks (read below about the "listen" REST API parameter), and also for user queries (either directly or via the REST API), as well as for the health checks done by the cluster members during leader elections (for example, to determine whether the master is still running, or if there is a node which has a WAL position that is ahead of the one doing the query; etc.) The connect_address is put in the member key in DCS, making it possible to translate the member name into the address to connect to its REST API.
|
||||
- **connect\_address**: IP address (or hostname) and port, to access the Patroni's :ref:`REST API <rest_api>`. All the members of the cluster must be able to connect to this address, so unless the Patroni setup is intended for a demo inside the localhost, this address must be a non "localhost" or loopback address (ie: "localhost" or "127.0.0.1"). It can serve as an endpoint for HTTP health checks (read below about the "listen" REST API parameter), and also for user queries (either directly or via the REST API), as well as for the health checks done by the cluster members during leader elections (for example, to determine whether the master is still running, or if there is a node which has a WAL position that is ahead of the one doing the query; etc.) The connect_address is put in the member key in DCS, making it possible to translate the member name into the address to connect to its REST API.
|
||||
|
||||
- **listen**: IP address (or hostname) and port that Patroni will listen to for the REST API - to provide also the same health checks and cluster messaging between the participating nodes, as described above. to provide health-check information for HAProxy (or any other load balancer capable of doing a HTTP "OPTION" or "GET" checks).
|
||||
|
||||
@@ -186,6 +216,8 @@ REST API
|
||||
|
||||
- **certfile**: Specifies the file with the certificate in the PEM format. If the certfile is not specified or is left empty, the API server will work without SSL.
|
||||
- **keyfile**: Specifies the file with the secret key in the PEM format.
|
||||
- **cafile**: Specifies the file with the CA_BUNDLE with certificates of trusted CAs to use while verifying client certs.
|
||||
- **verify\_client**: ``none``, ``optional`` or ``required``. When ``none`` REST API will not check client certificates. When ``required`` client certificates are required for all REST API calls. When ``optional`` client certificates are required for all unsafe REST API endpoints. If ``verify_client`` is set to ``optional`` or ``required`` basic-auth is not checked.
|
||||
|
||||
.. _patronictl_settings:
|
||||
|
||||
@@ -193,15 +225,20 @@ CTL
|
||||
---
|
||||
- **Optional**:
|
||||
- **insecure**: Allow connections to REST API without verifying SSL certs.
|
||||
- **cacert**: Specifices the file with the CA_BUNDLE file or directory with certificates of trusted CAs to use while verifying REST API SSL certs.
|
||||
- **certfile**: Specifies the file with the certificate in the PEM format to use while verifying REST API SSL certs. If not provided patronictl will use the value provided for REST API "certfile" parameter.
|
||||
|
||||
ZooKeeper
|
||||
----------
|
||||
- **hosts**: list of ZooKeeper cluster members in format: ['host1:port1', 'host2:port2', 'etc...'].
|
||||
- **cacert**: Specifies the file with the CA_BUNDLE file or directory with certificates of trusted CAs to use while verifying REST API SSL certs. If not provided patronictl will use the value provided for REST API "cafile" parameter.
|
||||
- **certfile**: Specifies the file with the client certificate in the PEM format. If not provided patronictl will use the value provided for REST API "certfile" parameter.
|
||||
- **keyfile**: Specifies the file with the client secret key in the PEM format. If not provided patronictl will use the value provided for REST API "keyfile" parameter.
|
||||
|
||||
Watchdog
|
||||
--------
|
||||
- **mode**: ``off``, ``automatic`` or ``required``. When ``off`` watchdog is disabled. When ``automatic`` watchdog will be used if available, but ignored if it is not. When ``required`` the node will not become a leader unless watchdog can be successfully enabled.
|
||||
- **device**: Path to watchdog device. Defaults to ``/dev/watchdog``.
|
||||
- **safety_margin**: Number of seconds of safety margin between watchdog triggering and leader key expiration.
|
||||
|
||||
Tags
|
||||
----
|
||||
- **nofailover**: ``true`` or ``false``, controls whether this node is allowed to participate in the leader race and become a leader. Defaults to ``false``
|
||||
- **clonefrom**: ``true`` or ``false``. If set to ``true`` other nodes might prefer to use this node for bootstrap (take ``pg_basebackup`` from). If there are several nodes with ``clonefrom`` tag set to ``true`` the node to bootstrap from will be chosen randomly. The default value is ``false``.
|
||||
- **noloadbalance**: ``true`` or ``false``. If set to ``true`` the node will return HTTP Status Code 503 for the ``GET /replica`` REST API health-check and therefore will be excluded from the load-balancing. Defaults to ``false``.
|
||||
- **replicatefrom**: The IP address/hostname of another replica. Used to support cascading replication.
|
||||
- **nosync**: ``true`` or ``false``. If set to ``true`` the node will never be selected as a synchronous replica.
|
||||
|
||||
@@ -14,7 +14,7 @@ Patroni configuration is stored in the DCS (Distributed Configuration Store). Th
|
||||
|
||||
- Local :ref:`configuration <settings>` (patroni.yml).
|
||||
These options are defined in the configuration file and take precedence over dynamic configuration.
|
||||
patroni.yml could be changed and reload in runtime (without restart of Patroni) by sending SIGHUP to the Patroni process or by performing ``POST /reload`` REST-API request.
|
||||
patroni.yml could be changed and reloaded in runtime (without restart of Patroni) by sending SIGHUP to the Patroni process, performing ``POST /reload`` REST-API request or executing ``patronictl reload``.
|
||||
|
||||
- Environment :ref:`configuration <environment>`.
|
||||
It is possible to set/override some of the "Local" configuration parameters with environment variables.
|
||||
@@ -76,6 +76,7 @@ Also, the following Patroni configuration options can be changed only dynamicall
|
||||
- loop_wait: 10
|
||||
- retry_timeouts: 10
|
||||
- maximum_lag_on_failover: 1048576
|
||||
- max_timelines_history: 0
|
||||
- check_timeline: false
|
||||
- postgresql.use_slots: true
|
||||
|
||||
@@ -83,147 +84,3 @@ Upon changing these options, Patroni will read the relevant section of the confi
|
||||
run-time values.
|
||||
|
||||
Patroni nodes are dumping the state of the DCS options to disk upon for every change of the configuration into the file ``patroni.dynamic.json`` located in the Postgres data directory. Only the master is allowed to restore these options from the on-disk dump if these are completely absent from the DCS or if they are invalid.
|
||||
|
||||
REST API
|
||||
========
|
||||
|
||||
We provide a REST API endpoint for working with dynamic configuration.
|
||||
|
||||
GET /config
|
||||
-----------
|
||||
Get current version of dynamic configuration.
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s localhost:8008/config | jq .
|
||||
{
|
||||
"ttl": 30,
|
||||
"loop_wait": 10,
|
||||
"retry_timeout": 10,
|
||||
"maximum_lag_on_failover": 1048576,
|
||||
"postgresql": {
|
||||
"use_slots": true,
|
||||
"use_pg_rewind": true,
|
||||
"parameters": {
|
||||
"hot_standby": "on",
|
||||
"wal_log_hints": "on",
|
||||
"wal_keep_segments": 8,
|
||||
"wal_level": "hot_standby",
|
||||
"max_wal_senders": 5,
|
||||
"max_replication_slots": 5,
|
||||
"max_connections": "100"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
PATCH /config
|
||||
-------------
|
||||
Change existing configuration.
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s -XPATCH -d \
|
||||
'{"loop_wait":5,"ttl":20,"postgresql":{"parameters":{"max_connections":"101"}}}' \
|
||||
http://localhost:8008/config | jq .
|
||||
{
|
||||
"ttl": 20,
|
||||
"loop_wait": 5,
|
||||
"maximum_lag_on_failover": 1048576,
|
||||
"retry_timeout": 10,
|
||||
"postgresql": {
|
||||
"use_slots": true,
|
||||
"use_pg_rewind": true,
|
||||
"parameters": {
|
||||
"hot_standby": "on",
|
||||
"wal_log_hints": "on",
|
||||
"wal_keep_segments": 8,
|
||||
"wal_level": "hot_standby",
|
||||
"max_wal_senders": 5,
|
||||
"max_replication_slots": 5,
|
||||
"max_connections": "101"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
The above REST API call patches the existing configuration and returns the new configuration.
|
||||
|
||||
Let's check that the node processed this configuration. First of all it should start printing log lines every 5 seconds (loop_wait=5). The change of "max_connections" requires a restart, so the "restart_pending" flag should be exposed:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s http://localhost:8008/patroni | jq .
|
||||
{
|
||||
"pending_restart": true,
|
||||
"database_system_identifier": "6287881213849985952",
|
||||
"postmaster_start_time": "2016-06-13 13:13:05.211 CEST",
|
||||
"xlog": {
|
||||
"location": 2197818976
|
||||
},
|
||||
"patroni": {
|
||||
"scope": "batman",
|
||||
"version": "1.0"
|
||||
},
|
||||
"state": "running",
|
||||
"role": "master",
|
||||
"server_version": 90503
|
||||
}
|
||||
|
||||
Removing parameters:
|
||||
|
||||
If you want to remove (reset) some setting just patch it with ``null``:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s -XPATCH -d \
|
||||
'{"postgresql":{"parameters":{"max_connections":null}}}' \
|
||||
http://localhost:8008/config | jq .
|
||||
{
|
||||
"ttl": 20,
|
||||
"loop_wait": 5,
|
||||
"retry_timeout": 10,
|
||||
"maximum_lag_on_failover": 1048576,
|
||||
"postgresql": {
|
||||
"use_slots": true,
|
||||
"use_pg_rewind": true,
|
||||
"parameters": {
|
||||
"hot_standby": "on",
|
||||
"unix_socket_directories": ".",
|
||||
"wal_keep_segments": 8,
|
||||
"wal_level": "hot_standby",
|
||||
"wal_log_hints": "on",
|
||||
"max_wal_senders": 5,
|
||||
"max_replication_slots": 5
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
Above call removes ``postgresql.parameters.max_connections`` from the dynamic configuration.
|
||||
|
||||
PUT /config
|
||||
-----------
|
||||
|
||||
It's also possible to perform the full rewrite of an existing dynamic configuration unconditionally:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s -XPUT -d \
|
||||
'{"maximum_lag_on_failover":1048576,"retry_timeout":10,"postgresql":{"use_slots":true,"use_pg_rewind":true,"parameters":{"hot_standby":"on","wal_log_hints":"on","wal_keep_segments":8,"wal_level":"hot_standby","unix_socket_directories":".","max_wal_senders":5}},"loop_wait":3,"ttl":20}' \
|
||||
http://localhost:8008/config | jq .
|
||||
{
|
||||
"ttl": 20,
|
||||
"maximum_lag_on_failover": 1048576,
|
||||
"retry_timeout": 10,
|
||||
"postgresql": {
|
||||
"use_slots": true,
|
||||
"parameters": {
|
||||
"hot_standby": "on",
|
||||
"unix_socket_directories": ".",
|
||||
"wal_keep_segments": 8,
|
||||
"wal_level": "hot_standby",
|
||||
"wal_log_hints": "on",
|
||||
"max_wal_senders": 5
|
||||
},
|
||||
"use_pg_rewind": true
|
||||
},
|
||||
"loop_wait": 3
|
||||
}
|
||||
|
||||
@@ -19,6 +19,7 @@ We call Patroni a "template" because it is far from being a one-size-fits-all or
|
||||
|
||||
README
|
||||
dynamic_configuration
|
||||
rest_api
|
||||
ENVIRONMENT
|
||||
SETTINGS
|
||||
replica_bootstrap
|
||||
|
||||
+377
-1
@@ -3,6 +3,382 @@
|
||||
Release notes
|
||||
=============
|
||||
|
||||
Version 1.6.5
|
||||
-------------
|
||||
|
||||
**New features**
|
||||
|
||||
- Master stop timeout (Krishna Sarabu)
|
||||
|
||||
The number of seconds Patroni is allowed to wait when stopping Postgres. Effective only when ``synchronous_mode`` is enabled. When set to value greater than 0 and the ``synchronous_mode`` is enabled, Patroni sends ``SIGKILL`` to the postmaster if the stop operation is running for more than the value set by ``master_stop_timeout``. Set the value according to your durability/availability tradeoff. If the parameter is not set or set to non-positive value, ``master_stop_timeout`` does not have an effect.
|
||||
|
||||
- Don't create permanent physical slot with name of the primary (Alexander Kukushkin)
|
||||
|
||||
It is a common problem that the primary recycles WAL segments while the replica is down. Now we have a good solution for static clusters, with a fixed number of nodes and names that never change. You just need to list the names of all nodes in the ``slots`` so the primary will not remove the slot when the node is down (not registered in DCS).
|
||||
|
||||
- First draft of Config Validator (Igor Yanchenko)
|
||||
|
||||
Use ``patroni --validate-config patroni.yaml`` in order to validate Patroni configuration.
|
||||
|
||||
- Possibility to configure max length of timelines history (Krishna)
|
||||
|
||||
Patroni writes the history of failovers/switchovers into the ``/history`` key in DCS. Over time the size of this key becomes big, but in most cases only the last few lines are interesting. The ``max_timelines_history`` parameter allows to specify the maximum number of timeline history items to be kept in DCS.
|
||||
|
||||
- Kazoo 2.7.0 compatibility (Danyal Prout)
|
||||
|
||||
Some non-public methods in Kazoo changed their signatures, but Patroni was relying on them.
|
||||
|
||||
|
||||
**Improvements in patronictl**
|
||||
|
||||
- Show member tags (Kostiantyn Nemchenko, Alexander)
|
||||
|
||||
Tags are configured individually for every node and there was no easy way to get an overview of them
|
||||
|
||||
- Improve members output (Alexander)
|
||||
|
||||
The redundant cluster name won't be shown anymore on every line, only in the table header.
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ patronictl list
|
||||
+ Cluster: batman (6813309862653668387) +---------+----+-----------+---------------------+
|
||||
| Member | Host | Role | State | TL | Lag in MB | Tags |
|
||||
+-------------+----------------+--------+---------+----+-----------+---------------------+
|
||||
| postgresql0 | 127.0.0.1:5432 | Leader | running | 3 | | clonefrom: true |
|
||||
| | | | | | | noloadbalance: true |
|
||||
| | | | | | | nosync: true |
|
||||
+-------------+----------------+--------+---------+----+-----------+---------------------+
|
||||
| postgresql1 | 127.0.0.1:5433 | | running | 3 | 0.0 | |
|
||||
+-------------+----------------+--------+---------+----+-----------+---------------------+
|
||||
|
||||
- Fail if a config file is specified explicitly but not found (Kaarel Moppel)
|
||||
|
||||
Previously ``patronictl`` was only reporting a ``DEBUG`` message.
|
||||
|
||||
- Solved the problem of not initialized K8s pod breaking patronictl (Alexander)
|
||||
|
||||
Patroni is relying on certain pod annotations on K8s. When one of the Patroni pods is stopping or starting there is no valid annotation yet and ``patronictl`` was failing with an exception.
|
||||
|
||||
|
||||
**Stability improvements**
|
||||
|
||||
- Apply 1 second backoff if LIST call to K8s API server failed (Alexander)
|
||||
|
||||
It is mostly necessary to avoid flooding logs, but also helps to prevent starvation of the main thread.
|
||||
|
||||
- Retry if the ``retry-after`` HTTP header is returned by K8s API (Alexander)
|
||||
|
||||
If the K8s API server is overwhelmed with requests it might ask to retry.
|
||||
|
||||
- Scrub ``KUBERNETES_`` environment from the postmaster (Feike Steenbergen)
|
||||
|
||||
The ``KUBERNETES_`` environment variables are not required for PostgreSQL, yet having them exposed to the postmaster will also expose them to backends and to regular database users (using pl/perl for example).
|
||||
|
||||
- Clean up tablespaces on reinitialize (Krishna)
|
||||
|
||||
During reinit, Patroni was removing only ``PGDATA`` and leaving user-defined tablespace directories. This is causing Patroni to loop in reinit. The previous workarond for the problem was implementing the :ref:`custom bootstrap <custom_bootstrap>` script.
|
||||
|
||||
- Explicitly execute ``CHECKPOINT`` after promote happened (Alexander)
|
||||
|
||||
It helps to reduce the time before the new primary is usable for ``pg_rewind``.
|
||||
|
||||
- Smart refresh of Etcd members (Alexander)
|
||||
|
||||
In case Patroni failed to execute a request on all members of the Etcd cluster, Patroni will re-check ``A`` or ``SRV`` records for changes of IPs/hosts before retrying the next time.
|
||||
|
||||
- Skip missing values from ``pg_controldata`` (Feike)
|
||||
|
||||
Values are missing when trying to use binaries of a version that doesn't match PGDATA. Patroni will try to start Postgres anyway, and Postgres will complain that the major version doesn't match and abort with an error.
|
||||
|
||||
|
||||
**Bugfixes**
|
||||
|
||||
- Disable SSL verification for Consul when required (Julien Riou)
|
||||
|
||||
Starting from a certain version of ``urllib3``, the ``cert_reqs`` must be explicitly set to ``ssl.CERT_NONE`` in order to effectively disable SSL verification.
|
||||
|
||||
- Avoid opening replication connection on every cycle of HA loop (Alexander)
|
||||
|
||||
Regression was introduced in 1.6.4.
|
||||
|
||||
- Call ``on_role_change`` callback on failed primary (Alexander)
|
||||
|
||||
In certain cases it could lead to the virtual IP remaining attached to the old primary. Regression was introduced in 1.4.5.
|
||||
|
||||
- Reset rewind state if postgres started after successful pg_rewind (Alexander)
|
||||
|
||||
As a result of this bug Patroni was starting up manually shut down postgres in the pause mode.
|
||||
|
||||
- Convert ``recovery_min_apply_delay`` to ``ms`` when checking ``recovery.conf``
|
||||
|
||||
Patroni was indefinitely restarting replica if ``recovery_min_apply_delay`` was configured on PostgreSQL older than 12.
|
||||
|
||||
- PyInstaller compatibility (Alexander)
|
||||
|
||||
PyInstaller freezes (packages) Python applications into stand-alone executables. The compatibility was broken when we switched to the ``spawn`` method instead of ``fork`` for ``multiprocessing``.
|
||||
|
||||
|
||||
Version 1.6.4
|
||||
-------------
|
||||
|
||||
**New features**
|
||||
|
||||
- Implemented ``--wait`` option for ``patronictl reinit`` (Igor Yanchenko)
|
||||
|
||||
Patronictl will wait for ``reinit`` to finish is the ``--wait`` option is used.
|
||||
|
||||
- Further improvements of Windows support (Igor Yanchenko, Alexander Kukushkin)
|
||||
|
||||
1. All shell scripts which are used for integration testing are rewritten in python
|
||||
2. The ``pg_ctl kill`` will be used to stop postgres on non posix systems
|
||||
3. Don't try to use unix-domain sockets
|
||||
|
||||
|
||||
**Stability improvements**
|
||||
|
||||
- Make sure ``unix_socket_directories`` and ``stats_temp_directory`` exist (Igor)
|
||||
|
||||
Upon the start of Patroni and Postgres make sure that ``unix_socket_directories`` and ``stats_temp_directory`` exist or try to create them. Patroni will exit if failed to create them.
|
||||
|
||||
- Make sure ``postgresql.pgpass`` is located in the place where Patroni has write access (Igor)
|
||||
|
||||
In case if it doesn't have a write access Patroni will exit with exception.
|
||||
|
||||
- Disable Consul ``serfHealth`` check by default (Kostiantyn Nemchenko)
|
||||
|
||||
Even in case of little network problems the failing ``serfHealth`` leads to invalidation of all sessions associated with the node. Therefore, the leader key is lost much earlier than ``ttl`` which causes unwanted restarts of replicas and maybe demotion of the primary.
|
||||
|
||||
- Configure tcp keepalives for connections to K8s API (Alexander)
|
||||
|
||||
In case if we get nothing from the socket after TTL seconds it can be considered dead.
|
||||
|
||||
- Avoid logging of passwords on user creation (Alexander)
|
||||
|
||||
If the password is rejected or logging is configured to verbose or not configured at all it might happen that the password is written into postgres logs. In order to avoid it Patroni will change ``log_statement``, ``log_min_duration_statement``, and ``log_min_error_statement`` to some safe values before doing the attempt to create/update user.
|
||||
|
||||
|
||||
**Bugfixes**
|
||||
|
||||
- Use ``restore_command`` from the ``standby_cluster`` config on cascading replicas (Alexander)
|
||||
|
||||
The ``standby_leader`` was already doing it from the beginning the feature existed. Not doing the same on replicas might prevent them from catching up with standby leader.
|
||||
|
||||
- Update timeline reported by the standby cluster (Alexander)
|
||||
|
||||
In case of timeline switch the standby cluster was correctly replicating from the primary but ``patronictl`` was reporting the old timeline.
|
||||
|
||||
- Allow certain recovery parameters be defined in the custom_conf (Alexander)
|
||||
|
||||
When doing validation of recovery parameters on replica Patroni will skip ``archive_cleanup_command``, ``promote_trigger_file``, ``recovery_end_command``, ``recovery_min_apply_delay``, and ``restore_command`` if they are not defined in the patroni config but in files other than ``postgresql.auto.conf`` or ``postgresql.conf``.
|
||||
|
||||
- Improve handling of postgresql parameters with period in its name (Alexander)
|
||||
|
||||
Such parameters could be defined by extensions where the unit is not necessarily a string. Changing the value might require a restart (for example ``pg_stat_statements.max``).
|
||||
|
||||
- Improve exception handling during shutdown (Alexander)
|
||||
|
||||
During shutdown Patroni is trying to update its status in the DCS. If the DCS is inaccessible an exception might be raised. Lack of exception handling was preventing logger thread from stopping.
|
||||
|
||||
|
||||
Version 1.6.3
|
||||
-------------
|
||||
|
||||
**Bugfixes**
|
||||
|
||||
- Don't expose password when running ``pg_rewind`` (Alexander Kukushkin)
|
||||
|
||||
Bug was introduced in the `#1301 <https://github.com/zalando/patroni/pull/1301>`__
|
||||
|
||||
- Apply connection parameters specified in the ``postgresql.authentication`` to ``pg_basebackup`` and custom replica creation methods (Alexander)
|
||||
|
||||
They were relying on url-like connection string and therefore parameters never applied.
|
||||
|
||||
|
||||
Version 1.6.2
|
||||
-------------
|
||||
|
||||
**New features**
|
||||
|
||||
- Implemented ``patroni --version`` (Igor Yanchenko)
|
||||
|
||||
It prints the current version of Patroni and exits.
|
||||
|
||||
- Set the ``user-agent`` http header for all http requests (Alexander Kukushkin)
|
||||
|
||||
Patroni is communicating with Consul, Etcd, and Kubernetes API via the http protocol. Having a specifically crafted ``user-agent`` (example: ``Patroni/1.6.2 Python/3.6.8 Linux``) might be useful for debugging and monitoring.
|
||||
|
||||
- Make it possible to configure log level for exception tracebacks (Igor)
|
||||
|
||||
If you set ``log.traceback_level=DEBUG`` the tracebacks will be visible only when ``log.level=DEBUG``. The default behavior remains the same.
|
||||
|
||||
|
||||
**Stability improvements**
|
||||
|
||||
- Avoid importing all DCS modules when searching for the module required by the config file (Alexander)
|
||||
|
||||
There is no need to import modules for Etcd, Consul, and Kubernetes if we need only e.g. Zookeeper. It helps to reduce memory usage and solves the problem of having INFO messages ``Failed to import smth``.
|
||||
|
||||
- Removed python ``requests`` module from explicit requirements (Alexander)
|
||||
|
||||
It wasn't used for anything critical, but causing a lot of problems when the new version of ``urllib3`` is released.
|
||||
|
||||
- Improve handling of ``etcd.hosts`` written as a comma-separated string instead of YAML array (Igor)
|
||||
|
||||
Previously it was failing when written in format ``host1:port1, host2:port2`` (the space character after the comma).
|
||||
|
||||
|
||||
**Usability improvements**
|
||||
|
||||
- Don't force users to choose members from an empty list in ``patronictl`` (Igor)
|
||||
|
||||
If the user provides a wrong cluster name, we will raise an exception rather than ask to choose a member from an empty list.
|
||||
|
||||
- Make the error message more helpful if the REST API cannot bind (Igor)
|
||||
|
||||
For an inexperienced user it might be hard to figure out what is wrong from the Python stacktrace.
|
||||
|
||||
|
||||
**Bugfixes**
|
||||
|
||||
- Fix calculation of ``wal_buffers`` (Alexander)
|
||||
|
||||
The base unit has been changed from 8 kB blocks to bytes in PostgreSQL 11.
|
||||
|
||||
- Use ``passfile`` in ``primary_conninfo`` only on PostgreSQL 10+ (Alexander)
|
||||
|
||||
On older versions there is no guarantee that ``passfile`` will work, unless the latest version of ``libpq`` is installed.
|
||||
|
||||
|
||||
Version 1.6.1
|
||||
-------------
|
||||
|
||||
**New features**
|
||||
|
||||
- Added ``PATRONICTL_CONFIG_FILE`` environment variable (msvechla)
|
||||
|
||||
It allows configuring the ``--config-file`` argument for ``patronictl`` from the environment.
|
||||
|
||||
- Implement ``patronictl history`` (Alexander Kukushkin)
|
||||
|
||||
It shows the history of failovers/switchovers.
|
||||
|
||||
- Pass ``-c statement_timeout=0`` in ``PGOPTIONS`` when doing ``pg_rewind`` (Alexander Kukushkin)
|
||||
|
||||
It protects from the case when ``statement_timeout`` on the server is set to some small value and one of the statements executed by pg_rewind is canceled.
|
||||
|
||||
- Allow lower values for PostgreSQL configuration (Soulou)
|
||||
|
||||
Patroni didn't allow some of the PostgreSQL configuration parameters be set smaller than some hardcoded values. Now the minimal allowed values are smaller, default values have not been changed.
|
||||
|
||||
- Allow for certificate-based authentication (Jonathan S. Katz)
|
||||
|
||||
This feature enables certificate-based authentication for superuser, replication, rewind accounts and allows the user to specify the ``sslmode`` they wish to connect with.
|
||||
|
||||
- Use the ``passfile`` in the ``primary_conninfo`` instead of password (Alexander Kukushkin)
|
||||
|
||||
It allows to avoid setting ``600`` permissions on postgresql.conf
|
||||
|
||||
- Perform ``pg_ctl reload`` regardless of config changes (Alexander Kukushkin)
|
||||
|
||||
It is possible that some config files are not controlled by Patroni. When somebody is doing a reload via the REST API or by sending SIGHUP to the Patroni process, the usual expectation is that Postgres will also be reloaded. Previously it didn't happen when there were no changes in the ``postgresql`` section of Patroni config.
|
||||
|
||||
- Compare all recovery parameters, not only ``primary_conninfo`` (Alexander Kukushkin)
|
||||
|
||||
Previously the ``check_recovery_conf()`` method was only checking whether ``primary_conninfo`` has changed, never taking into account all other recovery parameters.
|
||||
|
||||
- Make it possible to apply some recovery parameters without restart (Alexander Kukushkin)
|
||||
|
||||
Starting from PostgreSQL 12 the following recovery parameters could be changed without restart: ``archive_cleanup_command``, ``promote_trigger_file``, ``recovery_end_command``, and ``recovery_min_apply_delay``. In future Postgres releases this list will be extended and Patroni will support it automatically.
|
||||
|
||||
- Make it possible to change ``use_slots`` online (Alexander Kukushkin)
|
||||
|
||||
Previously it required restarting Patroni and removing slots manually.
|
||||
|
||||
- Remove only ``PATRONI_`` prefixed environment variables when starting up Postgres (Cody Coons)
|
||||
|
||||
It will solve a lot of problems with running different Foreign Data Wrappers.
|
||||
|
||||
|
||||
**Stability improvements**
|
||||
|
||||
- Use LIST + WATCH when working with K8s API (Alexander Kukushkin)
|
||||
|
||||
It allows to efficiently receive object changes (pods, endpoints/configmaps) and makes less stress on K8s master nodes.
|
||||
|
||||
- Improve the workflow when PGDATA is not empty during bootstrap (Alexander Kukushkin)
|
||||
|
||||
According to the ``initdb`` source code it might consider a PGDATA empty when there are only ``lost+found`` and ``.dotfiles`` in it. Now Patroni does the same. If ``PGDATA`` happens to be non-empty, and at the same time not valid from the ``pg_controldata`` point of view, Patroni will complain and exit.
|
||||
|
||||
- Avoid calling expensive ``os.listdir()`` on every HA loop (Alexander Kukushkin)
|
||||
|
||||
When the system is under IO stress, ``os.listdir()`` could take a few seconds (or even minutes) to execute, badly affecting the HA loop of Patroni. This could even cause the leader key to disappear from DCS due to the lack of updates. There is a better and less expensive way to check that the PGDATA is not empty. Now we check the presence of the ``global/pg_control`` file in the PGDATA.
|
||||
|
||||
- Some improvements in logging infrastructure (Alexander Kukushkin)
|
||||
|
||||
Previously threre was a possibility to loose the last few log lines on shutdown because the logging thread was a ``daemon`` thread.
|
||||
|
||||
- Use ``spawn`` multiprocessing start method on python 3.4+ (Maciej Kowalczyk)
|
||||
|
||||
It is a known `issue <https://bugs.python.org/issue6721>`__ in Python that threading and multiprocessing do not mix well. Switching from the default method ``fork`` to the ``spawn`` is a recommended workaround. Not doing so might result in the Postmaster starting process hanging and Patroni indefinitely reporting ``INFO: restarting after failure in progress``, while Postgres is actually up and running.
|
||||
|
||||
**Improvements in REST API**
|
||||
|
||||
- Make it possible to check client certificates in the REST API (Alexander Kukushkin)
|
||||
|
||||
If the ``verify_client`` is set to ``required``, Patroni will check client certificates for all REST API calls. When it is set to ``optional``, client certificates are checked for all unsafe REST API endpoints.
|
||||
|
||||
- Return the response code 503 for the ``GET /replica`` health check request if Postgres is not running (Alexander Anikin)
|
||||
|
||||
Postgres might spend significant time in recovery before it starts accepting client connections.
|
||||
|
||||
- Implement ``/history`` and ``/cluster`` endpoints (Alexander Kukushkin)
|
||||
|
||||
The ``/history`` endpoint shows the content of the ``history`` key in DCS. The ``/cluster`` endpoint shows all cluster members and some service info like pending and scheduled restarts or switchovers.
|
||||
|
||||
|
||||
**Improvements in Etcd support**
|
||||
|
||||
- Retry on Etcd RAFT internal error (Alexander Kukushkin)
|
||||
|
||||
When the Etcd node is being shut down, it sends ``response code=300, data='etcdserver: server stopped'``, which was causing Patroni to demote the primary.
|
||||
|
||||
- Don't give up on Etcd request retry too early (Alexander Kukushkin)
|
||||
|
||||
When there were some network problems, Patroni was quickly exhausting the list of Etcd nodes and giving up without using the whole ``retry_timeout``, potentially resulting in demoting the primary.
|
||||
|
||||
|
||||
**Bugfixes**
|
||||
|
||||
- Disable ``synchronous_commit`` when granting execute permissions to the ``pg_rewind`` user (kremius)
|
||||
|
||||
If the bootstrap is done with ``synchronous_mode_strict: true`` the `GRANT EXECUTE` statement was waiting indefinitely due to the non-synchronous nodes being available.
|
||||
|
||||
- Fix memory leak on python 3.7 (Alexander Kukushkin)
|
||||
|
||||
Patroni is using ``ThreadingMixIn`` to process REST API requests and python 3.7 made threads spawn for every request non-daemon by default.
|
||||
|
||||
- Fix race conditions in asynchronous actions (Alexander Kukushkin)
|
||||
|
||||
There was a chance that ``patronictl reinit --force`` could be overwritten by the attempt to recover stopped Postgres. This ended up in a situation when Patroni was trying to start Postgres while basebackup was running.
|
||||
|
||||
- Fix race condition in ``postmaster_start_time()`` method (Alexander Kukushkin)
|
||||
|
||||
If the method is executed from the REST API thread, it requires a separate cursor object to be created.
|
||||
|
||||
- Fix the problem of not promoting the sync standby that had a name contaning upper case letters (Alexander Kukushkin)
|
||||
|
||||
We converted the name to the lower case because Postgres was doing the same while comparing the ``application_name`` with the value in ``synchronous_standby_names``.
|
||||
|
||||
- Kill all children along with the callback process before starting the new one (Alexander Kukushkin)
|
||||
|
||||
Not doing so makes it hard to implement callbacks in bash and eventually can lead to the situation when two callbacks are running at the same time.
|
||||
|
||||
- Fix 'start failed' issue (Alexander Kukushkin)
|
||||
|
||||
Under certain conditions the Postgres state might be set to 'start failed' despite Postgres being up and running.
|
||||
|
||||
|
||||
Version 1.6.0
|
||||
-------------
|
||||
|
||||
@@ -40,7 +416,7 @@ This version adds compatibility with PostgreSQL 12, makes is possible to run pg_
|
||||
This functionality works similarly to ``pg_hba.conf``: if the ``postgresql.pg_ident`` is defined in the config file or DCS, Patroni will write its value to ``pg_ident.conf``, however, if ``postgresql.parameters.ident_file`` is defined, Patroni will assume that ``pg_ident`` is managed from outside and not update the file.
|
||||
|
||||
|
||||
**Improvements in REST API**
|
||||
**Improvements in REST API**
|
||||
|
||||
- Added ``/health`` endpoint (Wilfried Roset)
|
||||
|
||||
|
||||
@@ -127,7 +127,7 @@ A special ``no_master`` parameter, if defined, allows Patroni to call the replic
|
||||
running master or replicas. In that case, an empty string will be passed in a connection string. This is useful for
|
||||
restoring the formerly running cluster from the binary backup.
|
||||
|
||||
A special ``keep_data`` parameter, if defined, will instuct Patroni to not clean PGDATA folder before calling restore.
|
||||
A special ``keep_data`` parameter, if defined, will instruct Patroni to not clean PGDATA folder before calling restore.
|
||||
|
||||
A special ``no_params`` parameter, if defined, restricts passing parameters to custom command.
|
||||
|
||||
@@ -204,4 +204,4 @@ in a patroni configuration:
|
||||
Note, that these options will be applied only once during cluster bootstrap,
|
||||
and the only way to change them afterwards is through DCS.
|
||||
|
||||
If you use replication slots on the standby cluster, you must also create the corresponding replication slot on the primary cluster. It will not be done automatically by the standby cluster implementation. You can use Patroni's permenant replication slots feature on the primary cluster to maintain a replication slot with the same name as ``primary_slot_name``, or its default value if ``primary_slot_name`` is not provided.
|
||||
If you use replication slots on the standby cluster, you must also create the corresponding replication slot on the primary cluster. It will not be done automatically by the standby cluster implementation. You can use Patroni's permanent replication slots feature on the primary cluster to maintain a replication slot with the same name as ``primary_slot_name``, or its default value if ``primary_slot_name`` is not provided.
|
||||
|
||||
@@ -46,7 +46,7 @@ When ``synchronous_mode`` is on and a standby crashes, commits will block until
|
||||
|
||||
When it is absolutely necessary to guarantee that each write is stored durably
|
||||
on at least two nodes, enable ``synchronous_mode_strict`` in addition to the
|
||||
``synchronous_node``. This parameter prevents Patroni from switching off the
|
||||
``synchronous_mode``. This parameter prevents Patroni from switching off the
|
||||
synchronous replication on the primary when no synchronous standby candidates
|
||||
are available. As a downside, the primary is not be available for writes
|
||||
(unless the Postgres transaction explicitly turns of ``synchronous_mode``),
|
||||
@@ -57,6 +57,8 @@ You can ensure that a standby never becomes the synchronous standby by setting `
|
||||
|
||||
Synchronous mode can be switched on and off via Patroni REST interface. See :ref:`dynamic configuration <dynamic_configuration>` for instructions.
|
||||
|
||||
Note: Because of the way synchronous replication is implemented in PostgreSQL it is still possible to lose transactions even when using ``synchronous_mode_strict``. If the PostgreSQL backend is cancelled while waiting to acknowledge replication (as a result of packet cancellation due to client timeout or backend failure) transaction changes become visible for other backends. Such changes are not yet replicated and may be lost in case of standby promotion.
|
||||
|
||||
|
||||
Synchronous mode implementation
|
||||
-------------------------------
|
||||
|
||||
@@ -0,0 +1,336 @@
|
||||
.. _rest_api:
|
||||
|
||||
Patroni REST API
|
||||
================
|
||||
|
||||
Patroni has a rich REST API, which is used by Patroni itself during the leader race, by the ``patronictl`` tool in order to perform failovers/switchovers/reinitialize/restarts/reloads, by HAProxy or any other kind of load balancer to perform HTTP health checks, and of course could also be used for monitoring. Below you will find the list of Patroni REST API endpoints.
|
||||
|
||||
Health check endpoints
|
||||
----------------------
|
||||
For all health check ``GET`` requests Patroni returns a JSON document with the status of the node, along with the HTTP status code. If you don't want or don't need the JSON document, you might consider using the ``OPTIONS`` method instead of ``GET``.
|
||||
|
||||
- The following requests to Patroni REST API will return HTTP status code **200** only when the Patroni node is running as the leader:
|
||||
|
||||
- ``GET /``
|
||||
- ``GET /master``
|
||||
- ``GET /leader``
|
||||
- ``GET /primary``
|
||||
- ``GET /read-write``
|
||||
|
||||
- ``GET /replica``: replica health check endpoint. It returns HTTP status code **200** only when the Patroni node is in the state ``running``, the role is ``replica`` and ``noloadbalance`` tag is not set.
|
||||
|
||||
- ``GET /read-only``: like the above endpoint, but also includes the primary.
|
||||
|
||||
- ``GET /standby-leader``: returns HTTP status code **200** only when the Patroni node is running as the leader in a :ref:`standby cluster <standby_cluster>`.
|
||||
|
||||
- ``GET /synchronous`` or ``GET /sync``: returns HTTP status code **200** only when the Patroni node is running as a synchronous standby.
|
||||
|
||||
- ``GET /asynchronous`` or ``GET /async``: returns HTTP status code **200** only when the Patroni node is running as an asynchronous standby.
|
||||
|
||||
- ``GET /health``: returns HTTP status code **200** only when PostgreSQL is up and running.
|
||||
|
||||
Monitoring endpoint
|
||||
-------------------
|
||||
|
||||
The ``GET /patroni`` is used by Patroni during the leader race. It also could be used by your monitoring system. The JSON document produced by this endpoint has the same structure as the JSON produced by the health check endpoints.
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s http://localhost:8008/patroni | jq .
|
||||
{
|
||||
"state": "running",
|
||||
"postmaster_start_time": "2019-09-24 09:22:32.555 CEST",
|
||||
"role": "master",
|
||||
"server_version": 110005,
|
||||
"cluster_unlocked": false,
|
||||
"xlog": {
|
||||
"location": 25624640
|
||||
},
|
||||
"timeline": 3,
|
||||
"database_system_identifier": "6739877027151648096",
|
||||
"patroni": {
|
||||
"version": "1.6.0",
|
||||
"scope": "batman"
|
||||
}
|
||||
}
|
||||
|
||||
Cluster status endpoints
|
||||
------------------------
|
||||
|
||||
- The ``GET /cluster`` endpoint generates a JSON document describing the current cluster topology and state:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s http://localhost:8008/cluster | jq .
|
||||
{
|
||||
"members": [
|
||||
{
|
||||
"name": "postgresql0",
|
||||
"host": "127.0.0.1",
|
||||
"port": 5432,
|
||||
"role": "leader",
|
||||
"state": "running",
|
||||
"api_url": "http://127.0.0.1:8008/patroni",
|
||||
"timeline": 5,
|
||||
"tags": {
|
||||
"clonefrom": true
|
||||
}
|
||||
},
|
||||
{
|
||||
"name": "postgresql1",
|
||||
"host": "127.0.0.1",
|
||||
"port": 5433,
|
||||
"role": "replica",
|
||||
"state": "running",
|
||||
"api_url": "http://127.0.0.1:8009/patroni",
|
||||
"timeline": 5,
|
||||
"tags": {
|
||||
"clonefrom": true
|
||||
},
|
||||
"lag": 0
|
||||
}
|
||||
],
|
||||
"scheduled_switchover": {
|
||||
"at": "2019-09-24T10:36:00+02:00",
|
||||
"from": "postgresql0"
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
- The ``GET /history`` endpoint provides a view on the history of cluster switchovers/failovers. The format is very similar to the content of history files in the ``pg_wal`` directory. The only difference is the timestamp field showing when the new timeline was created.
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s http://localhost:8008/history | jq .
|
||||
[
|
||||
[
|
||||
1,
|
||||
25623960,
|
||||
"no recovery target specified",
|
||||
"2019-09-23T16:57:57+02:00"
|
||||
],
|
||||
[
|
||||
2,
|
||||
25624344,
|
||||
"no recovery target specified",
|
||||
"2019-09-24T09:22:33+02:00"
|
||||
],
|
||||
[
|
||||
3,
|
||||
25624752,
|
||||
"no recovery target specified",
|
||||
"2019-09-24T09:26:15+02:00"
|
||||
],
|
||||
[
|
||||
4,
|
||||
50331856,
|
||||
"no recovery target specified",
|
||||
"2019-09-24T09:35:52+02:00"
|
||||
]
|
||||
]
|
||||
|
||||
|
||||
Config endpoint
|
||||
---------------
|
||||
|
||||
``GET /config``: Get the current version of the dynamic configuration:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s localhost:8008/config | jq .
|
||||
{
|
||||
"ttl": 30,
|
||||
"loop_wait": 10,
|
||||
"retry_timeout": 10,
|
||||
"maximum_lag_on_failover": 1048576,
|
||||
"postgresql": {
|
||||
"use_slots": true,
|
||||
"use_pg_rewind": true,
|
||||
"parameters": {
|
||||
"hot_standby": "on",
|
||||
"wal_log_hints": "on",
|
||||
"wal_keep_segments": 8,
|
||||
"wal_level": "hot_standby",
|
||||
"max_wal_senders": 5,
|
||||
"max_replication_slots": 5,
|
||||
"max_connections": "100"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
``PATCH /config``: Change the existing configuration.
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s -XPATCH -d \
|
||||
'{"loop_wait":5,"ttl":20,"postgresql":{"parameters":{"max_connections":"101"}}}' \
|
||||
http://localhost:8008/config | jq .
|
||||
{
|
||||
"ttl": 20,
|
||||
"loop_wait": 5,
|
||||
"maximum_lag_on_failover": 1048576,
|
||||
"retry_timeout": 10,
|
||||
"postgresql": {
|
||||
"use_slots": true,
|
||||
"use_pg_rewind": true,
|
||||
"parameters": {
|
||||
"hot_standby": "on",
|
||||
"wal_log_hints": "on",
|
||||
"wal_keep_segments": 8,
|
||||
"wal_level": "hot_standby",
|
||||
"max_wal_senders": 5,
|
||||
"max_replication_slots": 5,
|
||||
"max_connections": "101"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
The above REST API call patches the existing configuration and returns the new configuration.
|
||||
|
||||
Let's check that the node processed this configuration. First of all it should start printing log lines every 5 seconds (loop_wait=5). The change of "max_connections" requires a restart, so the "pending_restart" flag should be exposed:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s http://localhost:8008/patroni | jq .
|
||||
{
|
||||
"pending_restart": true,
|
||||
"database_system_identifier": "6287881213849985952",
|
||||
"postmaster_start_time": "2016-06-13 13:13:05.211 CEST",
|
||||
"xlog": {
|
||||
"location": 2197818976
|
||||
},
|
||||
"patroni": {
|
||||
"scope": "batman",
|
||||
"version": "1.0"
|
||||
},
|
||||
"state": "running",
|
||||
"role": "master",
|
||||
"server_version": 90503
|
||||
}
|
||||
|
||||
Removing parameters:
|
||||
|
||||
If you want to remove (reset) some setting just patch it with ``null``:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s -XPATCH -d \
|
||||
'{"postgresql":{"parameters":{"max_connections":null}}}' \
|
||||
http://localhost:8008/config | jq .
|
||||
{
|
||||
"ttl": 20,
|
||||
"loop_wait": 5,
|
||||
"retry_timeout": 10,
|
||||
"maximum_lag_on_failover": 1048576,
|
||||
"postgresql": {
|
||||
"use_slots": true,
|
||||
"use_pg_rewind": true,
|
||||
"parameters": {
|
||||
"hot_standby": "on",
|
||||
"unix_socket_directories": ".",
|
||||
"wal_keep_segments": 8,
|
||||
"wal_level": "hot_standby",
|
||||
"wal_log_hints": "on",
|
||||
"max_wal_senders": 5,
|
||||
"max_replication_slots": 5
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
The above call removes ``postgresql.parameters.max_connections`` from the dynamic configuration.
|
||||
|
||||
``PUT /config``: It's also possible to perform the full rewrite of an existing dynamic configuration unconditionally:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s -XPUT -d \
|
||||
'{"maximum_lag_on_failover":1048576,"retry_timeout":10,"postgresql":{"use_slots":true,"use_pg_rewind":true,"parameters":{"hot_standby":"on","wal_log_hints":"on","wal_keep_segments":8,"wal_level":"hot_standby","unix_socket_directories":".","max_wal_senders":5}},"loop_wait":3,"ttl":20}' \
|
||||
http://localhost:8008/config | jq .
|
||||
{
|
||||
"ttl": 20,
|
||||
"maximum_lag_on_failover": 1048576,
|
||||
"retry_timeout": 10,
|
||||
"postgresql": {
|
||||
"use_slots": true,
|
||||
"parameters": {
|
||||
"hot_standby": "on",
|
||||
"unix_socket_directories": ".",
|
||||
"wal_keep_segments": 8,
|
||||
"wal_level": "hot_standby",
|
||||
"wal_log_hints": "on",
|
||||
"max_wal_senders": 5
|
||||
},
|
||||
"use_pg_rewind": true
|
||||
},
|
||||
"loop_wait": 3
|
||||
}
|
||||
|
||||
|
||||
Switchover and failover endpoints
|
||||
---------------------------------
|
||||
|
||||
``POST /switchover`` or ``POST /failover``. These endpoints are very similar to each other. There are a couple of minor differences though:
|
||||
|
||||
1. The failover endpoint allows to perform a manual failover when there are no healthy nodes, but at the same time it will not allow you to schedule a switchover.
|
||||
|
||||
2. The switchover endpoint is the opposite. It works only when the cluster is healthy (there is a leader) and allows to schedule a switchover at a given time.
|
||||
|
||||
|
||||
In the JSON body of the ``POST`` request you must specify at least the ``leader`` or ``candidate`` fields and optionally the ``scheduled_at`` field if you want to schedule a switchover at a specific time.
|
||||
|
||||
|
||||
Example: perform a failover to the specific node:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s http://localhost:8009/failover -XPOST -d '{"candidate":"postgresql1"}'
|
||||
Successfully failed over to "postgresql1"
|
||||
|
||||
|
||||
Example: schedule a switchover from the leader to any other healthy replica in the cluster at a specific time:
|
||||
|
||||
.. code-block:: bash
|
||||
|
||||
$ curl -s http://localhost:8008/switchover -XPOST -d \
|
||||
'{"leader":"postgresql0","scheduled_at":"2019-09-24T12:00+00"}'
|
||||
Switchover scheduled
|
||||
|
||||
|
||||
Depending on the situation the request might finish with a different HTTP status code and body. The status code **200** is returned when the switchover or failover successfully completed. If the switchover was successfully scheduled, Patroni will return HTTP status code **202**. In case something went wrong, the error status code (one of **400**, **412** or **503**) will be returned with some details in the response body. For more information please check the source code of ``patroni/api.py:do_POST_failover()`` method.
|
||||
|
||||
The switchover and failover endpoints are used by ``patronictl switchover`` and ``patronictl failover``, respectively.
|
||||
|
||||
|
||||
Restart endpoint
|
||||
----------------
|
||||
|
||||
- ``POST /restart``: You can restart Postgres on the specific node by performing the ``POST /restart`` call. In the JSON body of ``POST`` request it is possible to optionally specify some restart conditions:
|
||||
|
||||
- **restart_pending**: boolean, if set to ``true`` Patroni will restart PostgreSQL only when restart is pending in order to apply some changes in the PostgreSQL config.
|
||||
- **role**: perform restart only if the current role of the node matches with the role from the POST request.
|
||||
- **postgres_version**: perform restart only if the current version of postgres is smaller than specified in the POST request.
|
||||
- **timeout**: how long we should wait before PostgreSQL starts accepting connections. Overrides ``master_start_timeout``.
|
||||
- **schedule**: timestamp with time zone, schedule the restart somewhere in the future.
|
||||
|
||||
- ``DELETE /restart``: delete the scheduled restart
|
||||
|
||||
``POST /restart`` and ``DELETE /restart`` endpoints are used by ``patronictl restart`` and ``patronictl flush`` respectively.
|
||||
|
||||
|
||||
Reload endpoint
|
||||
---------------
|
||||
|
||||
The ``POST /reload`` call will order Patroni to re-read and apply the configuration file. This is the equivalent of sending the ``SIGHUP`` signal to the Patroni process. In case you changed some of the Postgres parameters which require a restart (like **shared_buffers**), you still have to explicitly do the restart of Postgres by either calling the ``POST /restart`` endpoint or with the help of ``patronictl restart``.
|
||||
|
||||
The reload endpoint is used by ``patronictl reload``.
|
||||
|
||||
|
||||
Reinitialize endpoint
|
||||
---------------------
|
||||
|
||||
``POST /reinitialize``: reinitialize the PostgreSQL data directory on the specified node. It is allowed to be executed only on replicas. Once called, it will remove the data directory and start ``pg_basebackup`` or some alternative :ref:`replica creation method <custom_replica_creation>`.
|
||||
|
||||
The call might fail if Patroni is in a loop trying to recover (restart) a failed Postgres. In order to overcome this problem one can specify ``{"force":true}`` in the request body.
|
||||
|
||||
The reinitialize endpoint is used by ``patronictl reinit``.
|
||||
@@ -0,0 +1,21 @@
|
||||
#!/usr/bin/env python
|
||||
import os
|
||||
import argparse
|
||||
import shutil
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--dirname", required=True)
|
||||
parser.add_argument("--pathname", required=True)
|
||||
parser.add_argument("--filename", required=True)
|
||||
parser.add_argument("--mode", required=True, choices=("archive", "restore"))
|
||||
args, _ = parser.parse_known_args()
|
||||
|
||||
full_filename = os.path.join(args.dirname, args.filename)
|
||||
if args.mode == "archive":
|
||||
if not os.path.isdir(args.dirname):
|
||||
os.makedirs(args.dirname)
|
||||
if not os.path.exists(full_filename):
|
||||
shutil.copy(args.pathname, full_filename)
|
||||
else:
|
||||
shutil.copy(full_filename, args.pathname)
|
||||
Executable
+14
@@ -0,0 +1,14 @@
|
||||
#!/usr/bin/env python
|
||||
import argparse
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--datadir", required=True)
|
||||
parser.add_argument("--dbname", required=True)
|
||||
parser.add_argument("--walmethod", required=True, choices=("fetch", "stream", "none"))
|
||||
args, _ = parser.parse_known_args()
|
||||
|
||||
walmethod = ["-X", args.walmethod] if args.walmethod != "none" else []
|
||||
sys.exit(subprocess.call(["pg_basebackup", "-D", args.datadir, "-c", "fast", "-d", args.dbname] + walmethod))
|
||||
@@ -1,22 +0,0 @@
|
||||
#!/bin/bash
|
||||
|
||||
while getopts ":-:" optchar; do
|
||||
[[ "${optchar}" == "-" ]] || continue
|
||||
case "${OPTARG}" in
|
||||
datadir=* )
|
||||
PGDATA=${OPTARG#*=}
|
||||
;;
|
||||
dbname=* )
|
||||
DBNAME=${OPTARG#*=}
|
||||
;;
|
||||
walmethod=* )
|
||||
WALMETHOD=${OPTARG#*=}
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
[[ -z $PGDATA || -z $DBNAME || -z $WALMETHOD ]] && exit 1
|
||||
|
||||
[[ $WALMETHOD != "none" ]] && WALMETHOD="-X $WALMETHOD" || WALMETHOD=""
|
||||
|
||||
exec pg_basebackup -D $PGDATA $WALMETHOD -c fast -d $DBNAME
|
||||
Executable
+11
@@ -0,0 +1,11 @@
|
||||
#!/usr/bin/env python
|
||||
import argparse
|
||||
import shutil
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--datadir", required=True)
|
||||
parser.add_argument("--sourcedir", required=True)
|
||||
args, _ = parser.parse_known_args()
|
||||
|
||||
shutil.copytree(args.sourcedir, args.datadir)
|
||||
@@ -1,21 +0,0 @@
|
||||
#!/bin/bash
|
||||
|
||||
set -x
|
||||
|
||||
while getopts ":-:" optchar; do
|
||||
[[ "${optchar}" == "-" ]] || continue
|
||||
case "${OPTARG}" in
|
||||
datadir=* )
|
||||
PGDATA=${OPTARG#*=}
|
||||
;;
|
||||
sourcedir=* )
|
||||
SOURCE=${OPTARG#*=}
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
[[ -z $PGDATA || -z $SOURCE ]] && exit 1
|
||||
|
||||
mkdir -p $(dirname $PGDATA)
|
||||
|
||||
exec cp -af $SOURCE $PGDATA
|
||||
@@ -4,6 +4,7 @@ Feature: basic replication
|
||||
Scenario: check replication of a single table
|
||||
Given I start postgres0
|
||||
Then postgres0 is a leader after 10 seconds
|
||||
And there is a non empty initialize key in DCS after 15 seconds
|
||||
When I issue a PATCH request to http://127.0.0.1:8008/config with {"ttl": 20, "loop_wait": 2, "synchronous_mode": true}
|
||||
Then I receive a response code 200
|
||||
When I start postgres1
|
||||
@@ -35,10 +36,12 @@ Feature: basic replication
|
||||
And I run patronictl.py resume batman
|
||||
Then I receive a response returncode 0
|
||||
And postgres2 role is the primary after 24 seconds
|
||||
And Response on GET http://127.0.0.1:8010/history contains recovery after 10 seconds
|
||||
When I issue a PATCH request to http://127.0.0.1:8010/config with {"synchronous_mode": null, "master_start_timeout": 0}
|
||||
Then I receive a response code 200
|
||||
When I add the table bar to postgres2
|
||||
Then table bar is present on postgres1 after 20 seconds
|
||||
And Response on GET http://127.0.0.1:8010/config contains master_start_timeout after 10 seconds
|
||||
|
||||
Scenario: check immediate failover when master_start_timeout=0
|
||||
Given I kill postmaster on postgres2
|
||||
|
||||
Executable
+17
@@ -0,0 +1,17 @@
|
||||
#!/usr/bin/env python
|
||||
import os
|
||||
import psycopg2
|
||||
import sys
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
if not (len(sys.argv) >= 3 and sys.argv[3] == "master"):
|
||||
sys.exit(1)
|
||||
|
||||
os.environ['PGPASSWORD'] = 'zalando'
|
||||
connection = psycopg2.connect(host='127.0.0.1', port=sys.argv[1], user='postgres')
|
||||
cursor = connection.cursor()
|
||||
cursor.execute("SELECT slot_name FROM pg_replication_slots WHERE slot_type = 'logical'")
|
||||
|
||||
with open("data/postgres0/label", "w") as label:
|
||||
label.write(next(iter(cursor.fetchone()), ""))
|
||||
@@ -1,5 +0,0 @@
|
||||
#!/bin/bash
|
||||
|
||||
[[ "$3" == "master" ]] || exit
|
||||
|
||||
PGPASSWORD=zalando psql -h localhost -U postgres -p $1 -w -tAc "SELECT slot_name FROM pg_replication_slots WHERE slot_type = 'logical'" >> data/postgres0/label
|
||||
Executable
+5
@@ -0,0 +1,5 @@
|
||||
#!/usr/bin/env python
|
||||
|
||||
import sys
|
||||
with open("data/{0}/{0}_cb.log".format(sys.argv[1]), "a+") as log:
|
||||
log.write(" ".join(sys.argv[-3:]) + "\n")
|
||||
@@ -13,5 +13,5 @@ Scenario: make a backup and do a restore into a new cluster
|
||||
Given I add the table bar to postgres1
|
||||
And I do a backup of postgres1
|
||||
When I start postgres2 in a cluster batman2 from backup
|
||||
Then postgres2 is a leader of batman2 after 10 seconds
|
||||
Then postgres2 is a leader of batman2 after 30 seconds
|
||||
And table bar is present on postgres2 after 10 seconds
|
||||
|
||||
+40
-65
@@ -1,11 +1,6 @@
|
||||
import abc
|
||||
import consul
|
||||
import datetime
|
||||
import etcd
|
||||
import kazoo.client
|
||||
import kazoo.exceptions
|
||||
import os
|
||||
import psutil
|
||||
import psycopg2
|
||||
import json
|
||||
import shutil
|
||||
@@ -176,7 +171,8 @@ class PatroniController(AbstractController):
|
||||
|
||||
config['name'] = name
|
||||
config['postgresql']['data_dir'] = self._data_dir
|
||||
config['postgresql']['use_unix_socket'] = True
|
||||
config['postgresql']['use_unix_socket'] = os.name != 'nt' # windows doesn't yet support unix-domain sockets
|
||||
config['postgresql']['pgpass'] = os.path.join(tempfile.gettempdir(), 'pgpass_' + name)
|
||||
config['postgresql']['parameters'].update({
|
||||
'logging_collector': 'on', 'log_destination': 'csvlog', 'log_directory': self._output_dir,
|
||||
'log_filename': name + '.log', 'log_statement': 'all', 'log_min_messages': 'debug1',
|
||||
@@ -249,59 +245,24 @@ class PatroniController(AbstractController):
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
def database_is_running(self):
|
||||
pid = self._get_pid()
|
||||
if not pid:
|
||||
return False
|
||||
try:
|
||||
os.kill(pid, 0)
|
||||
except OSError:
|
||||
return False
|
||||
return True
|
||||
|
||||
def patroni_hang(self, timeout):
|
||||
hang = ProcessHang(self._handle.pid, timeout)
|
||||
self._closables.append(hang)
|
||||
hang.start()
|
||||
|
||||
def checkpoint_hang(self, timeout):
|
||||
pid = self._get_pid()
|
||||
if not pid:
|
||||
return False
|
||||
proc = psutil.Process(pid)
|
||||
for child in proc.children():
|
||||
if 'checkpoint' in child.cmdline()[0]:
|
||||
checkpointer = child
|
||||
break
|
||||
else:
|
||||
return False
|
||||
hang = ProcessHang(checkpointer.pid, timeout)
|
||||
self._closables.append(hang)
|
||||
hang.start()
|
||||
return True
|
||||
|
||||
def cancel_background(self):
|
||||
for obj in self._closables:
|
||||
obj.close()
|
||||
self._closables = []
|
||||
|
||||
def terminate_backends(self):
|
||||
pid = self._get_pid()
|
||||
if not pid:
|
||||
return False
|
||||
proc = psutil.Process(pid)
|
||||
for p in proc.children():
|
||||
if 'process' not in p.cmdline()[0]:
|
||||
p.terminate()
|
||||
|
||||
@property
|
||||
def backup_source(self):
|
||||
return 'postgres://{username}:{password}@{host}:{port}/{database}'.format(**self._replication)
|
||||
|
||||
def backup(self, dest='data/basebackup'):
|
||||
subprocess.call([PatroniPoolController.BACKUP_SCRIPT, '--walmethod=none',
|
||||
'--datadir=' + os.path.join(self._work_directory, dest),
|
||||
'--dbname=' + self.backup_source])
|
||||
def backup(self, dest=os.path.join('data', 'basebackup')):
|
||||
subprocess.call(PatroniPoolController.BACKUP_SCRIPT + ['--walmethod=none',
|
||||
'--datadir=' + os.path.join(self._work_directory, dest),
|
||||
'--dbname=' + self.backup_source])
|
||||
|
||||
|
||||
class ProcessHang(object):
|
||||
@@ -375,9 +336,11 @@ class ConsulController(AbstractDcsController):
|
||||
super(ConsulController, self).__init__(context)
|
||||
os.environ['PATRONI_CONSUL_HOST'] = 'localhost:8500'
|
||||
os.environ['PATRONI_CONSUL_REGISTER_SERVICE'] = 'on'
|
||||
self._client = consul.Consul()
|
||||
self._config_file = None
|
||||
|
||||
import consul
|
||||
self._client = consul.Consul()
|
||||
|
||||
def _start(self):
|
||||
self._config_file = self._work_directory + '.json'
|
||||
with open(self._config_file, 'wb') as f:
|
||||
@@ -417,6 +380,8 @@ class EtcdController(AbstractDcsController):
|
||||
def __init__(self, context):
|
||||
super(EtcdController, self).__init__(context)
|
||||
os.environ['PATRONI_ETCD_HOST'] = 'localhost:2379'
|
||||
|
||||
import etcd
|
||||
self._client = etcd.Client(port=2379)
|
||||
|
||||
def _start(self):
|
||||
@@ -424,12 +389,14 @@ class EtcdController(AbstractDcsController):
|
||||
stdout=self._log, stderr=subprocess.STDOUT)
|
||||
|
||||
def query(self, key, scope='batman'):
|
||||
import etcd
|
||||
try:
|
||||
return self._client.get(self.path(key, scope)).value
|
||||
except etcd.EtcdKeyNotFound:
|
||||
return None
|
||||
|
||||
def cleanup_service_tree(self):
|
||||
import etcd
|
||||
try:
|
||||
self._client.delete(self.path(scope=''), recursive=True)
|
||||
except (etcd.EtcdKeyNotFound, etcd.EtcdConnectionFailed):
|
||||
@@ -473,13 +440,13 @@ class KubernetesController(AbstractDcsController):
|
||||
|
||||
def delete_pod(self, name):
|
||||
try:
|
||||
self._api.delete_namespaced_pod(name, self._namespace, self._client.V1DeleteOptions())
|
||||
except:
|
||||
self._api.delete_namespaced_pod(name, self._namespace, body=self._client.V1DeleteOptions())
|
||||
except Exception:
|
||||
pass
|
||||
while True:
|
||||
try:
|
||||
self._api.read_namespaced_pod(name, self._namespace)
|
||||
except:
|
||||
except Exception:
|
||||
break
|
||||
|
||||
def query(self, key, scope='batman'):
|
||||
@@ -488,22 +455,23 @@ class KubernetesController(AbstractDcsController):
|
||||
return (pod.metadata.annotations or {}).get('status', '')
|
||||
else:
|
||||
try:
|
||||
e = self._api.read_namespaced_endpoints(scope + ('' if key == 'leader' else '-' + key), self._namespace)
|
||||
if key == 'leader':
|
||||
ep = scope + {'leader': '', 'history': '-config', 'initialize': '-config'}.get(key, '-' + key)
|
||||
e = self._api.read_namespaced_endpoints(ep, self._namespace)
|
||||
if key != 'sync':
|
||||
return e.metadata.annotations[key]
|
||||
else:
|
||||
return json.dumps(e.metadata.annotations)
|
||||
except:
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
def cleanup_service_tree(self):
|
||||
try:
|
||||
self._api.delete_collection_namespaced_pod(self._namespace, label_selector=self._label_selector)
|
||||
except:
|
||||
except Exception:
|
||||
pass
|
||||
try:
|
||||
self._api.delete_collection_namespaced_endpoints(self._namespace, label_selector=self._label_selector)
|
||||
except:
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
while True:
|
||||
@@ -523,18 +491,22 @@ class ZooKeeperController(AbstractDcsController):
|
||||
super(ZooKeeperController, self).__init__(context, False)
|
||||
if export_env:
|
||||
os.environ['PATRONI_ZOOKEEPER_HOSTS'] = "'localhost:2181'"
|
||||
|
||||
import kazoo.client
|
||||
self._client = kazoo.client.KazooClient()
|
||||
|
||||
def _start(self):
|
||||
pass # TODO: implement later
|
||||
|
||||
def query(self, key, scope='batman'):
|
||||
import kazoo.exceptions
|
||||
try:
|
||||
return self._client.get(self.path(key, scope))[0].decode('utf-8')
|
||||
except kazoo.exceptions.NoNodeError:
|
||||
return None
|
||||
|
||||
def cleanup_service_tree(self):
|
||||
import kazoo.exceptions
|
||||
try:
|
||||
self._client.delete(self.path(scope=''), recursive=True)
|
||||
except (kazoo.exceptions.NoNodeError):
|
||||
@@ -561,7 +533,8 @@ class ExhibitorController(ZooKeeperController):
|
||||
|
||||
class PatroniPoolController(object):
|
||||
|
||||
BACKUP_SCRIPT = 'features/backup_create.sh'
|
||||
BACKUP_SCRIPT = [sys.executable, 'features/backup_create.py']
|
||||
ARCHIVE_RESTORE_SCRIPT = ' '.join((sys.executable, os.path.abspath('features/archive-restore.py')))
|
||||
|
||||
def __init__(self, context):
|
||||
self._context = context
|
||||
@@ -594,9 +567,8 @@ class PatroniPoolController(object):
|
||||
self._processes[name].start(max_wait_limit)
|
||||
|
||||
def __getattr__(self, func):
|
||||
if func not in ['stop', 'query', 'write_label', 'read_label', 'check_role_has_changed_to', 'add_tag_to_config',
|
||||
'get_watchdog', 'database_is_running', 'checkpoint_hang', 'patroni_hang',
|
||||
'terminate_backends', 'backup']:
|
||||
if func not in ['stop', 'query', 'write_label', 'read_label', 'check_role_has_changed_to',
|
||||
'add_tag_to_config', 'get_watchdog', 'patroni_hang', 'backup']:
|
||||
raise AttributeError("PatroniPoolController instance has no attribute '{0}'".format(func))
|
||||
|
||||
def wrapper(name, *args, **kwargs):
|
||||
@@ -610,7 +582,7 @@ class PatroniPoolController(object):
|
||||
self._processes.clear()
|
||||
|
||||
def create_and_set_output_directory(self, feature_name):
|
||||
feature_dir = os.path.join(self.patroni_path, 'features/output', feature_name.replace(' ', '_'))
|
||||
feature_dir = os.path.join(self.patroni_path, 'features', 'output', feature_name.replace(' ', '_'))
|
||||
if os.path.exists(feature_dir):
|
||||
shutil.rmtree(feature_dir)
|
||||
os.makedirs(feature_dir)
|
||||
@@ -623,7 +595,7 @@ class PatroniPoolController(object):
|
||||
'bootstrap': {
|
||||
'method': 'pg_basebackup',
|
||||
'pg_basebackup': {
|
||||
'command': self.BACKUP_SCRIPT + ' --walmethod=stream --dbname=' + f.backup_source
|
||||
'command': " ".join(self.BACKUP_SCRIPT) + ' --walmethod=stream --dbname=' + f.backup_source
|
||||
},
|
||||
'dcs': {
|
||||
'postgresql': {
|
||||
@@ -636,8 +608,9 @@ class PatroniPoolController(object):
|
||||
'postgresql': {
|
||||
'parameters': {
|
||||
'archive_mode': 'on',
|
||||
'archive_command': 'mkdir -p {0} && test ! -f {0}/%f && cp %p {0}/%f'.format(
|
||||
os.path.join(self.patroni_path, 'data/wal_archive'))
|
||||
'archive_command': (self.ARCHIVE_RESTORE_SCRIPT + ' --mode archive ' +
|
||||
'--dirname {} --filename %f --pathname %p').format(
|
||||
os.path.join(self.patroni_path, 'data', 'wal_archive'))
|
||||
},
|
||||
'authentication': {
|
||||
'superuser': {'password': 'zalando1'},
|
||||
@@ -653,12 +626,14 @@ class PatroniPoolController(object):
|
||||
'bootstrap': {
|
||||
'method': 'backup_restore',
|
||||
'backup_restore': {
|
||||
'command': 'features/backup_restore.sh --sourcedir=' + os.path.join(self.patroni_path,
|
||||
'data/basebackup'),
|
||||
'command': (sys.executable + ' features/backup_restore.py --sourcedir=' +
|
||||
os.path.join(self.patroni_path, 'data', 'basebackup')),
|
||||
'recovery_conf': {
|
||||
'recovery_target_action': 'promote',
|
||||
'recovery_target_timeline': 'latest',
|
||||
'restore_command': 'cp {0}/data/wal_archive/%f %p'.format(self.patroni_path)
|
||||
'restore_command': (self.ARCHIVE_RESTORE_SCRIPT + ' --mode restore ' +
|
||||
'--dirname {} --filename %f --pathname %p').format(
|
||||
os.path.join(self.patroni_path, 'data', 'wal_archive'))
|
||||
}
|
||||
}
|
||||
},
|
||||
|
||||
@@ -28,10 +28,7 @@ Scenario: check API requests on a stand-alone server
|
||||
And I receive a response text "Failover could be performed only to a specific candidate"
|
||||
|
||||
Scenario: check local configuration reload
|
||||
Given I issue an empty POST request to http://127.0.0.1:8008/reload
|
||||
Then I receive a response code 200
|
||||
And I receive a response text nothing changed
|
||||
When I add tag new_tag new_value to postgres0 config
|
||||
Given I add tag new_tag new_value to postgres0 config
|
||||
And I issue an empty POST request to http://127.0.0.1:8008/reload
|
||||
Then I receive a response code 202
|
||||
|
||||
@@ -46,6 +43,20 @@ Scenario: check dynamic configuration change via DCS
|
||||
When I issue a GET request to http://127.0.0.1:8008/patroni
|
||||
Then I receive a response code 200
|
||||
And I receive a response tags {'new_tag': 'new_value'}
|
||||
And I sleep for 4 seconds
|
||||
|
||||
Scenario: check the scheduled restart
|
||||
Given I issue a PATCH request to http://127.0.0.1:8008/config with {"postgresql": {"parameters": {"superuser_reserved_connections": "6"}}}
|
||||
Then I receive a response code 200
|
||||
And Response on GET http://127.0.0.1:8008/patroni contains pending_restart after 5 seconds
|
||||
Given I issue a scheduled restart at http://127.0.0.1:8008 in 3 seconds with {"role": "replica"}
|
||||
Then I receive a response code 202
|
||||
And I sleep for 4 seconds
|
||||
And Response on GET http://127.0.0.1:8008/patroni contains pending_restart after 10 seconds
|
||||
Given I issue a scheduled restart at http://127.0.0.1:8008 in 3 seconds with {"restart_pending": "True"}
|
||||
Then I receive a response code 202
|
||||
And Response on GET http://127.0.0.1:8008/patroni does not contain pending_restart after 10 seconds
|
||||
And postgres0 role is the primary after 10 seconds
|
||||
|
||||
Scenario: check API requests for the primary-replica pair in the pause mode
|
||||
Given I run patronictl.py pause batman
|
||||
@@ -73,6 +84,7 @@ Scenario: check the switchover via the API in the pause mode
|
||||
And postgres1 role is the primary after 10 seconds
|
||||
And postgres0 role is the secondary after 10 seconds
|
||||
And replication works from postgres1 to postgres0 after 20 seconds
|
||||
And "members/postgres0" key in DCS has state=running after 10 seconds
|
||||
When I issue a GET request to http://127.0.0.1:8008/master
|
||||
Then I receive a response code 503
|
||||
When I issue a GET request to http://127.0.0.1:8008/replica
|
||||
@@ -94,6 +106,7 @@ Scenario: check the scheduled switchover
|
||||
And postgres0 role is the primary after 10 seconds
|
||||
And postgres1 role is the secondary after 10 seconds
|
||||
And replication works from postgres0 to postgres1 after 25 seconds
|
||||
And "members/postgres1" key in DCS has state=running after 10 seconds
|
||||
When I issue a GET request to http://127.0.0.1:8008/master
|
||||
Then I receive a response code 200
|
||||
When I issue a GET request to http://127.0.0.1:8008/replica
|
||||
@@ -102,16 +115,3 @@ Scenario: check the scheduled switchover
|
||||
Then I receive a response code 503
|
||||
When I issue a GET request to http://127.0.0.1:8009/replica
|
||||
Then I receive a response code 200
|
||||
|
||||
Scenario: check the scheduled restart
|
||||
Given I issue a PATCH request to http://127.0.0.1:8008/config with {"postgresql": {"parameters": {"superuser_reserved_connections": "6"}}}
|
||||
Then I receive a response code 200
|
||||
And Response on GET http://127.0.0.1:8008/patroni contains pending_restart after 5 seconds
|
||||
Given I issue a scheduled restart at http://127.0.0.1:8008 in 3 seconds with {"role": "replica"}
|
||||
Then I receive a response code 202
|
||||
And I sleep for 4 seconds
|
||||
And Response on GET http://127.0.0.1:8008/patroni contains pending_restart after 10 seconds
|
||||
Given I issue a scheduled restart at http://127.0.0.1:8008 in 3 seconds with {"restart_pending": "True"}
|
||||
Then I receive a response code 202
|
||||
And Response on GET http://127.0.0.1:8008/patroni does not contain pending_restart after 10 seconds
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@ Feature: standby cluster
|
||||
Scenario: check permanent logical slots are preserved on failover/switchover
|
||||
Given I start postgres1
|
||||
Then postgres1 is a leader after 10 seconds
|
||||
And I sleep for 3 seconds
|
||||
And there is a non empty initialize key in DCS after 15 seconds
|
||||
When I issue a PATCH request to http://127.0.0.1:8009/config with {"loop_wait": 2, "slots": {"pm_1": {"type": "physical"}}, "postgresql": {"parameters": {"wal_level": "logical"}}}
|
||||
Then I receive a response code 200
|
||||
And Response on GET http://127.0.0.1:8009/config contains slots after 10 seconds
|
||||
@@ -30,7 +30,7 @@ Feature: standby cluster
|
||||
When I issue a GET request to http://127.0.0.1:8009/standby_leader
|
||||
Then I receive a response code 200
|
||||
And I receive a response role standby_leader
|
||||
And there is a postgres1_cb.log with "on_start replica batman1\non_role_change standby_leader batman1" in postgres1 data directory
|
||||
And there is a postgres1_cb.log with "on_role_change standby_leader batman1" in postgres1 data directory
|
||||
When I start postgres2 in a cluster batman1
|
||||
Then postgres2 role is the replica after 24 seconds
|
||||
And table foo is present on postgres2 after 20 seconds
|
||||
|
||||
@@ -13,7 +13,7 @@ def start_patroni_with_a_name_value_tag(context, name, tag_name, tag_value):
|
||||
def check_label(context, label, content, name):
|
||||
label = context.pctl.read_label(name, label)
|
||||
label = label.replace('\n', '\\n')
|
||||
assert label == content, "{0} is not equal to {1}".format(label, content)
|
||||
assert content in label, "{0} doesn't contain {1}".format(label, content)
|
||||
|
||||
|
||||
@step('I create label with "{content:w}" in {name:w} data directory')
|
||||
@@ -34,3 +34,17 @@ def check_member(context, name, key, value, time_limit):
|
||||
pass
|
||||
time.sleep(1)
|
||||
assert False, "{0} does not have {1}={2} in dcs after {3} seconds".format(name, key, value, time_limit)
|
||||
|
||||
|
||||
@step('there is a non empty {key:w} key in DCS after {time_limit:d} seconds')
|
||||
def check_initialize(context, key, time_limit):
|
||||
time_limit *= context.timeout_multiplier
|
||||
max_time = time.time() + int(time_limit)
|
||||
while time.time() < max_time:
|
||||
try:
|
||||
if context.dcs_ctl.query(key):
|
||||
return
|
||||
except Exception:
|
||||
pass
|
||||
time.sleep(1)
|
||||
assert False, "There is no {0} in dcs after {1} seconds".format(key, time_limit)
|
||||
|
||||
@@ -1,8 +1,6 @@
|
||||
import base64
|
||||
import json
|
||||
import os
|
||||
import parse
|
||||
import requests
|
||||
import shlex
|
||||
import subprocess
|
||||
import sys
|
||||
@@ -12,8 +10,10 @@ import yaml
|
||||
from behave import register_type, step, then
|
||||
from dateutil import tz
|
||||
from datetime import datetime, timedelta
|
||||
from patroni.request import PatroniRequest
|
||||
|
||||
tzutc = tz.tzutc()
|
||||
request_executor = PatroniRequest({'ctl': {'auth': 'username:password'}})
|
||||
|
||||
|
||||
@parse.with_pattern(r'https?://(?:\w|\.|:|/)+')
|
||||
@@ -45,9 +45,9 @@ def sleep_for_n_seconds(context, value):
|
||||
|
||||
|
||||
def _set_response(context, response):
|
||||
context.status_code = response.status_code
|
||||
data = response.content.decode('utf-8')
|
||||
ct = response.headers.get('content-type', '')
|
||||
context.status_code = response.status
|
||||
data = response.data.decode('utf-8')
|
||||
ct = response.getheader('content-type', '')
|
||||
if ct.startswith('application/json') or\
|
||||
ct.startswith('text/yaml') or\
|
||||
ct.startswith('text/x-yaml') or\
|
||||
@@ -63,13 +63,7 @@ def _set_response(context, response):
|
||||
|
||||
@step('I issue a GET request to {url:url}')
|
||||
def do_get(context, url):
|
||||
try:
|
||||
r = requests.get(url)
|
||||
except requests.exceptions.RequestException:
|
||||
context.status_code = None
|
||||
context.response = None
|
||||
else:
|
||||
_set_response(context, r)
|
||||
do_request(context, 'GET', url, None)
|
||||
|
||||
|
||||
@step('I issue an empty POST request to {url:url}')
|
||||
@@ -79,17 +73,11 @@ def do_post_empty(context, url):
|
||||
|
||||
@step('I issue a {request_method:w} request to {url:url} with {data}')
|
||||
def do_request(context, request_method, url, data):
|
||||
data = data and json.loads(data) or {}
|
||||
headers = {'Authorization': 'Basic ' + base64.b64encode('username:password'.encode('utf-8')).decode('utf-8'),
|
||||
'Content-Type': 'application/json'}
|
||||
data = data and json.loads(data)
|
||||
try:
|
||||
if request_method == 'PATCH':
|
||||
r = requests.patch(url, headers=headers, json=data)
|
||||
else:
|
||||
r = requests.post(url, headers=headers, json=data)
|
||||
except requests.exceptions.RequestException:
|
||||
context.status_code = None
|
||||
context.response = None
|
||||
r = request_executor.request(request_method, url, data)
|
||||
except Exception:
|
||||
context.status_code = context.response = None
|
||||
else:
|
||||
_set_response(context, r)
|
||||
|
||||
@@ -149,8 +137,8 @@ def add_tag_to_config(context, tag, value, pg_name):
|
||||
def check_http_response(context, url, value, timeout, negate=False):
|
||||
timeout *= context.timeout_multiplier
|
||||
for _ in range(int(timeout)):
|
||||
r = requests.get(url)
|
||||
if (value in r.content.decode('utf-8')) != negate:
|
||||
r = request_executor.request('GET', url)
|
||||
if (value in r.data.decode('utf-8')) != negate:
|
||||
break
|
||||
time.sleep(1)
|
||||
else:
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
|
||||
from behave import step
|
||||
@@ -9,7 +10,7 @@ SELECT * FROM pg_catalog.pg_stat_replication
|
||||
WHERE application_name = '{0}'
|
||||
"""
|
||||
|
||||
callback = "bash -c 'echo \"${*: -3:1} ${*: -2:1} ${*: -1:1}\" >> data/$1/$1_cb.log' -- "
|
||||
callback = sys.executable + " features/callback2.py "
|
||||
|
||||
|
||||
@step('I start {name:w} with callback configured')
|
||||
@@ -17,7 +18,7 @@ def start_patroni_with_callbacks(context, name):
|
||||
return context.pctl.start(name, custom_config={
|
||||
"postgresql": {
|
||||
"callbacks": {
|
||||
"on_role_change": "features/callback.sh"
|
||||
"on_role_change": sys.executable + " features/callback.py"
|
||||
}
|
||||
}
|
||||
})
|
||||
@@ -30,8 +31,8 @@ def start_patroni(context, name, cluster_name):
|
||||
"postgresql": {
|
||||
"callbacks": {c: callback + name for c in ('on_start', 'on_stop', 'on_restart', 'on_role_change')},
|
||||
"backup_restore": {
|
||||
"command": "features/backup_restore.sh --sourcedir=" + os.path.join(context.pctl.patroni_path,
|
||||
"data/basebackup")}
|
||||
"command": (sys.executable + " features/backup_restore.py --sourcedir=" +
|
||||
os.path.join(context.pctl.patroni_path, 'data', 'basebackup'))}
|
||||
}
|
||||
})
|
||||
|
||||
|
||||
@@ -44,31 +44,6 @@ def watchdog_was_triggered(context, name, timeout):
|
||||
assert False
|
||||
|
||||
|
||||
@then('{name:w} watchdog was not triggered')
|
||||
def watchdog_was_not_triggered(context, name):
|
||||
assert not context.pctl.get_watchdog(name).was_triggered
|
||||
|
||||
|
||||
@step('{name:w} checkpoint takes {timeout:d} seconds')
|
||||
def checkpoint_hang(context, name, timeout):
|
||||
assert context.pctl.checkpoint_hang(name, timeout)
|
||||
|
||||
|
||||
@step('{name:w} hangs for {timeout:d} seconds')
|
||||
def patroni_hang(context, name, timeout):
|
||||
return context.pctl.patroni_hang(name, timeout)
|
||||
|
||||
|
||||
@step('I terminate {name:w} user processes')
|
||||
def terminate_backends(context, name):
|
||||
return context.pctl.terminate_backends(name)
|
||||
|
||||
|
||||
@step('Sleep for {timeout:d} seconds')
|
||||
def dcs_connection_lost(context, timeout):
|
||||
time.sleep(timeout)
|
||||
|
||||
|
||||
@then('{name:w} database is running')
|
||||
def database_is_running(context, name):
|
||||
assert context.pctl.database_is_running(name)
|
||||
|
||||
@@ -106,6 +106,20 @@ objects:
|
||||
application: ${APPLICATION_NAME}
|
||||
cluster-name: ${PATRONI_CLUSTER_NAME}
|
||||
spec:
|
||||
initContainers:
|
||||
- command:
|
||||
- sh
|
||||
- -c
|
||||
- "mkdir -p /home/postgres/pgdata/pgroot/data && chmod 0700 /home/postgres/pgdata/pgroot/data"
|
||||
image: docker-registry.default.svc:5000/${NAMESPACE}/patroni:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
name: fix-perms
|
||||
resources: {}
|
||||
terminationMessagePath: /dev/termination-log
|
||||
terminationMessagePolicy: File
|
||||
volumeMounts:
|
||||
- mountPath: /home/postgres/pgdata
|
||||
name: ${APPLICATION_NAME}
|
||||
containers:
|
||||
- env:
|
||||
- name: PATRONI_KUBERNETES_POD_IP
|
||||
|
||||
@@ -1,14 +1,6 @@
|
||||
#!/usr/bin/env python
|
||||
from patroni import main
|
||||
|
||||
import os
|
||||
|
||||
if __name__ == '__main__':
|
||||
|
||||
if os.getenv("PATRONI_DEBUG_MODE"):
|
||||
# XXX Visual Code specific https://github.com/microsoft/ptvsd/issues/1443
|
||||
# create processes by spawning new Python interpreters instead of forking the current one
|
||||
import multiprocessing
|
||||
multiprocessing.set_start_method('spawn', True)
|
||||
|
||||
main()
|
||||
|
||||
+52
-15
@@ -4,26 +4,30 @@ import signal
|
||||
import sys
|
||||
import time
|
||||
|
||||
from patroni.version import __version__
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
PATRONI_ENV_PREFIX = 'PATRONI_'
|
||||
KUBERNETES_ENV_PREFIX = 'KUBERNETES_'
|
||||
|
||||
|
||||
class Patroni(object):
|
||||
|
||||
def __init__(self):
|
||||
def __init__(self, conf):
|
||||
from patroni.api import RestApiServer
|
||||
from patroni.config import Config
|
||||
from patroni.dcs import get_dcs
|
||||
from patroni.ha import Ha
|
||||
from patroni.log import PatroniLogger
|
||||
from patroni.postgresql import Postgresql
|
||||
from patroni.version import __version__
|
||||
from patroni.request import PatroniRequest
|
||||
from patroni.watchdog import Watchdog
|
||||
|
||||
self.setup_signal_handlers()
|
||||
|
||||
self.version = __version__
|
||||
self.logger = PatroniLogger()
|
||||
self.config = Config()
|
||||
self.config = conf
|
||||
self.logger.reload_config(self.config.get('log', {}))
|
||||
self.dcs = get_dcs(self.config)
|
||||
self.watchdog = Watchdog(self.config)
|
||||
@@ -31,8 +35,8 @@ class Patroni(object):
|
||||
|
||||
self.postgresql = Postgresql(self.config['postgresql'])
|
||||
self.api = RestApiServer(self, self.config['restapi'])
|
||||
self.request = PatroniRequest(self.config, True)
|
||||
self.ha = Ha(self)
|
||||
self.is_in_debug_mode = bool(os.getenv("PATRONI_DEBUG_MODE"))
|
||||
|
||||
self.tags = self.get_tags()
|
||||
self.next_run = time.time()
|
||||
@@ -67,14 +71,16 @@ class Patroni(object):
|
||||
def nosync(self):
|
||||
return bool(self.tags.get('nosync', False))
|
||||
|
||||
def reload_config(self):
|
||||
def reload_config(self, sighup=False):
|
||||
try:
|
||||
self.tags = self.get_tags()
|
||||
self.logger.reload_config(self.config.get('log', {}))
|
||||
self.dcs.reload_config(self.config)
|
||||
self.watchdog.reload_config(self.config)
|
||||
if sighup:
|
||||
self.request.reload_config(self.config)
|
||||
self.api.reload_config(self.config['restapi'])
|
||||
self.postgresql.reload_config(self.config['postgresql'])
|
||||
self.postgresql.reload_config(self.config['postgresql'], sighup)
|
||||
self.dcs.reload_config(self.config)
|
||||
except Exception:
|
||||
logger.exception('Failed to reload config_file=%s', self.config.config_file)
|
||||
|
||||
@@ -103,9 +109,8 @@ class Patroni(object):
|
||||
self.next_run = current_time
|
||||
# Release the GIL so we don't starve anyone waiting on async_executor lock
|
||||
time.sleep(0.001)
|
||||
# Warn user that Patroni is not keeping up or runs in debug
|
||||
msg = "Patroni runs in the debug mode: keys' TTL is infinite, loop wait disabled" if self.is_in_debug_mode else "Loop time exceeded, rescheduling immediately."
|
||||
logger.warning(msg)
|
||||
# Warn user that Patroni is not keeping up
|
||||
logger.warning("Loop time exceeded, rescheduling immediately.")
|
||||
elif self.ha.watch(nap_time):
|
||||
self.next_run = time.time()
|
||||
|
||||
@@ -116,13 +121,16 @@ class Patroni(object):
|
||||
|
||||
def run(self):
|
||||
self.api.start()
|
||||
self.logger.start()
|
||||
self.next_run = time.time()
|
||||
|
||||
while not self.received_sigterm:
|
||||
if self._received_sighup:
|
||||
self._received_sighup = False
|
||||
if self.config.reload_local_configuration():
|
||||
self.reload_config()
|
||||
self.reload_config(True)
|
||||
else:
|
||||
self.postgresql.config.reload_config(self.config['postgresql'], True)
|
||||
|
||||
logger.info(self.ha.run_cycle())
|
||||
|
||||
@@ -152,12 +160,41 @@ class Patroni(object):
|
||||
self.api.shutdown()
|
||||
except Exception:
|
||||
logger.exception('Exception during RestApi.shutdown')
|
||||
self.ha.shutdown()
|
||||
try:
|
||||
self.ha.shutdown()
|
||||
except Exception:
|
||||
logger.exception('Exception during Ha.shutdown')
|
||||
self.logger.shutdown()
|
||||
|
||||
|
||||
def patroni_main():
|
||||
patroni = Patroni()
|
||||
import argparse
|
||||
|
||||
from multiprocessing import freeze_support
|
||||
from patroni.config import Config, ConfigParseError
|
||||
from patroni.validator import schema
|
||||
|
||||
freeze_support()
|
||||
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument('--version', action='version', version='%(prog)s {0}'.format(__version__))
|
||||
parser.add_argument('--validate-config', action='store_true', help='Run config validator and exit')
|
||||
parser.add_argument('configfile', nargs='?', default='',
|
||||
help='Patroni may also read the configuration from the {0} environment variable'
|
||||
.format(Config.PATRONI_CONFIG_VARIABLE))
|
||||
args = parser.parse_args()
|
||||
try:
|
||||
if args.validate_config:
|
||||
conf = Config(args.configfile, validator=schema)
|
||||
sys.exit()
|
||||
else:
|
||||
conf = Config(args.configfile)
|
||||
except ConfigParseError as e:
|
||||
if e.value:
|
||||
print(e.value)
|
||||
parser.print_help()
|
||||
sys.exit(1)
|
||||
patroni = Patroni(conf)
|
||||
try:
|
||||
patroni.run()
|
||||
except KeyboardInterrupt:
|
||||
@@ -193,8 +230,8 @@ def check_psycopg2():
|
||||
|
||||
|
||||
def main():
|
||||
check_psycopg2()
|
||||
if os.getpid() != 1:
|
||||
check_psycopg2()
|
||||
return patroni_main()
|
||||
|
||||
# Patroni started with PID=1, it looks like we are in the container
|
||||
|
||||
+83
-70
@@ -7,12 +7,13 @@ import traceback
|
||||
import dateutil.parser
|
||||
import datetime
|
||||
import os
|
||||
import six
|
||||
import socket
|
||||
|
||||
from patroni.postgresql import PostgresConnectionException
|
||||
from patroni.postgresql.misc import postgres_version_to_int, PostgresException
|
||||
from patroni.exceptions import PostgresConnectionException, PostgresException
|
||||
from patroni.postgresql.misc import postgres_version_to_int
|
||||
from patroni.utils import deep_compare, parse_bool, patch_config, Retry, \
|
||||
RetryFailedError, parse_int, split_host_port, tzutc, uri
|
||||
RetryFailedError, parse_int, split_host_port, tzutc, uri, cluster_as_json
|
||||
from six.moves.BaseHTTPServer import BaseHTTPRequestHandler, HTTPServer
|
||||
from six.moves.socketserver import ThreadingMixIn
|
||||
from threading import Thread
|
||||
@@ -20,20 +21,6 @@ from threading import Thread
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def check_auth(func):
|
||||
"""Decorator function to check authorization header.
|
||||
|
||||
Usage example:
|
||||
@check_auth
|
||||
def do_PUT_foo():
|
||||
pass
|
||||
"""
|
||||
def wrapper(handler, *args, **kwargs):
|
||||
if handler.check_auth_header():
|
||||
return func(handler, *args, **kwargs)
|
||||
return wrapper
|
||||
|
||||
|
||||
class RestApiHandler(BaseHTTPRequestHandler):
|
||||
|
||||
def _write_response(self, status_code, body, content_type='text/html', headers=None):
|
||||
@@ -49,14 +36,20 @@ class RestApiHandler(BaseHTTPRequestHandler):
|
||||
def _write_json_response(self, status_code, response):
|
||||
self._write_response(status_code, json.dumps(response), content_type='application/json')
|
||||
|
||||
def send_auth_request(self, body):
|
||||
headers = {'WWW-Authenticate': 'Basic realm="' + self.server.patroni.__class__.__name__ + '"'}
|
||||
self._write_response(401, body, headers=headers)
|
||||
def check_auth(func):
|
||||
"""Decorator function to check authorization header or client certificates
|
||||
|
||||
def check_auth_header(self):
|
||||
auth_header = self.headers.get('Authorization')
|
||||
status = self.server.check_auth_header(auth_header)
|
||||
return not status or self.send_auth_request(status)
|
||||
Usage example:
|
||||
@check_auth
|
||||
def do_PUT_foo():
|
||||
pass
|
||||
"""
|
||||
|
||||
def wrapper(self, *args, **kwargs):
|
||||
if self.server.check_auth(self):
|
||||
return func(self, *args, **kwargs)
|
||||
|
||||
return wrapper
|
||||
|
||||
def _write_status_response(self, status_code, response):
|
||||
patroni = self.server.patroni
|
||||
@@ -140,6 +133,14 @@ class RestApiHandler(BaseHTTPRequestHandler):
|
||||
response = self.get_postgresql_status(True)
|
||||
self._write_status_response(200, response)
|
||||
|
||||
def do_GET_cluster(self):
|
||||
cluster = self.server.patroni.dcs.cluster or self.server.patroni.dcs.get_cluster()
|
||||
self._write_json_response(200, cluster_as_json(cluster))
|
||||
|
||||
def do_GET_history(self):
|
||||
cluster = self.server.patroni.dcs.cluster or self.server.patroni.dcs.get_cluster()
|
||||
self._write_json_response(200, cluster.history and cluster.history.lines or [])
|
||||
|
||||
def do_GET_config(self):
|
||||
cluster = self.server.patroni.dcs.cluster or self.server.patroni.dcs.get_cluster()
|
||||
if cluster.config:
|
||||
@@ -187,18 +188,8 @@ class RestApiHandler(BaseHTTPRequestHandler):
|
||||
|
||||
@check_auth
|
||||
def do_POST_reload(self):
|
||||
try:
|
||||
if self.server.patroni.config.reload_local_configuration(True):
|
||||
status_code = 202
|
||||
response = 'reload scheduled'
|
||||
self.server.patroni.sighup_handler()
|
||||
else:
|
||||
status_code = 200
|
||||
response = 'nothing changed'
|
||||
except Exception as e:
|
||||
status_code = 500
|
||||
response = str(e)
|
||||
self._write_response(status_code, response)
|
||||
self.server.patroni.sighup_handler()
|
||||
self._write_response(202, 'reload scheduled')
|
||||
|
||||
@staticmethod
|
||||
def parse_schedule(schedule, action):
|
||||
@@ -432,10 +423,7 @@ class RestApiHandler(BaseHTTPRequestHandler):
|
||||
|
||||
if self.server.patroni.postgresql.state not in ('running', 'restarting', 'starting'):
|
||||
raise RetryFailedError('')
|
||||
stmt = ("WITH replication_info AS ("
|
||||
"SELECT usename, application_name, client_addr, state, sync_state, sync_priority"
|
||||
" FROM pg_catalog.pg_stat_replication) SELECT"
|
||||
" pg_catalog.to_char(pg_catalog.pg_postmaster_start_time(), 'YYYY-MM-DD HH24:MI:SS.MS TZ'),"
|
||||
stmt = ("SELECT pg_catalog.to_char(pg_catalog.pg_postmaster_start_time(), 'YYYY-MM-DD HH24:MI:SS.MS TZ'),"
|
||||
" CASE WHEN pg_catalog.pg_is_in_recovery() THEN 0"
|
||||
" ELSE ('x' || pg_catalog.substr(pg_catalog.pg_{0}file_name("
|
||||
"pg_catalog.pg_current_{0}_{1}()), 1, 8))::bit(32)::int END,"
|
||||
@@ -446,8 +434,10 @@ class RestApiHandler(BaseHTTPRequestHandler):
|
||||
" pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_last_{0}_replay_{1}(), '0/0')::bigint,"
|
||||
" pg_catalog.to_char(pg_catalog.pg_last_xact_replay_timestamp(), 'YYYY-MM-DD HH24:MI:SS.MS TZ'),"
|
||||
" pg_catalog.pg_is_in_recovery() AND pg_catalog.pg_is_{0}_replay_paused(), "
|
||||
"(SELECT pg_catalog.array_to_json(pg_catalog.array_agg("
|
||||
"pg_catalog.row_to_json(ri))) FROM replication_info ri)")
|
||||
" pg_catalog.array_to_json(pg_catalog.array_agg(pg_catalog.row_to_json(ri))) "
|
||||
"FROM (SELECT (SELECT rolname FROM pg_authid WHERE oid = usesysid) AS usename,"
|
||||
" application_name, client_addr, w.state, sync_state, sync_priority"
|
||||
" FROM pg_catalog.pg_stat_get_wal_senders() w, pg_catalog.pg_stat_get_activity(pid)) AS ri")
|
||||
|
||||
row = self.query(stmt.format(self.server.patroni.postgresql.wal_name,
|
||||
self.server.patroni.postgresql.lsn_name), retry=retry)[0]
|
||||
@@ -492,12 +482,14 @@ class RestApiHandler(BaseHTTPRequestHandler):
|
||||
|
||||
|
||||
class RestApiServer(ThreadingMixIn, HTTPServer, Thread):
|
||||
# On 3.7+ the `ThreadingMixIn` gathers all non-daemon worker threads in order to join on them at server close.
|
||||
daemon_threads = True # Make worker threads "fire and forget" to prevent a memory leak.
|
||||
|
||||
def __init__(self, patroni, config):
|
||||
self.patroni = patroni
|
||||
self.__listen = None
|
||||
self.__initialize(config)
|
||||
self.__set_config_parameters(config)
|
||||
self.__ssl_options = None
|
||||
self.reload_config(config)
|
||||
self.daemon = True
|
||||
|
||||
def query(self, sql, *params):
|
||||
@@ -528,13 +520,16 @@ class RestApiServer(ThreadingMixIn, HTTPServer, Thread):
|
||||
if not auth_header.startswith('Basic ') or not self.check_basic_auth_key(auth_header[6:]):
|
||||
return 'not authenticated'
|
||||
|
||||
@staticmethod
|
||||
def __get_ssl_options(config):
|
||||
return {option: config[option] for option in ['certfile', 'keyfile'] if option in config}
|
||||
def check_auth(self, rh):
|
||||
if not hasattr(rh.request, 'getpeercert') or not rh.request.getpeercert(): # valid client cert isn't present
|
||||
if self.__protocol == 'https' and self.__ssl_options.get('verify_client') in ('required', 'optional'):
|
||||
return rh._write_response(403, 'client certificate required')
|
||||
|
||||
def __set_config_parameters(self, config):
|
||||
self.__auth_key = base64.b64encode(config['auth'].encode('utf-8')).decode('utf-8') if 'auth' in config else None
|
||||
self.connection_string = uri(self.__protocol, config.get('connect_address') or self.__listen, 'patroni')
|
||||
reason = self.check_auth_header(rh.headers.get('Authorization'))
|
||||
if reason:
|
||||
headers = {'WWW-Authenticate': 'Basic realm="' + self.patroni.__class__.__name__ + '"'}
|
||||
return rh._write_response(401, reason, headers=headers)
|
||||
return True
|
||||
|
||||
@staticmethod
|
||||
def __has_dual_stack():
|
||||
@@ -553,7 +548,7 @@ class RestApiServer(ThreadingMixIn, HTTPServer, Thread):
|
||||
|
||||
def __httpserver_init(self, host, port):
|
||||
dual_stack = self.__has_dual_stack()
|
||||
if host == '':
|
||||
if host in ('', '*'):
|
||||
host = None
|
||||
|
||||
info = socket.getaddrinfo(host, port, socket.AF_UNSPEC, socket.SOCK_STREAM, 0, socket.AI_PASSIVE)
|
||||
@@ -561,45 +556,63 @@ class RestApiServer(ThreadingMixIn, HTTPServer, Thread):
|
||||
info.sort(key=lambda x: x[0] == socket.AF_INET, reverse=not dual_stack)
|
||||
|
||||
self.address_family = info[0][0]
|
||||
HTTPServer.__init__(self, info[0][-1][:2], RestApiHandler)
|
||||
|
||||
def __initialize(self, config):
|
||||
try:
|
||||
host, port = split_host_port(config['listen'], None)
|
||||
HTTPServer.__init__(self, info[0][-1][:2], RestApiHandler)
|
||||
except socket.error:
|
||||
logger.error(
|
||||
"Couldn't start a service on '%s:%s', please check your `restapi.listen` configuration", host, port)
|
||||
raise
|
||||
|
||||
def __initialize(self, listen, ssl_options):
|
||||
try:
|
||||
host, port = split_host_port(listen, None)
|
||||
except Exception:
|
||||
raise ValueError('Invalid "restapi" config: expected <HOST>:<PORT> for "listen", but got "{0}"'
|
||||
.format(config['listen']))
|
||||
.format(listen))
|
||||
|
||||
if self.__listen is not None: # changing config in runtime
|
||||
reloading_config = self.__listen is not None # changing config in runtime
|
||||
if reloading_config:
|
||||
self.shutdown()
|
||||
|
||||
self.__listen = config['listen']
|
||||
self.__ssl_options = self.__get_ssl_options(config)
|
||||
self.__listen = listen
|
||||
self.__ssl_options = ssl_options
|
||||
|
||||
self.__httpserver_init(host, port)
|
||||
Thread.__init__(self, target=self.serve_forever)
|
||||
self._set_fd_cloexec(self.socket)
|
||||
|
||||
self.__protocol = 'http'
|
||||
|
||||
# wrap socket with ssl if 'certfile' is defined in a config.yaml
|
||||
# Sometime it's also needed to pass reference to a 'keyfile'.
|
||||
if self.__ssl_options.get('certfile'):
|
||||
self.__protocol = 'https' if ssl_options.get('certfile') else 'http'
|
||||
if self.__protocol == 'https':
|
||||
import ssl
|
||||
ctx = ssl.create_default_context(ssl.Purpose.CLIENT_AUTH)
|
||||
ctx.load_cert_chain(**self.__ssl_options)
|
||||
ctx = ssl.create_default_context(ssl.Purpose.CLIENT_AUTH, cafile=ssl_options.get('cafile'))
|
||||
ctx.load_cert_chain(certfile=ssl_options['certfile'], keyfile=ssl_options.get('keyfile'))
|
||||
verify_client = ssl_options.get('verify_client')
|
||||
if verify_client:
|
||||
modes = {'none': ssl.CERT_NONE, 'optional': ssl.CERT_OPTIONAL, 'required': ssl.CERT_REQUIRED}
|
||||
if verify_client in modes:
|
||||
ctx.verify_mode = modes[verify_client]
|
||||
else:
|
||||
logger.error('Bad value in the "restapi.verify_client": %s', verify_client)
|
||||
self.socket = ctx.wrap_socket(self.socket, server_side=True)
|
||||
self.__protocol = 'https'
|
||||
return True
|
||||
if reloading_config:
|
||||
self.start()
|
||||
|
||||
def reload_config(self, config):
|
||||
if 'listen' not in config: # changing config in runtime
|
||||
raise ValueError('Can not find "restapi.listen" config')
|
||||
|
||||
elif (self.__listen != config['listen'] or self.__ssl_options != self.__get_ssl_options(config)) \
|
||||
and self.__initialize(config):
|
||||
self.start()
|
||||
self.__set_config_parameters(config)
|
||||
ssl_options = {n: config[n] for n in ('certfile', 'keyfile', 'cafile') if n in config}
|
||||
|
||||
if isinstance(config.get('verify_client'), six.string_types):
|
||||
ssl_options['verify_client'] = config['verify_client'].lower()
|
||||
|
||||
if self.__listen != config['listen'] or self.__ssl_options != ssl_options:
|
||||
self.__initialize(config['listen'], ssl_options)
|
||||
|
||||
self.__auth_key = base64.b64encode(config['auth'].encode('utf-8')).decode('utf-8') if 'auth' in config else None
|
||||
self.connection_string = uri(self.__protocol, config.get('connect_address') or self.__listen, 'patroni')
|
||||
|
||||
@staticmethod
|
||||
def handle_error(request, client_address):
|
||||
|
||||
@@ -110,6 +110,12 @@ class AsyncExecutor(object):
|
||||
def run_async(self, func, args=()):
|
||||
Thread(target=self.run, args=(func, args)).start()
|
||||
|
||||
def try_run_async(self, action, func, args=()):
|
||||
prev = self.schedule(action)
|
||||
if prev is None:
|
||||
return self.run_async(func, args)
|
||||
return 'Failed to run {0}, {1} is already in progress'.format(action, prev)
|
||||
|
||||
def cancel(self):
|
||||
with self:
|
||||
with self._scheduled_action_lock:
|
||||
|
||||
+44
-33
@@ -2,19 +2,34 @@ import json
|
||||
import logging
|
||||
import os
|
||||
import shutil
|
||||
import sys
|
||||
import tempfile
|
||||
import yaml
|
||||
|
||||
from collections import defaultdict
|
||||
from copy import deepcopy
|
||||
from patroni import PATRONI_ENV_PREFIX
|
||||
from patroni.exceptions import ConfigParseError
|
||||
from patroni.dcs import ClusterConfig
|
||||
from patroni.postgresql.config import ConfigHandler
|
||||
from patroni.postgresql.config import CaseInsensitiveDict, ConfigHandler
|
||||
from patroni.utils import deep_compare, parse_bool, parse_int, patch_config
|
||||
from requests.structures import CaseInsensitiveDict
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
_AUTH_ALLOWED_PARAMETERS = (
|
||||
'username',
|
||||
'password',
|
||||
'sslmode',
|
||||
'sslcert',
|
||||
'sslkey',
|
||||
'sslrootcert',
|
||||
'sslcrl'
|
||||
)
|
||||
|
||||
|
||||
def default_validator(conf):
|
||||
if not conf:
|
||||
return "Config is empty."
|
||||
|
||||
|
||||
class Config(object):
|
||||
"""
|
||||
@@ -36,7 +51,6 @@ class Config(object):
|
||||
to work with it as with the old `config` object.
|
||||
"""
|
||||
|
||||
PATRONI_ENV_PREFIX = 'PATRONI_'
|
||||
PATRONI_CONFIG_VARIABLE = PATRONI_ENV_PREFIX + 'CONFIGURATION'
|
||||
|
||||
__CACHE_FILENAME = 'patroni.dynamic.json'
|
||||
@@ -45,6 +59,7 @@ class Config(object):
|
||||
'maximum_lag_on_failover': 1048576,
|
||||
'check_timeline': False,
|
||||
'master_start_timeout': 300,
|
||||
'master_stop_timeout': 0,
|
||||
'synchronous_mode': False,
|
||||
'synchronous_mode_strict': False,
|
||||
'standby_cluster': {
|
||||
@@ -66,27 +81,26 @@ class Config(object):
|
||||
}
|
||||
}
|
||||
|
||||
def __init__(self):
|
||||
def __init__(self, configfile, validator=default_validator):
|
||||
self._modify_index = -1
|
||||
self._dynamic_configuration = {}
|
||||
|
||||
self.__environment_configuration = self._build_environment_configuration()
|
||||
|
||||
# Patroni reads the configuration from the command-line argument if it exists, otherwise from the environment
|
||||
self._config_file = len(sys.argv) >= 2 and os.path.isfile(sys.argv[1]) and sys.argv[1]
|
||||
self._config_file = configfile and os.path.isfile(configfile) and configfile
|
||||
if self._config_file:
|
||||
self._local_configuration = self._load_config_file()
|
||||
else:
|
||||
config_env = os.environ.pop(self.PATRONI_CONFIG_VARIABLE, None)
|
||||
self._local_configuration = config_env and yaml.safe_load(config_env) or self.__environment_configuration
|
||||
if not self._local_configuration:
|
||||
print('Usage: {0} config.yml'.format(sys.argv[0]))
|
||||
print('\tPatroni may also read the configuration from the {0} environment variable'.
|
||||
format(self.PATRONI_CONFIG_VARIABLE))
|
||||
sys.exit(1)
|
||||
if validator:
|
||||
error = validator(self._local_configuration)
|
||||
if error:
|
||||
raise ConfigParseError(error)
|
||||
|
||||
self.__effective_configuration = self._build_effective_configuration({}, self._local_configuration)
|
||||
self._data_dir = self.__effective_configuration['postgresql']['data_dir']
|
||||
self._data_dir = self.__effective_configuration.get('postgresql', {}).get('data_dir', "")
|
||||
self._cache_file = os.path.join(self._data_dir, self.__CACHE_FILENAME)
|
||||
self._load_cache()
|
||||
self._cache_needs_saving = False
|
||||
@@ -158,23 +172,19 @@ class Config(object):
|
||||
except Exception:
|
||||
logger.exception('Exception when setting dynamic_configuration')
|
||||
|
||||
def reload_local_configuration(self, dry_run=False):
|
||||
def reload_local_configuration(self):
|
||||
if self.config_file:
|
||||
try:
|
||||
configuration = self._load_config_file()
|
||||
if not deep_compare(self._local_configuration, configuration):
|
||||
new_configuration = self._build_effective_configuration(self._dynamic_configuration, configuration)
|
||||
if dry_run:
|
||||
return not deep_compare(new_configuration, self.__effective_configuration)
|
||||
self._local_configuration = configuration
|
||||
self.__effective_configuration = new_configuration
|
||||
return True
|
||||
else:
|
||||
logger.info('No configuration items changed, nothing to reload.')
|
||||
logger.info('No local configuration items changed.')
|
||||
except Exception:
|
||||
logger.exception('Exception when reloading local configuration from %s', self.config_file)
|
||||
if dry_run:
|
||||
raise
|
||||
|
||||
@staticmethod
|
||||
def _process_postgresql_parameters(parameters, is_local=False):
|
||||
@@ -208,7 +218,7 @@ class Config(object):
|
||||
ret = defaultdict(dict)
|
||||
|
||||
def _popenv(name):
|
||||
return os.environ.pop(Config.PATRONI_ENV_PREFIX + name.upper(), None)
|
||||
return os.environ.pop(PATRONI_ENV_PREFIX + name.upper(), None)
|
||||
|
||||
for param in ('name', 'namespace', 'scope'):
|
||||
value = _popenv(param)
|
||||
@@ -217,7 +227,7 @@ class Config(object):
|
||||
|
||||
def _fix_log_env(name, oldname):
|
||||
value = _popenv(oldname)
|
||||
name = Config.PATRONI_ENV_PREFIX + 'LOG_' + name.upper()
|
||||
name = PATRONI_ENV_PREFIX + 'LOG_' + name.upper()
|
||||
if value and name not in os.environ:
|
||||
os.environ[name] = value
|
||||
|
||||
@@ -230,9 +240,10 @@ class Config(object):
|
||||
if value:
|
||||
ret[section][param] = value
|
||||
|
||||
_set_section_values('restapi', ['listen', 'connect_address', 'certfile', 'keyfile'])
|
||||
_set_section_values('restapi', ['listen', 'connect_address', 'certfile', 'keyfile', 'cafile', 'verify_client'])
|
||||
_set_section_values('ctl', ['insecure', 'cacert', 'certfile', 'keyfile'])
|
||||
_set_section_values('postgresql', ['listen', 'connect_address', 'config_dir', 'data_dir', 'pgpass', 'bin_dir'])
|
||||
_set_section_values('log', ['level', 'format', 'dateformat', 'max_queue_size',
|
||||
_set_section_values('log', ['level', 'traceback_level', 'format', 'dateformat', 'max_queue_size',
|
||||
'dir', 'file_size', 'file_num', 'loggers'])
|
||||
|
||||
def _parse_dict(value):
|
||||
@@ -250,9 +261,9 @@ class Config(object):
|
||||
if value:
|
||||
ret['log']['loggers'] = value
|
||||
|
||||
def _get_auth(name):
|
||||
def _get_auth(name, params=None):
|
||||
ret = {}
|
||||
for param in ('username', 'password'):
|
||||
for param in params or _AUTH_ALLOWED_PARAMETERS[:2]:
|
||||
value = _popenv(name + '_' + param)
|
||||
if value:
|
||||
ret[param] = value
|
||||
@@ -264,7 +275,7 @@ class Config(object):
|
||||
|
||||
authentication = {}
|
||||
for user_type in ('replication', 'superuser', 'rewind'):
|
||||
entry = _get_auth(user_type)
|
||||
entry = _get_auth(user_type, _AUTH_ALLOWED_PARAMETERS)
|
||||
if entry:
|
||||
authentication[user_type] = entry
|
||||
|
||||
@@ -281,7 +292,7 @@ class Config(object):
|
||||
return None
|
||||
|
||||
for param in list(os.environ.keys()):
|
||||
if param.startswith(Config.PATRONI_ENV_PREFIX):
|
||||
if param.startswith(PATRONI_ENV_PREFIX):
|
||||
# PATRONI_(ETCD|CONSUL|ZOOKEEPER|EXHIBITOR|...)_(HOSTS?|PORT|..)
|
||||
name, suffix = (param[8:].split('_', 1) + [''])[:2]
|
||||
if suffix in ('HOST', 'HOSTS', 'PORT', 'USE_PROXIES', 'PROTOCOL', 'SRV', 'URL', 'PROXY',
|
||||
@@ -304,7 +315,7 @@ class Config(object):
|
||||
|
||||
users = {}
|
||||
for param in list(os.environ.keys()):
|
||||
if param.startswith(Config.PATRONI_ENV_PREFIX):
|
||||
if param.startswith(PATRONI_ENV_PREFIX):
|
||||
name, suffix = (param[8:].rsplit('_', 1) + [''])[:2]
|
||||
# PATRONI_<username>_PASSWORD=<password>, PATRONI_<username>_OPTIONS=<option1,option2,...>
|
||||
# CREATE USER "<username>" WITH <OPTIONS> PASSWORD '<password>'
|
||||
@@ -332,14 +343,9 @@ class Config(object):
|
||||
config['postgresql'][name] = deepcopy(value)
|
||||
elif name not in config or name in ['watchdog']:
|
||||
config[name] = deepcopy(value) if value else {}
|
||||
|
||||
if os.getenv("PATRONI_DEBUG_MODE"):
|
||||
config['ttl'] = 24 * 60 * 60 # practical infinity for a debugging session
|
||||
config['loop_wait'] = -1
|
||||
|
||||
|
||||
# restapi server expects to get restapi.auth = 'username:password'
|
||||
if 'authentication' in config['restapi']:
|
||||
if 'restapi' in config and 'authentication' in config['restapi']:
|
||||
config['restapi']['auth'] = '{username}:{password}'.format(**config['restapi']['authentication'])
|
||||
|
||||
# special treatment for old config
|
||||
@@ -358,6 +364,11 @@ class Config(object):
|
||||
if 'superuser' not in pg_config['authentication'] and 'pg_rewind' in pg_config:
|
||||
pg_config['authentication']['superuser'] = pg_config['pg_rewind']
|
||||
|
||||
# handle setting additional connection parameters that may be available
|
||||
# in the configuration file, such as SSL connection parameters
|
||||
for name, value in pg_config['authentication'].items():
|
||||
pg_config['authentication'][name] = {n: v for n, v in value.items() if n in _AUTH_ALLOWED_PARAMETERS}
|
||||
|
||||
# no 'name' in config
|
||||
if 'name' not in config and 'name' in pg_config:
|
||||
config['name'] = pg_config['name']
|
||||
|
||||
+196
-191
@@ -2,11 +2,11 @@
|
||||
Patroni Control
|
||||
'''
|
||||
|
||||
import base64
|
||||
import click
|
||||
import codecs
|
||||
import datetime
|
||||
import dateutil.parser
|
||||
import dateutil.tz
|
||||
import cdiff
|
||||
import copy
|
||||
import difflib
|
||||
@@ -15,26 +15,24 @@ import json
|
||||
import logging
|
||||
import os
|
||||
import random
|
||||
import requests
|
||||
import six
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
import tzlocal
|
||||
import yaml
|
||||
|
||||
from click import ClickException
|
||||
from contextlib import contextmanager
|
||||
from patroni.config import Config
|
||||
from patroni.dcs import get_dcs as _get_dcs
|
||||
from patroni.exceptions import PatroniException
|
||||
from patroni.postgresql import Postgresql
|
||||
from patroni.postgresql.misc import postgres_version_to_int
|
||||
from patroni.utils import patch_config, polling_loop
|
||||
from patroni.utils import cluster_as_json, patch_config, polling_loop
|
||||
from patroni.request import PatroniRequest
|
||||
from patroni.version import __version__
|
||||
from prettytable import PrettyTable
|
||||
from prettytable import ALL, FRAME, PrettyTable
|
||||
from six.moves.urllib_parse import urlparse
|
||||
from six import text_type
|
||||
|
||||
CONFIG_DIR_PATH = click.get_app_dir('patroni')
|
||||
CONFIG_FILE_PATH = os.path.join(CONFIG_DIR_PATH, 'patronictl.yaml')
|
||||
@@ -48,14 +46,42 @@ class PatroniCtlException(ClickException):
|
||||
pass
|
||||
|
||||
|
||||
class PatronictlPrettyTable(PrettyTable):
|
||||
|
||||
def __init__(self, header, *args, **kwargs):
|
||||
PrettyTable.__init__(self, *args, **kwargs)
|
||||
self.__table_header = header
|
||||
self.__hline_num = 0
|
||||
self.__hline = None
|
||||
|
||||
def _is_first_hline(self):
|
||||
return self.__hline_num == 0
|
||||
|
||||
def _set_hline(self, value):
|
||||
self.__hline = value
|
||||
|
||||
def _get_hline(self):
|
||||
ret = self.__hline
|
||||
|
||||
# Inject nice table header
|
||||
if self._is_first_hline() and self.__table_header:
|
||||
header = self.__table_header[:len(ret) - 2]
|
||||
ret = "".join([ret[0], header, ret[1 + len(header):]])
|
||||
|
||||
self.__hline_num += 1
|
||||
return ret
|
||||
|
||||
_hrule = property(_get_hline, _set_hline)
|
||||
|
||||
|
||||
def parse_dcs(dcs):
|
||||
if dcs is None:
|
||||
return None
|
||||
elif '//' not in dcs:
|
||||
dcs = '//' + dcs
|
||||
|
||||
parsed = urlparse(dcs)
|
||||
scheme = parsed.scheme
|
||||
if scheme == '' and parsed.netloc == '':
|
||||
parsed = urlparse('//' + dcs)
|
||||
port = int(parsed.port) if parsed.port else None
|
||||
|
||||
if scheme == '':
|
||||
@@ -68,21 +94,17 @@ def parse_dcs(dcs):
|
||||
|
||||
|
||||
def load_config(path, dcs):
|
||||
from patroni.config import Config
|
||||
|
||||
if not (os.path.exists(path) and os.access(path, os.R_OK)):
|
||||
logging.debug('Ignoring configuration file "%s". It does not exists or is not readable.', path)
|
||||
if path != CONFIG_FILE_PATH: # bail if non-default config location specified but file not found / readable
|
||||
raise PatroniCtlException('Provided config file {0} not existing or no read rights.'
|
||||
' Check the -c/--config-file parameter'.format(path))
|
||||
else:
|
||||
logging.debug('Ignoring configuration file "%s". It does not exists or is not readable.', path)
|
||||
else:
|
||||
logging.debug('Loading configuration from file %s', path)
|
||||
config = {}
|
||||
old_argv = list(sys.argv)
|
||||
try:
|
||||
sys.argv[1] = path
|
||||
if Config.PATRONI_CONFIG_VARIABLE not in os.environ:
|
||||
for p in ('PATRONI_RESTAPI_LISTEN', 'PATRONI_POSTGRESQL_DATA_DIR'):
|
||||
if p not in os.environ:
|
||||
os.environ[p] = '.'
|
||||
config = Config().copy()
|
||||
finally:
|
||||
sys.argv = old_argv
|
||||
config = Config(path, validator=None).copy()
|
||||
|
||||
dcs = parse_dcs(dcs) or parse_dcs(config.get('dcs_api')) or {}
|
||||
if dcs:
|
||||
@@ -100,7 +122,7 @@ def store_config(config, path):
|
||||
yaml.dump(config, fd)
|
||||
|
||||
|
||||
option_format = click.option('--format', '-f', 'fmt', help='Output format (pretty, json, yaml)', default='pretty')
|
||||
option_format = click.option('--format', '-f', 'fmt', help='Output format (pretty, tsv, json, yaml)', default='pretty')
|
||||
option_watchrefresh = click.option('-w', '--watch', type=float, help='Auto update the screen every X seconds')
|
||||
option_watch = click.option('-W', is_flag=True, help='Auto update the screen every 2 seconds')
|
||||
option_force = click.option('--force', is_flag=True, help='Do not ask for confirmation at any point')
|
||||
@@ -135,63 +157,41 @@ def get_dcs(config, scope):
|
||||
raise PatroniCtlException(str(e))
|
||||
|
||||
|
||||
def auth_header(config):
|
||||
if config.get('restapi', {}).get('auth', ''):
|
||||
return {'Authorization': 'Basic ' + base64.b64encode(config['restapi']['auth'].encode('utf-8')).decode('utf-8')}
|
||||
|
||||
|
||||
def request_patroni(member, request_type, endpoint, content=None, headers=None):
|
||||
def request_patroni(member, method='GET', endpoint=None, data=None):
|
||||
ctx = click.get_current_context() # the current click context
|
||||
headers = headers or {}
|
||||
url_parts = urlparse(member.api_url)
|
||||
logging.debug(url_parts)
|
||||
if 'Content-Type' not in headers:
|
||||
headers['Content-Type'] = 'application/json'
|
||||
|
||||
url = '{0}://{1}/{2}'.format(url_parts.scheme, url_parts.netloc, endpoint)
|
||||
|
||||
insecure = ctx.obj.get('ctl', {}).get('insecure', False)
|
||||
# Get certfile if any from several configuration namespace
|
||||
cert = ctx.obj.get('ctl', {}).get('cacert') or \
|
||||
ctx.obj.get('restapi', {}).get('cacert') or \
|
||||
ctx.obj.get('restapi', {}).get('certfile')
|
||||
# In the case we specificaly disable SSL cert verification we don't want to have the warning
|
||||
if insecure:
|
||||
verify = False
|
||||
elif cert:
|
||||
verify = cert
|
||||
else:
|
||||
verify = True
|
||||
return getattr(requests, request_type)(url, headers=headers,
|
||||
data=json.dumps(content) if content else None, timeout=60,
|
||||
verify=verify)
|
||||
request_executor = ctx.obj.get('__request_patroni')
|
||||
if not request_executor:
|
||||
request_executor = ctx.obj['__request_patroni'] = PatroniRequest(ctx.obj)
|
||||
return request_executor(member, method, endpoint, data)
|
||||
|
||||
|
||||
def print_output(columns, rows=None, alignment=None, fmt='pretty', header=True, delimiter='\t'):
|
||||
rows = rows or []
|
||||
if fmt == 'pretty':
|
||||
t = PrettyTable(columns)
|
||||
for k, v in (alignment or {}).items():
|
||||
t.align[k] = v
|
||||
for r in rows:
|
||||
t.add_row(r)
|
||||
click.echo(t)
|
||||
return
|
||||
def print_output(columns, rows, alignment=None, fmt='pretty', header=None, delimiter='\t'):
|
||||
if fmt in {'json', 'yaml', 'yml'}:
|
||||
elements = [{k: v for k, v in zip(columns, r) if not header or str(v)} for r in rows]
|
||||
func = json.dumps if fmt == 'json' else format_config_for_editing
|
||||
click.echo(func(elements))
|
||||
elif fmt in {'pretty', 'tsv'}:
|
||||
list_cluster = bool(header and columns and columns[0] == 'Cluster')
|
||||
if list_cluster and 'Tags' in columns: # we want to format member tags as YAML
|
||||
i = columns.index('Tags')
|
||||
for row in rows:
|
||||
if row[i]:
|
||||
row[i] = format_config_for_editing(row[i], fmt == 'tsv').strip()
|
||||
if list_cluster and fmt == 'pretty': # skip cluster name if pretty-printing
|
||||
columns = columns[1:] if columns else []
|
||||
rows = [row[1:] for row in rows]
|
||||
|
||||
if fmt in ['json', 'yaml', 'yml']:
|
||||
elements = [dict(zip(columns, r)) for r in rows]
|
||||
if fmt == 'json':
|
||||
click.echo(json.dumps(elements))
|
||||
elif fmt in ('yaml', 'yml'):
|
||||
click.echo(yaml.safe_dump(elements, encoding=None, default_flow_style=False, allow_unicode=True, width=200))
|
||||
|
||||
if fmt == 'tsv':
|
||||
if columns is not None and header:
|
||||
click.echo(delimiter.join(columns))
|
||||
|
||||
for r in rows:
|
||||
c = [str(c) for c in r]
|
||||
click.echo(delimiter.join(c))
|
||||
if fmt == 'tsv':
|
||||
for r in ([columns] if columns else []) + rows:
|
||||
click.echo(delimiter.join(map(str, r)))
|
||||
else:
|
||||
hrules = ALL if any(any(isinstance(c, six.string_types) and '\n' in c for c in r) for r in rows) else FRAME
|
||||
table = PatronictlPrettyTable(header, columns, hrules=hrules)
|
||||
for k, v in (alignment or {}).items():
|
||||
table.align[k] = v
|
||||
for r in rows:
|
||||
table.add_row(r)
|
||||
click.echo(table)
|
||||
|
||||
|
||||
def watching(w, watch, max_count=None, clear=True):
|
||||
@@ -271,10 +271,12 @@ def get_cursor(cluster, connect_parameters, role='master', member=None):
|
||||
return None
|
||||
|
||||
|
||||
def get_members(cluster, cluster_name, member_names, role, force, action, scheduled_at=None):
|
||||
def get_members(cluster, cluster_name, member_names, role, force, action, ask_confirmation=True):
|
||||
candidates = {m.name: m for m in cluster.members}
|
||||
|
||||
if not force or role:
|
||||
if not member_names and not candidates:
|
||||
raise PatroniCtlException('{0} cluster doesn\'t have any members'.format(cluster_name))
|
||||
output_members(cluster, cluster_name)
|
||||
|
||||
if role:
|
||||
@@ -294,21 +296,26 @@ def get_members(cluster, cluster_name, member_names, role, force, action, schedu
|
||||
if member_name not in candidates:
|
||||
raise PatroniCtlException('{0} is not a member of cluster'.format(member_name))
|
||||
|
||||
members = [candidates[n] for n in member_names]
|
||||
if ask_confirmation:
|
||||
confirm_members_action(members, force, action)
|
||||
return members
|
||||
|
||||
|
||||
def confirm_members_action(members, force, action, scheduled_at=None):
|
||||
if scheduled_at:
|
||||
if not force:
|
||||
confirm = click.confirm('Are you sure you want to schedule {0} of members {1} at {2}?'
|
||||
.format(action, ', '.join(member_names), scheduled_at))
|
||||
.format(action, ', '.join([m.name for m in members]), scheduled_at))
|
||||
if not confirm:
|
||||
raise PatroniCtlException('Aborted scheduled {0}'.format(action))
|
||||
else:
|
||||
if not force:
|
||||
confirm = click.confirm('Are you sure you want to {0} members {1}?'
|
||||
.format(action, ', '.join(member_names)))
|
||||
.format(action, ', '.join([m.name for m in members])))
|
||||
if not confirm:
|
||||
raise PatroniCtlException('Aborted {0}'.format(action))
|
||||
|
||||
return [candidates[n] for n in member_names]
|
||||
|
||||
|
||||
@ctl.command('dsn', help='Generate a dsn for the provided member, defaults to a dsn of the master')
|
||||
@click.option('--role', '-r', help='Give a dsn of any member with this role', type=click.Choice(['master', 'replica',
|
||||
@@ -333,8 +340,7 @@ def dsn(obj, cluster_name, role, member):
|
||||
|
||||
@ctl.command('query', help='Query a Patroni PostgreSQL member')
|
||||
@arg_cluster_name
|
||||
@option_format
|
||||
@click.option('--format', 'fmt', help='Output format (pretty, json)', default='tsv')
|
||||
@click.option('--format', 'fmt', help='Output format (pretty, tsv, json, yaml)', default='tsv')
|
||||
@click.option('--file', '-f', 'p_file', help='Execute the SQL commands from this file', type=click.File('rb'))
|
||||
@click.option('--password', help='force password prompt', is_flag=True)
|
||||
@click.option('-U', '--username', help='database user name', type=str)
|
||||
@@ -391,8 +397,8 @@ def query(
|
||||
if cursor is None:
|
||||
cluster = dcs.get_cluster()
|
||||
|
||||
output, cursor = query_member(cluster, cursor, member, role, command, connect_parameters)
|
||||
print_output(None, output, fmt=fmt, delimiter=delimiter)
|
||||
output, header = query_member(cluster, cursor, member, role, command, connect_parameters)
|
||||
print_output(header, output, fmt=fmt, delimiter=delimiter)
|
||||
|
||||
|
||||
def query_member(cluster, cursor, member, role, command, connect_parameters):
|
||||
@@ -409,15 +415,8 @@ def query_member(cluster, cursor, member, role, command, connect_parameters):
|
||||
logging.debug(message)
|
||||
return [[timestamp(0), message]], None
|
||||
|
||||
cursor.execute('SELECT pg_catalog.pg_is_in_recovery()')
|
||||
in_recovery = cursor.fetchone()[0]
|
||||
|
||||
if in_recovery and role == 'master' or not in_recovery and role == 'replica':
|
||||
cursor.connection.close()
|
||||
return None, None
|
||||
|
||||
cursor.execute(command)
|
||||
return cursor.fetchall(), cursor
|
||||
return cursor.fetchall(), [d.name for d in cursor.description]
|
||||
except (psycopg2.OperationalError, psycopg2.DatabaseError) as oe:
|
||||
logging.debug(oe)
|
||||
if cursor is not None and not cursor.connection.closed:
|
||||
@@ -457,9 +456,9 @@ def remove(obj, cluster_name, fmt):
|
||||
|
||||
|
||||
def check_response(response, member_name, action_name, silent_success=False):
|
||||
if response.status_code >= 400:
|
||||
if response.status >= 400:
|
||||
click.echo('Failed: {0} for member {1}, status code={2}, ({3})'.format(
|
||||
action_name, member_name, response.status_code, response.text
|
||||
action_name, member_name, response.status, response.data.decode('utf-8')
|
||||
))
|
||||
return False
|
||||
elif not silent_success:
|
||||
@@ -472,7 +471,7 @@ def parse_scheduled(scheduled):
|
||||
try:
|
||||
scheduled_at = dateutil.parser.parse(scheduled)
|
||||
if scheduled_at.tzinfo is None:
|
||||
scheduled_at = tzlocal.get_localzone().localize(scheduled_at)
|
||||
scheduled_at = scheduled_at.replace(tzinfo=dateutil.tz.tzlocal())
|
||||
except (ValueError, TypeError):
|
||||
message = 'Unable to parse scheduled timestamp ({0}). It should be in an unambiguous format (e.g. ISO 8601)'
|
||||
raise PatroniCtlException(message.format(scheduled))
|
||||
@@ -493,18 +492,17 @@ def reload(obj, cluster_name, member_names, force, role):
|
||||
|
||||
members = get_members(cluster, cluster_name, member_names, role, force, 'reload')
|
||||
|
||||
content = {}
|
||||
for member in members:
|
||||
r = request_patroni(member, 'post', 'reload', content, auth_header(obj))
|
||||
if r.status_code == 200:
|
||||
r = request_patroni(member, 'post', 'reload')
|
||||
if r.status == 200:
|
||||
click.echo('No changes to apply on member {0}'.format(member.name))
|
||||
elif r.status_code == 202:
|
||||
elif r.status == 202:
|
||||
click.echo('Reload request received for member {0} and will be processed within {1} seconds'.format(
|
||||
member.name, cluster.config.data.get('loop_wait'))
|
||||
)
|
||||
else:
|
||||
click.echo('Failed: reload for member {0}, status code={1}, ({2})'.format(
|
||||
member.name, r.status_code, r.text)
|
||||
member.name, r.status, r.data.decode('utf-8'))
|
||||
)
|
||||
|
||||
|
||||
@@ -526,14 +524,15 @@ def reload(obj, cluster_name, member_names, force, role):
|
||||
def restart(obj, cluster_name, member_names, force, role, p_any, scheduled, version, pending, timeout):
|
||||
cluster = get_dcs(obj, cluster_name).get_cluster()
|
||||
|
||||
members = get_members(cluster, cluster_name, member_names, role, force, 'restart', False)
|
||||
if scheduled is None and not force:
|
||||
next_hour = (datetime.datetime.now() + datetime.timedelta(hours=1)).strftime('%Y-%m-%dT%H:%M')
|
||||
scheduled = click.prompt('When should the restart take place (e.g. ' + next_hour + ') ',
|
||||
type=str, default='now')
|
||||
|
||||
scheduled_at = parse_scheduled(scheduled)
|
||||
confirm_members_action(members, force, 'restart', scheduled_at)
|
||||
|
||||
members = get_members(cluster, cluster_name, member_names, role, force, 'restart', scheduled_at)
|
||||
if p_any:
|
||||
random.shuffle(members)
|
||||
members = members[:1]
|
||||
@@ -565,19 +564,19 @@ def restart(obj, cluster_name, member_names, force, role, p_any, scheduled, vers
|
||||
for member in members:
|
||||
if 'schedule' in content:
|
||||
if force and member.data.get('scheduled_restart'):
|
||||
r = request_patroni(member, 'delete', 'restart', headers=auth_header(obj))
|
||||
r = request_patroni(member, 'delete', 'restart')
|
||||
check_response(r, member.name, 'flush scheduled restart', True)
|
||||
|
||||
r = request_patroni(member, 'post', 'restart', content, auth_header(obj))
|
||||
if r.status_code == 200:
|
||||
r = request_patroni(member, 'post', 'restart', content)
|
||||
if r.status == 200:
|
||||
click.echo('Success: restart on member {0}'.format(member.name))
|
||||
elif r.status_code == 202:
|
||||
elif r.status == 202:
|
||||
click.echo('Success: restart scheduled on member {0}'.format(member.name))
|
||||
elif r.status_code == 409:
|
||||
elif r.status == 409:
|
||||
click.echo('Failed: another restart is already scheduled on member {0}'.format(member.name))
|
||||
else:
|
||||
click.echo('Failed: restart for member {0}, status code={1}, ({2})'.format(
|
||||
member.name, r.status_code, r.text)
|
||||
member.name, r.status, r.data.decode('utf-8'))
|
||||
)
|
||||
|
||||
|
||||
@@ -585,20 +584,39 @@ def restart(obj, cluster_name, member_names, force, role, p_any, scheduled, vers
|
||||
@click.argument('cluster_name')
|
||||
@click.argument('member_names', nargs=-1)
|
||||
@option_force
|
||||
@click.option('--wait', help='Wait until reinitialization completes', is_flag=True)
|
||||
@click.pass_obj
|
||||
def reinit(obj, cluster_name, member_names, force):
|
||||
def reinit(obj, cluster_name, member_names, force, wait):
|
||||
cluster = get_dcs(obj, cluster_name).get_cluster()
|
||||
members = get_members(cluster, cluster_name, member_names, None, force, 'reinitialize')
|
||||
|
||||
wait_on_members = []
|
||||
for member in members:
|
||||
body = {'force': force}
|
||||
while True:
|
||||
r = request_patroni(member, 'post', 'reinitialize', body, auth_header(obj))
|
||||
if not check_response(r, member.name, 'reinitialize') and r.text.endswith(' already in progress') \
|
||||
r = request_patroni(member, 'post', 'reinitialize', body)
|
||||
started = check_response(r, member.name, 'reinitialize')
|
||||
if not started and r.data.endswith(b' already in progress') \
|
||||
and not force and click.confirm('Do you want to cancel it and reinitialize anyway?'):
|
||||
body['force'] = True
|
||||
continue
|
||||
break
|
||||
if started and wait:
|
||||
wait_on_members.append(member)
|
||||
|
||||
last_display = []
|
||||
while wait_on_members:
|
||||
if wait_on_members != last_display:
|
||||
click.echo('Waiting for reinitialize to complete on: {0}'.format(
|
||||
", ".join(member.name for member in wait_on_members))
|
||||
)
|
||||
last_display[:] = wait_on_members
|
||||
time.sleep(2)
|
||||
for member in wait_on_members:
|
||||
data = json.loads(request_patroni(member, 'get', 'patroni').data.decode('utf-8'))
|
||||
if data.get('state') != 'creating replica':
|
||||
click.echo('Reinitialize is completed on: {0}'.format(member.name))
|
||||
wait_on_members.remove(member)
|
||||
|
||||
|
||||
def _do_failover_or_switchover(obj, action, cluster_name, master, candidate, force, scheduled=None):
|
||||
@@ -682,19 +700,19 @@ def _do_failover_or_switchover(obj, action, cluster_name, master, candidate, for
|
||||
try:
|
||||
member = cluster.leader.member if cluster.leader else cluster.get_member(candidate, False)
|
||||
|
||||
r = request_patroni(member, 'post', action, failover_value, auth_header(obj))
|
||||
r = request_patroni(member, 'post', action, failover_value)
|
||||
|
||||
# probably old patroni, which doesn't support switchover yet
|
||||
if r.status_code == 501 and action == 'switchover' and 'Server does not support this operation' in r.text:
|
||||
r = request_patroni(member, 'post', 'failover', failover_value, auth_header(obj))
|
||||
if r.status == 501 and action == 'switchover' and b'Server does not support this operation' in r.data:
|
||||
r = request_patroni(member, 'post', 'failover', failover_value)
|
||||
|
||||
if r.status_code in (200, 202):
|
||||
if r.status in (200, 202):
|
||||
logging.debug(r)
|
||||
cluster = dcs.get_cluster()
|
||||
logging.debug(cluster)
|
||||
click.echo('{0} {1}'.format(timestamp(), r.text))
|
||||
click.echo('{0} {1}'.format(timestamp(), r.data.decode('utf-8')))
|
||||
else:
|
||||
click.echo('{0} failed, details: {1}, {2}'.format(action.title(), r.status_code, r.text))
|
||||
click.echo('{0} failed, details: {1}, {2}'.format(action.title(), r.status, r.data.decode('utf-8')))
|
||||
return
|
||||
except Exception:
|
||||
logging.exception(r)
|
||||
@@ -731,80 +749,54 @@ def switchover(obj, cluster_name, master, candidate, force, scheduled):
|
||||
def output_members(cluster, name, extended=False, fmt='pretty'):
|
||||
rows = []
|
||||
logging.debug(cluster)
|
||||
leader_name = None
|
||||
if cluster.leader:
|
||||
leader_name = cluster.leader.name
|
||||
|
||||
xlog_location_cluster = cluster.last_leader_operation or 0
|
||||
|
||||
# Mainly for consistent pretty printing and watching we sort the output
|
||||
cluster.members.sort(key=lambda x: x.name)
|
||||
|
||||
has_scheduled_restarts = any(m.data.get('scheduled_restart') for m in cluster.members)
|
||||
has_pending_restarts = any(m.data.get('pending_restart') for m in cluster.members)
|
||||
|
||||
# Show Host as 'host:port' if somebody is running on non-standard port or two nodes are running on the same host
|
||||
append_port = any(str(m.conn_kwargs()['port']) != '5432' for m in cluster.members) or\
|
||||
len(set(m.conn_kwargs()['host'] for m in cluster.members)) < len(cluster.members)
|
||||
|
||||
for m in cluster.members:
|
||||
logging.debug(m)
|
||||
|
||||
role = ''
|
||||
if m.name == leader_name:
|
||||
role = 'Leader'
|
||||
elif m.name == cluster.sync.sync_standby:
|
||||
role = 'Sync standby'
|
||||
|
||||
xlog_location = m.data.get('xlog_location')
|
||||
lag = ''
|
||||
if xlog_location is None:
|
||||
lag = 'unknown'
|
||||
elif xlog_location_cluster >= xlog_location:
|
||||
lag = round((xlog_location_cluster - xlog_location)/1024/1024)
|
||||
|
||||
host = m.conn_kwargs()['host']
|
||||
if append_port:
|
||||
host += ':{0}'.format(m.conn_kwargs()['port'])
|
||||
|
||||
row = [name, m.name, host, role, m.data.get('state', ''), m.data.get('timeline', ''), lag]
|
||||
|
||||
if extended or has_pending_restarts:
|
||||
row.append('*' if m.data.get('pending_restart') else '')
|
||||
|
||||
if extended or has_scheduled_restarts:
|
||||
value = ''
|
||||
scheduled_restart = m.data.get('scheduled_restart')
|
||||
if scheduled_restart:
|
||||
value = scheduled_restart['schedule']
|
||||
if 'postgres_version' in scheduled_restart:
|
||||
value += ' if version < {0}'.format(scheduled_restart['postgres_version'])
|
||||
|
||||
row.append(value)
|
||||
|
||||
rows.append(row)
|
||||
initialize = {None: 'uninitialized', '': 'initializing'}.get(cluster.initialize, cluster.initialize)
|
||||
cluster = cluster_as_json(cluster)
|
||||
|
||||
columns = ['Cluster', 'Member', 'Host', 'Role', 'State', 'TL', 'Lag in MB']
|
||||
alignment = {'Lag in MB': 'r', 'TL': 'r'}
|
||||
for c in ('Pending restart', 'Scheduled restart', 'Tags'):
|
||||
if extended or any(m.get(c.lower().replace(' ', '_')) for m in cluster['members']):
|
||||
columns.append(c)
|
||||
|
||||
if extended or has_pending_restarts:
|
||||
columns.append('Pending restart')
|
||||
# Show Host as 'host:port' if somebody is running on non-standard port or two nodes are running on the same host
|
||||
members = [m for m in cluster['members'] if 'host' in m]
|
||||
append_port = any('port' in m and m['port'] != 5432 for m in members) or\
|
||||
len(set(m['host'] for m in cluster['members'])) < len(members)
|
||||
|
||||
if extended or has_scheduled_restarts:
|
||||
columns.append('Scheduled restart')
|
||||
for m in cluster['members']:
|
||||
logging.debug(m)
|
||||
|
||||
print_output(columns, rows, alignment, fmt)
|
||||
lag = m.get('lag', '')
|
||||
m.update(cluster=name, member=m['name'], host=m.get('host'), tl=m.get('timeline', ''),
|
||||
role='' if m['role'] == 'replica' else m['role'].replace('_', ' ').title(),
|
||||
lag_in_mb=round(lag/1024/1024) if isinstance(lag, six.integer_types) else lag,
|
||||
pending_restart='*' if m.get('pending_restart') else '')
|
||||
|
||||
if append_port and m['host'] and m.get('port'):
|
||||
m['host'] = ':'.join([m['host'], str(m['port'])])
|
||||
|
||||
if 'scheduled_restart' in m:
|
||||
value = m['scheduled_restart']['schedule']
|
||||
if 'postgres_version' in m['scheduled_restart']:
|
||||
value += ' if version < {0}'.format(m['scheduled_restart']['postgres_version'])
|
||||
m['scheduled_restart'] = value
|
||||
|
||||
rows.append([m.get(n.lower().replace(' ', '_'), '') for n in columns])
|
||||
|
||||
print_output(columns, rows, {'Lag in MB': 'r', 'TL': 'r', 'Tags': 'l'},
|
||||
fmt, ' Cluster: {0} ({1}) '.format(name, initialize))
|
||||
|
||||
if fmt != 'pretty': # Omit service info when using machine-readable formats
|
||||
return
|
||||
|
||||
service_info = []
|
||||
if cluster.is_paused():
|
||||
if cluster.get('pause'):
|
||||
service_info.append('Maintenance mode: on')
|
||||
|
||||
if cluster.failover and cluster.failover.scheduled_at:
|
||||
info = 'Switchover scheduled at: ' + cluster.failover.scheduled_at.isoformat()
|
||||
if cluster.failover.leader:
|
||||
info += '\n from: ' + cluster.failover.leader
|
||||
if cluster.failover.candidate:
|
||||
info += '\n to: ' + cluster.failover.candidate
|
||||
if 'scheduled_switchover' in cluster:
|
||||
info = 'Switchover scheduled at: ' + cluster['scheduled_switchover']['at']
|
||||
for name in ('from', 'to'):
|
||||
if name in cluster['scheduled_switchover']:
|
||||
info += '\n{0:>24}: {1}'.format(name, cluster['scheduled_switchover'][name])
|
||||
service_info.append(info)
|
||||
|
||||
if service_info:
|
||||
@@ -904,7 +896,7 @@ def scaffold(obj, cluster_name, sysid):
|
||||
click.echo("Cluster {0} has been created successfully".format(cluster_name))
|
||||
|
||||
|
||||
@ctl.command('flush', help='Flush scheduled events')
|
||||
@ctl.command('flush', help='Discard scheduled events (restarts only currently)')
|
||||
@click.argument('cluster_name')
|
||||
@click.argument('member_names', nargs=-1)
|
||||
@click.argument('target', type=click.Choice(['restart']))
|
||||
@@ -919,7 +911,7 @@ def flush(obj, cluster_name, member_names, force, role, target):
|
||||
for member in members:
|
||||
if target == 'restart':
|
||||
if member.data.get('scheduled_restart'):
|
||||
r = request_patroni(member, 'delete', 'restart', None, auth_header(obj))
|
||||
r = request_patroni(member, 'delete', 'restart')
|
||||
check_response(r, member.name, 'flush scheduled restart')
|
||||
else:
|
||||
click.echo('No scheduled restart for member {0}'.format(member.name))
|
||||
@@ -956,20 +948,20 @@ def toggle_pause(config, cluster_name, paused, wait):
|
||||
|
||||
for member in members:
|
||||
try:
|
||||
r = request_patroni(member, 'patch', 'config', {'pause': paused or None}, auth_header(config))
|
||||
r = request_patroni(member, 'patch', 'config', {'pause': paused or None})
|
||||
except Exception as err:
|
||||
logging.warning(str(err))
|
||||
logging.warning('Member %s is not accessible', member.name)
|
||||
continue
|
||||
|
||||
if r.status_code == 200:
|
||||
if r.status == 200:
|
||||
if wait:
|
||||
wait_until_pause_is_applied(dcs, paused, cluster)
|
||||
else:
|
||||
click.echo('Success: cluster management is {0}'.format(paused and 'paused' or 'resumed'))
|
||||
else:
|
||||
click.echo('Failed: {0} cluster management status code={1}, ({2})'.format(
|
||||
paused and 'pause' or 'resume', r.status_code, r.text))
|
||||
paused and 'pause' or 'resume', r.status, r.data.decode('utf-8')))
|
||||
break
|
||||
else:
|
||||
raise PatroniCtlException('Can not find accessible cluster member')
|
||||
@@ -1024,7 +1016,7 @@ def show_diff(before_editing, after_editing):
|
||||
buf = io.StringIO()
|
||||
for line in unified_diff:
|
||||
# Force cast to unicode as difflib on Python 2.7 returns a mix of unicode and str.
|
||||
buf.write(text_type(line))
|
||||
buf.write(six.text_type(line))
|
||||
buf.seek(0)
|
||||
|
||||
class opts:
|
||||
@@ -1037,12 +1029,12 @@ def show_diff(before_editing, after_editing):
|
||||
click.echo(line.rstrip('\n'))
|
||||
|
||||
|
||||
def format_config_for_editing(data):
|
||||
def format_config_for_editing(data, default_flow_style=False):
|
||||
"""Formats configuration as YAML for human consumption.
|
||||
|
||||
:param data: configuration as nested dictionaries
|
||||
:returns unicode YAML of the configuration"""
|
||||
return yaml.safe_dump(data, default_flow_style=False, encoding=None, allow_unicode=True)
|
||||
return yaml.safe_dump(data, default_flow_style=default_flow_style, encoding=None, allow_unicode=True, width=200)
|
||||
|
||||
|
||||
def apply_config_changes(before_editing, data, kvpairs):
|
||||
@@ -1233,8 +1225,8 @@ def version(obj, cluster_name, member_names):
|
||||
if m.api_url:
|
||||
if not member_names or m.name in member_names:
|
||||
try:
|
||||
response = request_patroni(m, 'get', 'patroni')
|
||||
data = response.json()
|
||||
response = request_patroni(m)
|
||||
data = json.loads(response.data.decode('utf-8'))
|
||||
version = data.get('patroni', {}).get('version')
|
||||
pg_version = data.get('server_version')
|
||||
pg_version_str = " PostgreSQL {0}".format(format_pg_version(pg_version)) if pg_version else ""
|
||||
@@ -1243,6 +1235,19 @@ def version(obj, cluster_name, member_names):
|
||||
click.echo("{0}: failed to get version: {1}".format(m.name, e))
|
||||
|
||||
|
||||
@ctl.command('history', help="Show the history of failovers/switchovers")
|
||||
@arg_cluster_name
|
||||
@option_format
|
||||
@click.pass_obj
|
||||
def history(obj, cluster_name, fmt):
|
||||
cluster = get_dcs(obj, cluster_name).get_cluster()
|
||||
history = cluster.history and cluster.history.lines or []
|
||||
for line in history:
|
||||
if len(line) < 4:
|
||||
line.append('')
|
||||
print_output(['TL', 'LSN', 'Reason', 'Timestamp'], history, {'TL': 'r', 'LSN': 'r'}, fmt)
|
||||
|
||||
|
||||
def format_pg_version(version):
|
||||
if version < 100000:
|
||||
return "{0}.{1}.{2}".format(version // 10000, version // 100 % 100, version % 100)
|
||||
|
||||
+66
-43
@@ -66,33 +66,44 @@ def dcs_modules():
|
||||
module_prefix = __package__ + '.'
|
||||
|
||||
if getattr(sys, 'frozen', False):
|
||||
importer = pkgutil.get_importer(dcs_dirname)
|
||||
return [module for module in list(importer.toc) if module.startswith(module_prefix) and module.count('.') == 2]
|
||||
toc = set()
|
||||
for importer in pkgutil.iter_importers(dcs_dirname):
|
||||
if hasattr(importer, 'toc'):
|
||||
toc |= importer.toc
|
||||
return [module for module in toc if module.startswith(module_prefix) and module.count('.') == 2]
|
||||
else:
|
||||
return [module_prefix + name for _, name, is_pkg in pkgutil.iter_modules([dcs_dirname]) if not is_pkg]
|
||||
|
||||
|
||||
def get_dcs(config):
|
||||
available_implementations = set()
|
||||
for module_name in dcs_modules():
|
||||
try:
|
||||
module = importlib.import_module(module_name)
|
||||
for name in filter(lambda name: not name.startswith('__'), dir(module)): # iterate through module content
|
||||
item = getattr(module, name)
|
||||
name = name.lower()
|
||||
# try to find implementation of AbstractDCS interface, class name must match with module_name
|
||||
if inspect.isclass(item) and issubclass(item, AbstractDCS) and __package__ + '.' + name == module_name:
|
||||
available_implementations.add(name)
|
||||
if name in config: # which has configuration section in the config file
|
||||
modules = dcs_modules()
|
||||
|
||||
for module_name in modules:
|
||||
name = module_name.split('.')[-1]
|
||||
if name in config: # we will try to import only modules which have configuration section in the config file
|
||||
try:
|
||||
module = importlib.import_module(module_name)
|
||||
for key, item in module.__dict__.items(): # iterate through the module content
|
||||
# try to find implementation of AbstractDCS interface, class name must match with module_name
|
||||
if key.lower() == name and inspect.isclass(item) and issubclass(item, AbstractDCS):
|
||||
# propagate some parameters
|
||||
config[name].update({p: config[p] for p in ('namespace', 'name', 'scope', 'loop_wait',
|
||||
'patronictl', 'ttl', 'retry_timeout') if p in config})
|
||||
return item(config[name])
|
||||
except ImportError:
|
||||
logger.debug('Failed to import %s', module_name)
|
||||
|
||||
available_implementations = []
|
||||
for module_name in modules:
|
||||
name = module_name.split('.')[-1]
|
||||
try:
|
||||
module = importlib.import_module(module_name)
|
||||
available_implementations.extend(name for key, item in module.__dict__.items() if key.lower() == name
|
||||
and inspect.isclass(item) and issubclass(item, AbstractDCS))
|
||||
except ImportError:
|
||||
if not config.get('patronictl'):
|
||||
logger.info('Failed to import %s', module_name)
|
||||
logger.info('Failed to import %s', module_name)
|
||||
raise PatroniException("""Can not find suitable configuration of distributed configuration store
|
||||
Available implementations: """ + ', '.join(available_implementations))
|
||||
Available implementations: """ + ', '.join(sorted(set(available_implementations))))
|
||||
|
||||
|
||||
class Member(namedtuple('Member', 'index,name,session,data')):
|
||||
@@ -129,10 +140,10 @@ class Member(namedtuple('Member', 'index,name,session,data')):
|
||||
@property
|
||||
def conn_url(self):
|
||||
conn_url = self.data.get('conn_url')
|
||||
conn_kwargs = self.data.get('conn_kwargs')
|
||||
if conn_url:
|
||||
return conn_url
|
||||
|
||||
conn_kwargs = self.data.get('conn_kwargs')
|
||||
if conn_kwargs:
|
||||
conn_url = uri('postgresql', (conn_kwargs.get('host'), conn_kwargs.get('port', 5432)))
|
||||
self.data['conn_url'] = conn_url
|
||||
@@ -140,16 +151,19 @@ class Member(namedtuple('Member', 'index,name,session,data')):
|
||||
|
||||
def conn_kwargs(self, auth=None):
|
||||
defaults = {
|
||||
"host": "",
|
||||
"port": "",
|
||||
"database": ""
|
||||
"host": None,
|
||||
"port": None,
|
||||
"database": None
|
||||
}
|
||||
ret = self.data.get('conn_kwargs')
|
||||
if ret:
|
||||
defaults.update(ret)
|
||||
ret = defaults
|
||||
else:
|
||||
r = urlparse(self.conn_url)
|
||||
conn_url = self.conn_url
|
||||
if not conn_url:
|
||||
return {} # due to the invalid conn_url we don't care about authentication parameters
|
||||
r = urlparse(conn_url)
|
||||
ret = {
|
||||
'host': r.hostname,
|
||||
'port': r.port or 5432,
|
||||
@@ -157,11 +171,11 @@ class Member(namedtuple('Member', 'index,name,session,data')):
|
||||
}
|
||||
self.data['conn_kwargs'] = ret.copy()
|
||||
|
||||
# apply any remaining authentication parameters
|
||||
if auth and isinstance(auth, dict):
|
||||
ret.update({k: v for k, v in auth.items() if v is not None})
|
||||
if 'username' in auth:
|
||||
ret['user'] = auth['username']
|
||||
if 'password' in auth:
|
||||
ret['password'] = auth['password']
|
||||
ret['user'] = ret.pop('username')
|
||||
return ret
|
||||
|
||||
@property
|
||||
@@ -232,21 +246,25 @@ class Leader(namedtuple('Leader', 'index,session,member')):
|
||||
def conn_url(self):
|
||||
return self.member.conn_url
|
||||
|
||||
@property
|
||||
def data(self):
|
||||
return self.member.data
|
||||
|
||||
@property
|
||||
def timeline(self):
|
||||
return self.member.data.get('timeline')
|
||||
return self.data.get('timeline')
|
||||
|
||||
@property
|
||||
def checkpoint_after_promote(self):
|
||||
"""
|
||||
>>> Leader(1, '', Member.from_node(1, '', '', '{"version":"z"}')).checkpoint_after_promote
|
||||
"""
|
||||
version = self.member.data.get('version')
|
||||
version = self.data.get('version')
|
||||
if version:
|
||||
try:
|
||||
# 1.5.6 is the last version which doesn't expose checkpoint_after_promote: false
|
||||
if tuple(map(int, version.split('.'))) > (1, 5, 6):
|
||||
return self.member.data['role'] == 'master' and 'checkpoint_after_promote' not in self.member.data
|
||||
return self.data['role'] == 'master' and 'checkpoint_after_promote' not in self.data
|
||||
except Exception:
|
||||
logger.debug('Failed to parse Patroni version %s', version)
|
||||
|
||||
@@ -323,6 +341,10 @@ class ClusterConfig(namedtuple('ClusterConfig', 'index,data,modify_index')):
|
||||
self.data.get('permanent_slots') or self.data.get('slots')
|
||||
) or {}
|
||||
|
||||
@property
|
||||
def max_timelines_history(self):
|
||||
return self.data.get('max_timelines_history', 0)
|
||||
|
||||
|
||||
class SyncState(namedtuple('SyncState', 'index,leader,sync_standby')):
|
||||
"""Immutable object (namedtuple) which represents last observed synhcronous replication state
|
||||
@@ -437,19 +459,21 @@ class Cluster(namedtuple('Cluster', 'initialize,config,leader,last_leader_operat
|
||||
def is_synchronous_mode(self):
|
||||
return self.check_mode('synchronous_mode')
|
||||
|
||||
def get_replication_slots(self, name, role):
|
||||
def get_replication_slots(self, my_name, role):
|
||||
# if the replicatefrom tag is set on the member - we should not create the replication slot for it on
|
||||
# the current master, because that member would replicate from elsewhere. We still create the slot if
|
||||
# the replicatefrom destination member is currently not a member of the cluster (fallback to the
|
||||
# master), or if replicatefrom destination member happens to be the current master
|
||||
use_slots = self.config and self.config.data.get('postgresql', {}).get('use_slots', True)
|
||||
if role in ('master', 'standby_leader'):
|
||||
slot_members = [m.name for m in self.members if m.name != name and
|
||||
(m.replicatefrom is None or m.replicatefrom == name or
|
||||
slot_members = [m.name for m in self.members if use_slots and m.name != my_name and
|
||||
(m.replicatefrom is None or m.replicatefrom == my_name or
|
||||
not self.has_member(m.replicatefrom))]
|
||||
permanent_slots = (self.config and self.config.permanent_slots or {}).copy()
|
||||
else:
|
||||
# only manage slots for replicas that replicate from this one, except for the leader among them
|
||||
slot_members = [m.name for m in self.members if m.replicatefrom == name and m.name != self.leader.name]
|
||||
slot_members = [m.name for m in self.members if use_slots and
|
||||
m.replicatefrom == my_name and m.name != self.leader.name]
|
||||
permanent_slots = {}
|
||||
|
||||
slots = {slot_name_from_member_name(name): {'type': 'physical'} for name in slot_members}
|
||||
@@ -470,22 +494,21 @@ class Cluster(namedtuple('Cluster', 'initialize,config,leader,last_leader_operat
|
||||
logger.error("Slot name may only contain lower case letters, numbers, and the underscore chars")
|
||||
continue
|
||||
|
||||
if name in slots:
|
||||
logger.error("Permanent replication slot {'%s': %s} is conflicting with" +
|
||||
" physical replication slot for cluster member", name, value)
|
||||
continue
|
||||
|
||||
value = deepcopy(value)
|
||||
if not value:
|
||||
value = {'type': 'physical'}
|
||||
|
||||
value = deepcopy(value) if value else {'type': 'physical'}
|
||||
if isinstance(value, dict):
|
||||
if 'type' not in value:
|
||||
value['type'] = 'logical' if value.get('database') and value.get('plugin') else 'physical'
|
||||
|
||||
if value['type'] == 'physical' or value['type'] == 'logical' \
|
||||
and value.get('database') and value.get('plugin'):
|
||||
slots[name] = value
|
||||
if value['type'] == 'physical':
|
||||
if name != my_name: # Don't try to create permanent physical replication slot for yourself
|
||||
slots[name] = value
|
||||
continue
|
||||
elif value['type'] == 'logical' and value.get('database') and value.get('plugin'):
|
||||
if name in slots:
|
||||
logger.error("Permanent logical replication slot {'%s': %s} is conflicting with" +
|
||||
" physical replication slot for cluster member", name, value)
|
||||
else:
|
||||
slots[name] = value
|
||||
continue
|
||||
|
||||
logger.error("Bad value for slot '%s' in permanent_slots: %s", name, permanent_slots[name])
|
||||
|
||||
@@ -9,13 +9,14 @@ import time
|
||||
import urllib3
|
||||
|
||||
from consul import ConsulException, NotFound, base
|
||||
from patroni.dcs import AbstractDCS, ClusterConfig, Cluster, Failover, Leader, Member, SyncState, TimelineHistory
|
||||
from patroni.exceptions import DCSError
|
||||
from patroni.utils import deep_compare, parse_bool, Retry, RetryFailedError, split_host_port, uri
|
||||
from urllib3.exceptions import HTTPError
|
||||
from six.moves.urllib.parse import urlencode, urlparse, quote
|
||||
from six.moves.http_client import HTTPException
|
||||
|
||||
from . import AbstractDCS, Cluster, ClusterConfig, Failover, Leader, Member, SyncState, TimelineHistory
|
||||
from ..exceptions import DCSError
|
||||
from ..utils import deep_compare, parse_bool, Retry, RetryFailedError, split_host_port, uri, USER_AGENT
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
@@ -52,8 +53,7 @@ class HTTPClient(object):
|
||||
kwargs['cert_file'] = cert
|
||||
if ca_cert:
|
||||
kwargs['ca_certs'] = ca_cert
|
||||
if verify or ca_cert:
|
||||
kwargs['cert_reqs'] = ssl.CERT_REQUIRED
|
||||
kwargs['cert_reqs'] = ssl.CERT_REQUIRED if verify or ca_cert else ssl.CERT_NONE
|
||||
self.http = urllib3.PoolManager(num_pools=10, **kwargs)
|
||||
self._ttl = None
|
||||
|
||||
@@ -111,8 +111,9 @@ class HTTPClient(object):
|
||||
else:
|
||||
kwargs['timeout'] = self._read_timeout
|
||||
token = params.pop('token', self.token) if isinstance(params, dict) else self.token
|
||||
kwargs['headers'] = urllib3.make_headers(user_agent=USER_AGENT)
|
||||
if token:
|
||||
kwargs['headers'] = {'X-Consul-Token': token}
|
||||
kwargs['headers']['X-Consul-Token'] = token
|
||||
return callback(self.response(self.http.request(method.upper(), self.uri(path, params), **kwargs)))
|
||||
return wrapper
|
||||
|
||||
@@ -215,7 +216,7 @@ class Consul(AbstractDCS):
|
||||
self.set_retry_timeout(config['retry_timeout'])
|
||||
self.set_ttl(config.get('ttl') or 30)
|
||||
self._last_session_refresh = 0
|
||||
self.__session_checks = config.get('checks')
|
||||
self.__session_checks = config.get('checks', [])
|
||||
self._register_service = config.get('register_service', False)
|
||||
if self._register_service:
|
||||
self._service_name = service_name_from_scope_name(self._scope)
|
||||
|
||||
+189
-104
@@ -5,26 +5,31 @@ import logging
|
||||
import os
|
||||
import urllib3.util.connection
|
||||
import random
|
||||
import requests
|
||||
import six
|
||||
import socket
|
||||
import time
|
||||
|
||||
from dns.exception import DNSException
|
||||
from dns import resolver
|
||||
from patroni.dcs import AbstractDCS, ClusterConfig, Cluster, Failover, Leader, Member, SyncState, TimelineHistory
|
||||
from patroni.exceptions import DCSError
|
||||
from patroni.utils import Retry, RetryFailedError, split_host_port, uri
|
||||
from urllib3.exceptions import HTTPError, ReadTimeoutError
|
||||
from requests.exceptions import RequestException
|
||||
from urllib3 import Timeout
|
||||
from urllib3.exceptions import HTTPError, ReadTimeoutError, ProtocolError
|
||||
from six.moves.queue import Queue
|
||||
from six.moves.http_client import HTTPException
|
||||
from six.moves.urllib_parse import urlparse
|
||||
from threading import Thread
|
||||
|
||||
from . import AbstractDCS, Cluster, ClusterConfig, Failover, Leader, Member, SyncState, TimelineHistory
|
||||
from ..exceptions import DCSError
|
||||
from ..request import get as requests_get
|
||||
from ..utils import Retry, RetryFailedError, split_host_port, uri, USER_AGENT
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class EtcdRaftInternal(etcd.EtcdException):
|
||||
"""Raft Internal Error"""
|
||||
|
||||
|
||||
class EtcdError(DCSError):
|
||||
pass
|
||||
|
||||
@@ -65,12 +70,15 @@ class DnsCachingResolver(Thread):
|
||||
def resolve_async(self, host, port, attempt=0):
|
||||
self._resolve_queue.put(((host, port), attempt))
|
||||
|
||||
def remove(self, host, port):
|
||||
self._cache.pop((host, port), None)
|
||||
|
||||
@staticmethod
|
||||
def _do_resolve(host, port):
|
||||
try:
|
||||
return socket.getaddrinfo(host, port, 0, socket.SOCK_STREAM, socket.IPPROTO_TCP)
|
||||
except socket.gaierror:
|
||||
logger.warning('failed to resolve host %s', host)
|
||||
except Exception as e:
|
||||
logger.warning('failed to resolve host %s: %s', host, e)
|
||||
return []
|
||||
|
||||
|
||||
@@ -87,29 +95,64 @@ class Client(etcd.Client):
|
||||
# Workaround for the case when https://github.com/jplana/python-etcd/pull/196 is not applied
|
||||
self.http.connection_pool_kw.pop('ssl_version', None)
|
||||
self._config = config
|
||||
self._initial_machines_cache = []
|
||||
self._load_machines_cache()
|
||||
self._allow_reconnect = True
|
||||
# allow passing retry argument to api_execute in params
|
||||
self._comparison_conditions.add('retry')
|
||||
self._read_options.add('retry')
|
||||
self._del_conditions.add('retry')
|
||||
|
||||
def _build_request_parameters(self):
|
||||
def _calculate_timeouts(self, etcd_nodes, timeout=None):
|
||||
"""Calculate a request timeout and number of retries per single etcd node.
|
||||
In case if the timeout per node is too small (less than one second) we will reduce the number of nodes.
|
||||
For the cluster with only one node we will try to do 2 retries.
|
||||
For clusters with 2 nodes we will try to do 1 retry for every node.
|
||||
No retries for clusters with 3 or more nodes. We better rely on switching to a different node."""
|
||||
|
||||
per_node_timeout = timeout = float(timeout or self.read_timeout)
|
||||
|
||||
max_retries = 4 - min(etcd_nodes, 3)
|
||||
per_node_retries = 1
|
||||
min_timeout = 1.0
|
||||
|
||||
while etcd_nodes > 0:
|
||||
per_node_timeout = float(timeout) / etcd_nodes
|
||||
if per_node_timeout >= min_timeout:
|
||||
# for small clusters we will try to do more than on try on every node
|
||||
while per_node_retries < max_retries and per_node_timeout / (per_node_retries + 1) >= min_timeout:
|
||||
per_node_retries += 1
|
||||
per_node_timeout /= per_node_retries
|
||||
break
|
||||
# if the timeout per one node is to small try to reduce number of nodes
|
||||
etcd_nodes -= 1
|
||||
max_retries = 1
|
||||
|
||||
return etcd_nodes, per_node_timeout, per_node_retries - 1
|
||||
|
||||
def _get_headers(self):
|
||||
basic_auth = ':'.join((self.username, self.password)) if self.username and self.password else None
|
||||
return urllib3.make_headers(basic_auth=basic_auth, user_agent=USER_AGENT)
|
||||
|
||||
def _build_request_parameters(self, etcd_nodes, timeout=None):
|
||||
kwargs = {'headers': self._get_headers(), 'redirect': self.allow_redirect}
|
||||
|
||||
# calculate the number of retries and timeout *per node*
|
||||
# actual number of retries depends on the number of nodes
|
||||
etcd_nodes = len(self._machines_cache) + 1
|
||||
kwargs['retries'] = 0 if etcd_nodes > 3 else (1 if etcd_nodes > 1 else 2)
|
||||
|
||||
# if etcd_nodes > 3:
|
||||
# kwargs.update({'retries': 0, 'timeout': float(self.read_timeout)/etcd_nodes})
|
||||
# elif etcd_nodes > 1:
|
||||
# kwargs.update({'retries': 1, 'timeout': self.read_timeout/2.0/etcd_nodes})
|
||||
# else:
|
||||
# kwargs.update({'retries': 2, 'timeout': self.read_timeout/3.0})
|
||||
kwargs['timeout'] = self.read_timeout/float(kwargs['retries'] + 1)/etcd_nodes
|
||||
if timeout is not None:
|
||||
kwargs.update(retries=0, timeout=timeout)
|
||||
else:
|
||||
_, per_node_timeout, per_node_retries = self._calculate_timeouts(etcd_nodes)
|
||||
connect_timeout = max(1, per_node_timeout/2)
|
||||
kwargs.update(timeout=Timeout(connect=connect_timeout, total=per_node_timeout), retries=per_node_retries)
|
||||
return kwargs
|
||||
|
||||
def set_machines_cache_ttl(self, cache_ttl):
|
||||
self._machines_cache_ttl = cache_ttl
|
||||
|
||||
@property
|
||||
def machines_cache(self):
|
||||
base_uri, cache = self._base_uri, self._machines_cache
|
||||
return ([base_uri] if base_uri in cache else []) + [machine for machine in cache if machine != base_uri]
|
||||
|
||||
@property
|
||||
def machines(self):
|
||||
"""Original `machines` method(property) of `etcd.Client` class raise exception
|
||||
@@ -122,59 +165,64 @@ class Client(etcd.Client):
|
||||
Also this method implements the same timeout-retry logic as `api_execute`, because
|
||||
the original method was retrying 2 times with the `read_timeout` on each node."""
|
||||
|
||||
kwargs = self._build_request_parameters()
|
||||
machines_cache = self.machines_cache
|
||||
kwargs = self._build_request_parameters(len(machines_cache))
|
||||
|
||||
while True:
|
||||
for base_uri in machines_cache:
|
||||
try:
|
||||
response = self.http.request(self._MGET, self._base_uri + self.version_prefix + '/machines', **kwargs)
|
||||
response = self.http.request(self._MGET, base_uri + self.version_prefix + '/machines', **kwargs)
|
||||
data = self._handle_server_response(response).data.decode('utf-8')
|
||||
machines = [m.strip() for m in data.split(',') if m.strip()]
|
||||
logger.debug("Retrieved list of machines: %s", machines)
|
||||
if not machines:
|
||||
raise etcd.EtcdException
|
||||
random.shuffle(machines)
|
||||
for url in machines:
|
||||
r = urlparse(url)
|
||||
port = r.port or (443 if r.scheme == 'https' else 80)
|
||||
self._dns_resolver.resolve_async(r.hostname, port)
|
||||
return machines
|
||||
if machines:
|
||||
random.shuffle(machines)
|
||||
self._update_dns_cache(self._dns_resolver.resolve_async, machines)
|
||||
return machines
|
||||
except Exception as e:
|
||||
# We can't get the list of machines, if one server is in the
|
||||
# machines cache, try on it
|
||||
logger.error("Failed to get list of machines from %s%s: %r", self._base_uri, self.version_prefix, e)
|
||||
if self._machines_cache:
|
||||
self._base_uri = self._machines_cache.pop(0)
|
||||
logger.info("Retrying on %s", self._base_uri)
|
||||
elif self._update_machines_cache:
|
||||
raise etcd.EtcdException("Could not get the list of servers, "
|
||||
"maybe you provided the wrong "
|
||||
"host(s) to connect to?")
|
||||
else:
|
||||
return []
|
||||
self.http.clear()
|
||||
logger.error("Failed to get list of machines from %s%s: %r", base_uri, self.version_prefix, e)
|
||||
|
||||
raise etcd.EtcdConnectionFailed('No more machines in the cluster')
|
||||
|
||||
def set_read_timeout(self, timeout):
|
||||
self._read_timeout = timeout
|
||||
|
||||
def _do_http_request(self, request_executor, method, url, fields=None, **kwargs):
|
||||
try:
|
||||
response = request_executor(method, url, fields=fields, **kwargs)
|
||||
response.data.decode('utf-8')
|
||||
self._check_cluster_id(response)
|
||||
except (HTTPError, HTTPException, socket.error, socket.timeout) as e:
|
||||
if (isinstance(fields, dict) and fields.get("wait") == "true" and
|
||||
isinstance(e, ReadTimeoutError)):
|
||||
logger.debug("Watch timed out.")
|
||||
raise etcd.EtcdWatchTimedOut("Watch timed out: {0}".format(e), cause=e)
|
||||
logger.error("Request to server %s failed: %r", self._base_uri, e)
|
||||
logger.info("Reconnection allowed, looking for another server.")
|
||||
self._base_uri = self._next_server(cause=e)
|
||||
response = False
|
||||
return response
|
||||
def _do_http_request(self, retry, machines_cache, request_executor, method, path, fields=None, **kwargs):
|
||||
some_request_failed = False
|
||||
for i, base_uri in enumerate(machines_cache):
|
||||
if i > 0:
|
||||
logger.info("Retrying on %s", base_uri)
|
||||
try:
|
||||
response = request_executor(method, base_uri + path, fields=fields, **kwargs)
|
||||
response.data.decode('utf-8')
|
||||
self._check_cluster_id(response)
|
||||
if some_request_failed:
|
||||
self.set_base_uri(base_uri)
|
||||
self._refresh_machines_cache()
|
||||
return response
|
||||
except (HTTPError, HTTPException, socket.error, socket.timeout) as e:
|
||||
self.http.clear()
|
||||
# switch to the next etcd node because we don't know exactly what happened,
|
||||
# whether the key didn't received an update or there is a network problem.
|
||||
if not retry and i + 1 < len(machines_cache):
|
||||
self.set_base_uri(machines_cache[i + 1])
|
||||
if (isinstance(fields, dict) and fields.get("wait") == "true" and
|
||||
isinstance(e, (ReadTimeoutError, ProtocolError))):
|
||||
logger.debug("Watch timed out.")
|
||||
raise etcd.EtcdWatchTimedOut("Watch timed out: {0}".format(e), cause=e)
|
||||
logger.error("Request to server %s failed: %r", base_uri, e)
|
||||
logger.info("Reconnection allowed, looking for another server.")
|
||||
if not retry:
|
||||
raise etcd.EtcdException('{0} {1} request failed'.format(method, path))
|
||||
some_request_failed = True
|
||||
|
||||
raise etcd.EtcdConnectionFailed('No more machines in the cluster')
|
||||
|
||||
def api_execute(self, path, method, params=None, timeout=None):
|
||||
if not path.startswith('/'):
|
||||
raise ValueError('Path does not start with /')
|
||||
|
||||
retry = params.pop('retry', None) if isinstance(params, dict) else None
|
||||
kwargs = {'fields': params, 'preload_content': False}
|
||||
|
||||
if method in [self._MGET, self._MDELETE]:
|
||||
@@ -190,32 +238,35 @@ class Client(etcd.Client):
|
||||
self._load_machines_cache()
|
||||
elif not self._use_proxies and time.time() - self._machines_cache_updated > self._machines_cache_ttl:
|
||||
self._refresh_machines_cache()
|
||||
self._machines_cache_updated = time.time()
|
||||
|
||||
kwargs.update(self._build_request_parameters())
|
||||
machines_cache = self.machines_cache
|
||||
etcd_nodes = len(machines_cache)
|
||||
kwargs.update(self._build_request_parameters(etcd_nodes, timeout))
|
||||
|
||||
if timeout is not None:
|
||||
kwargs.update({'retries': 0, 'timeout': timeout})
|
||||
|
||||
response = False
|
||||
|
||||
try:
|
||||
some_request_failed = False
|
||||
while not response:
|
||||
response = self._do_http_request(request_executor, method, self._base_uri + path, **kwargs)
|
||||
|
||||
if response is False:
|
||||
some_request_failed = True
|
||||
if some_request_failed:
|
||||
self._refresh_machines_cache()
|
||||
except etcd.EtcdConnectionFailed as e:
|
||||
if isinstance(e, etcd.EtcdWatchTimedOut) and self._machines_cache:
|
||||
self._base_uri = self._next_server()
|
||||
else:
|
||||
self._update_machines_cache = True
|
||||
if not response:
|
||||
while True:
|
||||
try:
|
||||
response = self._do_http_request(retry, machines_cache, request_executor, method, path, **kwargs)
|
||||
return self._handle_server_response(response)
|
||||
except etcd.EtcdWatchTimedOut:
|
||||
raise
|
||||
return self._handle_server_response(response)
|
||||
except etcd.EtcdConnectionFailed as ex:
|
||||
try:
|
||||
if self._load_machines_cache():
|
||||
machines_cache = self.machines_cache
|
||||
etcd_nodes = len(machines_cache)
|
||||
except Exception as e:
|
||||
logger.debug('Failed to update list of etcd nodes: %r', e)
|
||||
sleeptime = retry.sleeptime
|
||||
remaining_time = retry.stoptime - sleeptime - time.time()
|
||||
nodes, timeout, retries = self._calculate_timeouts(etcd_nodes, remaining_time)
|
||||
if nodes == 0:
|
||||
self._update_machines_cache = True
|
||||
raise ex
|
||||
retry.sleep_func(sleeptime)
|
||||
retry.update_delay()
|
||||
# We still have some time left. Partially reduce `machines_cache` and retry request
|
||||
kwargs.update(timeout=Timeout(connect=max(1, timeout/2), total=timeout), retries=retries)
|
||||
machines_cache = machines_cache[:nodes]
|
||||
|
||||
@staticmethod
|
||||
def get_srv_record(host):
|
||||
@@ -237,12 +288,12 @@ class Client(etcd.Client):
|
||||
url = uri(protocol, (host, port), endpoint)
|
||||
if endpoint:
|
||||
try:
|
||||
response = requests.get(url, timeout=self.read_timeout, verify=False)
|
||||
if response.ok:
|
||||
for member in response.json():
|
||||
response = requests_get(url, timeout=self.read_timeout, verify=False)
|
||||
if response.status < 400:
|
||||
for member in json.loads(response.data.decode('utf-8')):
|
||||
ret.extend(member['clientURLs'])
|
||||
break
|
||||
except RequestException:
|
||||
except Exception:
|
||||
logger.exception('GET %s', url)
|
||||
else:
|
||||
ret.append(url)
|
||||
@@ -276,6 +327,13 @@ class Client(etcd.Client):
|
||||
machines_cache = self._get_machines_cache_from_dns(self._config['host'], self._config['port'])
|
||||
return machines_cache
|
||||
|
||||
@staticmethod
|
||||
def _update_dns_cache(func, machines):
|
||||
for url in machines:
|
||||
r = urlparse(url)
|
||||
port = r.port or (443 if r.scheme == 'https' else 80)
|
||||
func(r.hostname, port)
|
||||
|
||||
def _load_machines_cache(self):
|
||||
"""This method should fill up `_machines_cache` from scratch.
|
||||
It could happen only in two cases:
|
||||
@@ -287,23 +345,48 @@ class Client(etcd.Client):
|
||||
if 'srv' not in self._config and 'host' not in self._config and 'hosts' not in self._config:
|
||||
raise Exception('Neither srv, hosts, host nor url are defined in etcd section of config')
|
||||
|
||||
self._machines_cache = self._get_machines_cache_from_config()
|
||||
|
||||
machines_cache = self._get_machines_cache_from_config()
|
||||
# Can not bootstrap list of etcd-cluster members, giving up
|
||||
if not self._machines_cache:
|
||||
if not machines_cache:
|
||||
raise etcd.EtcdException
|
||||
|
||||
# After filling up initial list of machines_cache we should ask etcd-cluster about actual list
|
||||
self._base_uri = self._next_server()
|
||||
self._refresh_machines_cache()
|
||||
# enforce resolving dns name,they might get new ips
|
||||
self._update_dns_cache(self._dns_resolver.remove, machines_cache)
|
||||
|
||||
# The etcd cluster could change its topology over time and depending on how we resolve the initial
|
||||
# topology (list of hosts in the Patroni config or DNS records, A or SRV) we might get into the situation
|
||||
# the the real topology doesn't match anymore with the topology resolved from the configuration file.
|
||||
# In case if the "initial" topology is the same as before we will not override the `_machines_cache`.
|
||||
ret = set(machines_cache) != set(self._initial_machines_cache)
|
||||
if ret:
|
||||
self._initial_machines_cache = self._machines_cache = machines_cache
|
||||
|
||||
# After filling up the initial list of machines_cache we should ask etcd-cluster about actual list
|
||||
self._refresh_machines_cache(True)
|
||||
|
||||
self._update_machines_cache = False
|
||||
return ret
|
||||
|
||||
def _refresh_machines_cache(self, updating_cache=False):
|
||||
if self._use_proxies:
|
||||
self._machines_cache = self._get_machines_cache_from_config()
|
||||
else:
|
||||
try:
|
||||
self._machines_cache = self.machines
|
||||
except etcd.EtcdConnectionFailed:
|
||||
if updating_cache:
|
||||
raise etcd.EtcdException("Could not get the list of servers, "
|
||||
"maybe you provided the wrong "
|
||||
"host(s) to connect to?")
|
||||
return
|
||||
|
||||
if self._base_uri not in self._machines_cache:
|
||||
self.set_base_uri(self._machines_cache[0])
|
||||
self._machines_cache_updated = time.time()
|
||||
|
||||
def _refresh_machines_cache(self):
|
||||
self._machines_cache = self._get_machines_cache_from_config() if self._use_proxies else self.machines
|
||||
if self._base_uri in self._machines_cache:
|
||||
self._machines_cache.remove(self._base_uri)
|
||||
def set_base_uri(self, value):
|
||||
logger.info('Selected new etcd server %s', value)
|
||||
self._base_uri = value
|
||||
|
||||
|
||||
class Etcd(AbstractDCS):
|
||||
@@ -312,15 +395,15 @@ class Etcd(AbstractDCS):
|
||||
super(Etcd, self).__init__(config)
|
||||
self._ttl = int(config.get('ttl') or 30)
|
||||
self._retry = Retry(deadline=config['retry_timeout'], max_delay=1, max_tries=-1,
|
||||
retry_exceptions=(etcd.EtcdLeaderElectionInProgress,
|
||||
etcd.EtcdWatcherCleared,
|
||||
etcd.EtcdEventIndexCleared))
|
||||
retry_exceptions=(etcd.EtcdLeaderElectionInProgress, EtcdRaftInternal))
|
||||
self._client = self.get_etcd_client(config)
|
||||
self.__do_not_watch = False
|
||||
self._has_failed = False
|
||||
|
||||
def retry(self, *args, **kwargs):
|
||||
return self._retry.copy()(*args, **kwargs)
|
||||
retry = self._retry.copy()
|
||||
kwargs['retry'] = retry
|
||||
return retry(*args, **kwargs)
|
||||
|
||||
def _handle_exception(self, e, name='', do_sleep=False, raise_ex=None):
|
||||
if not self._has_failed:
|
||||
@@ -368,7 +451,7 @@ class Etcd(AbstractDCS):
|
||||
config['hosts'] = []
|
||||
for value in hosts:
|
||||
if isinstance(value, six.string_types):
|
||||
config['hosts'].append(uri(protocol, split_host_port(value, default_port)))
|
||||
config['hosts'].append(uri(protocol, split_host_port(value.strip(), default_port)))
|
||||
elif 'host' in config:
|
||||
host, port = split_host_port(config['host'], 2379)
|
||||
config['host'] = host
|
||||
@@ -506,7 +589,7 @@ class Etcd(AbstractDCS):
|
||||
|
||||
@catch_etcd_errors
|
||||
def take_leader(self):
|
||||
return self.retry(self._client.set, self.leader_path, self._name, self._ttl)
|
||||
return self.retry(self._client.write, self.leader_path, self._name, ttl=self._ttl)
|
||||
|
||||
def attempt_to_acquire_leader(self, permanent=False):
|
||||
try:
|
||||
@@ -535,7 +618,7 @@ class Etcd(AbstractDCS):
|
||||
|
||||
@catch_etcd_errors
|
||||
def _update_leader(self):
|
||||
return self.retry(self._client.test_and_set, self.leader_path, self._name, self._name, self._ttl)
|
||||
return self.retry(self._client.write, self.leader_path, self._name, prevValue=self._name, ttl=self._ttl)
|
||||
|
||||
@catch_etcd_errors
|
||||
def initialize(self, create_new=True, sysid=""):
|
||||
@@ -581,7 +664,6 @@ class Etcd(AbstractDCS):
|
||||
# than reestablishing http connection every time from every replica.
|
||||
return True
|
||||
except etcd.EtcdWatchTimedOut:
|
||||
self._client.http.clear()
|
||||
self._has_failed = False
|
||||
return False
|
||||
except (etcd.EtcdEventIndexCleared, etcd.EtcdWatcherCleared): # Watch failed
|
||||
@@ -596,3 +678,6 @@ class Etcd(AbstractDCS):
|
||||
return super(Etcd, self).watch(None, timeout)
|
||||
finally:
|
||||
self.event.clear()
|
||||
|
||||
|
||||
etcd.EtcdError.error_exceptions[300] = EtcdRaftInternal
|
||||
|
||||
@@ -1,11 +1,11 @@
|
||||
import json
|
||||
import logging
|
||||
import random
|
||||
import requests
|
||||
import time
|
||||
|
||||
from patroni.dcs.zookeeper import ZooKeeper
|
||||
from patroni.request import get as requests_get
|
||||
from patroni.utils import uri
|
||||
from requests.exceptions import RequestException
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -48,10 +48,10 @@ class ExhibitorEnsembleProvider(object):
|
||||
random.shuffle(exhibitors)
|
||||
for host in exhibitors:
|
||||
try:
|
||||
response = requests.get(uri('http', (host, self._exhibitor_port), self._uri_path), timeout=self.TIMEOUT)
|
||||
return response.json()
|
||||
except RequestException:
|
||||
pass
|
||||
response = requests_get(uri('http', (host, self._exhibitor_port), self._uri_path), timeout=self.TIMEOUT)
|
||||
return json.loads(response.data.decode('utf-8'))
|
||||
except Exception:
|
||||
logging.debug('Request to %s failed', host)
|
||||
return None
|
||||
|
||||
@property
|
||||
|
||||
+176
-36
@@ -8,11 +8,14 @@ import sys
|
||||
import time
|
||||
|
||||
from kubernetes import client as k8s_client, config as k8s_config, watch as k8s_watch
|
||||
from patroni.dcs import AbstractDCS, ClusterConfig, Cluster, Failover, Leader, Member, SyncState, TimelineHistory
|
||||
from patroni.exceptions import DCSError
|
||||
from patroni.utils import deep_compare, tzutc, Retry, RetryFailedError
|
||||
from urllib3 import Timeout
|
||||
from urllib3.exceptions import HTTPError
|
||||
from six.moves.http_client import HTTPException
|
||||
from threading import Condition, Lock, Thread
|
||||
|
||||
from . import AbstractDCS, Cluster, ClusterConfig, Failover, Leader, Member, SyncState, TimelineHistory
|
||||
from ..exceptions import DCSError
|
||||
from ..utils import deep_compare, Retry, RetryFailedError, tzutc, USER_AGENT
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -28,16 +31,39 @@ class KubernetesRetriableException(k8s_client.rest.ApiException):
|
||||
self.body = orig.body
|
||||
self.headers = orig.headers
|
||||
|
||||
@property
|
||||
def sleeptime(self):
|
||||
try:
|
||||
return int(self.headers['retry-after'])
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
class CoreV1ApiProxy(object):
|
||||
|
||||
def __init__(self, use_endpoints=False):
|
||||
self._api = k8s_client.CoreV1Api()
|
||||
self._api.api_client.user_agent = USER_AGENT
|
||||
self._api.api_client.rest_client.pool_manager.connection_pool_kw['maxsize'] = 10
|
||||
self._request_timeout = None
|
||||
self._use_endpoints = use_endpoints
|
||||
|
||||
def set_timeout(self, timeout):
|
||||
self._request_timeout = (1, timeout / 3.0)
|
||||
def configure_timeouts(self, loop_wait, retry_timeout, ttl):
|
||||
# Normally every loop_wait seconds we should have receive something from the socket.
|
||||
# If we didn't received anything after the loop_wait + retry_timeout it is a time
|
||||
# to start worrying (send keepalive messages). Finally, the connection should be
|
||||
# considered as dead if we received nothing from the socket after the ttl seconds.
|
||||
cnt = 3
|
||||
idle = int(loop_wait + retry_timeout)
|
||||
intvl = max(1, int(float(ttl - idle) / cnt))
|
||||
self._api.api_client.rest_client.pool_manager.connection_pool_kw['socket_options'] = [
|
||||
(socket.SOL_SOCKET, socket.SO_KEEPALIVE, 1),
|
||||
(socket.IPPROTO_TCP, socket.TCP_KEEPIDLE, idle),
|
||||
(socket.IPPROTO_TCP, socket.TCP_KEEPINTVL, intvl),
|
||||
(socket.IPPROTO_TCP, socket.TCP_KEEPCNT, cnt),
|
||||
(socket.IPPROTO_TCP, 18, int(ttl * 1000)) # TCP_USER_TIMEOUT
|
||||
]
|
||||
self._request_timeout = (1, retry_timeout / 3.0)
|
||||
|
||||
def __getattr__(self, func):
|
||||
if func.endswith('_kind'):
|
||||
@@ -49,7 +75,7 @@ class CoreV1ApiProxy(object):
|
||||
try:
|
||||
return getattr(self._api, func)(*args, **kwargs)
|
||||
except k8s_client.rest.ApiException as e:
|
||||
if e.status in (502, 503, 504): # XXX
|
||||
if e.status in (502, 503, 504) or e.headers and 'retry-after' in e.headers: # XXX
|
||||
raise KubernetesRetriableException(e)
|
||||
raise
|
||||
return wrapper
|
||||
@@ -70,6 +96,114 @@ def catch_kubernetes_errors(func):
|
||||
return wrapper
|
||||
|
||||
|
||||
class ObjectCache(Thread):
|
||||
|
||||
def __init__(self, dcs, func, retry, condition):
|
||||
Thread.__init__(self)
|
||||
self.daemon = True
|
||||
self._api_client = k8s_client.ApiClient()
|
||||
self._dcs = dcs
|
||||
self._func = func
|
||||
self._retry = retry
|
||||
self._condition = condition
|
||||
self._is_ready = False
|
||||
self._object_cache = {}
|
||||
self._object_cache_lock = Lock()
|
||||
self._annotations_map = {self._dcs.leader_path: self._dcs._LEADER, self._dcs.config_path: self._dcs._CONFIG}
|
||||
self.start()
|
||||
|
||||
def _list(self):
|
||||
try:
|
||||
return self._func(_request_timeout=(self._retry.deadline, Timeout.DEFAULT_TIMEOUT))
|
||||
except Exception:
|
||||
time.sleep(1)
|
||||
raise
|
||||
|
||||
def _watch(self, resource_version):
|
||||
return self._func(_request_timeout=(self._retry.deadline, Timeout.DEFAULT_TIMEOUT),
|
||||
_preload_content=False, watch=True, resource_version=resource_version)
|
||||
|
||||
def set(self, name, value):
|
||||
with self._object_cache_lock:
|
||||
old_value = self._object_cache.get(name)
|
||||
ret = not old_value or int(old_value.metadata.resource_version) < int(value.metadata.resource_version)
|
||||
if ret:
|
||||
self._object_cache[name] = value
|
||||
return ret, old_value
|
||||
|
||||
def delete(self, name, resource_version):
|
||||
with self._object_cache_lock:
|
||||
old_value = self._object_cache.get(name)
|
||||
ret = old_value and int(old_value.metadata.resource_version) < int(resource_version)
|
||||
if ret:
|
||||
del self._object_cache[name]
|
||||
return not old_value or ret, old_value
|
||||
|
||||
def copy(self):
|
||||
with self._object_cache_lock:
|
||||
return self._object_cache.copy()
|
||||
|
||||
def _build_cache(self):
|
||||
objects = self._list()
|
||||
return_type = 'V1' + objects.kind[:-4]
|
||||
with self._object_cache_lock:
|
||||
self._object_cache = {item.metadata.name: item for item in objects.items}
|
||||
with self._condition:
|
||||
self._is_ready = True
|
||||
self._condition.notify()
|
||||
|
||||
response = self._watch(objects.metadata.resource_version)
|
||||
try:
|
||||
for line in k8s_watch.watch.iter_resp_lines(response):
|
||||
event = json.loads(line)
|
||||
obj = event['object']
|
||||
if obj.get('code') == 410:
|
||||
break
|
||||
|
||||
ev_type = event['type']
|
||||
name = obj['metadata']['name']
|
||||
|
||||
if ev_type in ('ADDED', 'MODIFIED'):
|
||||
obj = k8s_watch.watch.SimpleNamespace(data=json.dumps(obj))
|
||||
obj = self._api_client.deserialize(obj, return_type)
|
||||
success, old_value = self.set(name, obj)
|
||||
if success:
|
||||
new_value = (obj.metadata.annotations or {}).get(self._annotations_map.get(name))
|
||||
elif ev_type == 'DELETED':
|
||||
success, old_value = self.delete(name, obj['metadata']['resourceVersion'])
|
||||
new_value = None
|
||||
else:
|
||||
logger.warning('Unexpected event type: %s', ev_type)
|
||||
continue
|
||||
|
||||
if success and return_type != 'V1Pod':
|
||||
if old_value:
|
||||
old_value = (old_value.metadata.annotations or {}).get(self._annotations_map.get(name))
|
||||
|
||||
if old_value != new_value and \
|
||||
(name != self._dcs.config_path or old_value is not None and new_value is not None):
|
||||
logger.debug('%s changed from %s to %s', name, old_value, new_value)
|
||||
self._dcs.event.set()
|
||||
finally:
|
||||
with self._condition:
|
||||
self._is_ready = False
|
||||
response.close()
|
||||
response.release_conn()
|
||||
|
||||
def run(self):
|
||||
while True:
|
||||
try:
|
||||
self._build_cache()
|
||||
except Exception as e:
|
||||
with self._condition:
|
||||
self._is_ready = False
|
||||
logger.error('ObjectCache.run %r', e)
|
||||
|
||||
def is_ready(self):
|
||||
"""Must be called only when holding the lock on `_condition`"""
|
||||
return self._is_ready
|
||||
|
||||
|
||||
class Kubernetes(AbstractDCS):
|
||||
|
||||
def __init__(self, config):
|
||||
@@ -101,8 +235,7 @@ class Kubernetes(AbstractDCS):
|
||||
self.__subsets = [k8s_client.V1EndpointSubset(addresses=addresses, ports=ports)]
|
||||
self._should_create_config_service = True
|
||||
self._api = CoreV1ApiProxy(use_endpoints)
|
||||
self.set_retry_timeout(config['retry_timeout'])
|
||||
self.set_ttl(config.get('ttl') or 30)
|
||||
self.reload_config(config)
|
||||
self._leader_observed_record = {}
|
||||
self._leader_observed_time = None
|
||||
self._leader_resource_version = None
|
||||
@@ -110,6 +243,16 @@ class Kubernetes(AbstractDCS):
|
||||
self._config_resource_version = None
|
||||
self.__do_not_watch = False
|
||||
|
||||
self._condition = Condition()
|
||||
|
||||
pods_func = functools.partial(self._api.list_namespaced_pod, self._namespace,
|
||||
label_selector=self._label_selector)
|
||||
self._pods = ObjectCache(self, pods_func, self._retry, self._condition)
|
||||
|
||||
kinds_func = functools.partial(self._api.list_namespaced_kind, self._namespace,
|
||||
label_selector=self._label_selector)
|
||||
self._kinds = ObjectCache(self, kinds_func, self._retry, self._condition)
|
||||
|
||||
def retry(self, *args, **kwargs):
|
||||
return self._retry.copy()(*args, **kwargs)
|
||||
|
||||
@@ -131,7 +274,10 @@ class Kubernetes(AbstractDCS):
|
||||
|
||||
def set_retry_timeout(self, retry_timeout):
|
||||
self._retry.deadline = retry_timeout
|
||||
self._api.set_timeout(retry_timeout)
|
||||
|
||||
def reload_config(self, config):
|
||||
super(Kubernetes, self).reload_config(config)
|
||||
self._api.configure_timeouts(self.loop_wait, self._retry.deadline, self.ttl)
|
||||
|
||||
@staticmethod
|
||||
def member(pod):
|
||||
@@ -140,14 +286,21 @@ class Kubernetes(AbstractDCS):
|
||||
member.data['pod_labels'] = pod.metadata.labels
|
||||
return member
|
||||
|
||||
def _wait_caches(self):
|
||||
stop_time = time.time() + self._retry.deadline
|
||||
while not (self._pods.is_ready() and self._kinds.is_ready()):
|
||||
timeout = stop_time - time.time()
|
||||
if timeout <= 0:
|
||||
raise RetryFailedError('Exceeded retry deadline')
|
||||
self._condition.wait(timeout)
|
||||
|
||||
def _load_cluster(self):
|
||||
try:
|
||||
# get list of members
|
||||
response = self.retry(self._api.list_namespaced_pod, self._namespace, label_selector=self._label_selector)
|
||||
members = [self.member(pod) for pod in response.items]
|
||||
with self._condition:
|
||||
self._wait_caches()
|
||||
|
||||
response = self.retry(self._api.list_namespaced_kind, self._namespace, label_selector=self._label_selector)
|
||||
nodes = {item.metadata.name: item for item in response.items}
|
||||
members = [self.member(pod) for pod in self._pods.copy().values()]
|
||||
nodes = self._kinds.copy()
|
||||
|
||||
config = nodes.get(self.config_path)
|
||||
metadata = config and config.metadata
|
||||
@@ -195,12 +348,13 @@ class Kubernetes(AbstractDCS):
|
||||
if metadata:
|
||||
member = Member(-1, leader, None, {})
|
||||
member = ([m for m in members if m.name == leader] or [member])[0]
|
||||
leader = Leader(response.metadata.resource_version, None, member)
|
||||
leader = Leader(metadata.resource_version, None, member)
|
||||
|
||||
# failover key
|
||||
failover = nodes.get(self.failover_path)
|
||||
metadata = failover and failover.metadata
|
||||
failover = Failover.from_node(metadata and metadata.resource_version, metadata and metadata.annotations)
|
||||
failover = Failover.from_node(metadata and metadata.resource_version,
|
||||
metadata and (metadata.annotations or {}).copy())
|
||||
|
||||
# get synchronization state
|
||||
sync = nodes.get(self.sync_path)
|
||||
@@ -279,7 +433,10 @@ class Kubernetes(AbstractDCS):
|
||||
body = k8s_client.V1Endpoints(**endpoints)
|
||||
else:
|
||||
body = k8s_client.V1ConfigMap(metadata=metadata)
|
||||
return self.retry(func, self._namespace, body) if retry else func(self._namespace, body)
|
||||
ret = self.retry(func, self._namespace, body) if retry else func(self._namespace, body)
|
||||
if ret:
|
||||
self._kinds.set(name, ret)
|
||||
return ret
|
||||
|
||||
def patch_or_create_config(self, annotations, resource_version=None, patch=False, retry=True):
|
||||
# SCOPE-config endpoint requires corresponding service otherwise it might be "cleaned" by k8s master
|
||||
@@ -353,7 +510,8 @@ class Kubernetes(AbstractDCS):
|
||||
"""Unused"""
|
||||
|
||||
def manual_failover(self, leader, candidate, scheduled_at=None, index=None):
|
||||
annotations = {'leader': leader or None, 'member': candidate or None, 'scheduled_at': scheduled_at}
|
||||
annotations = {'leader': leader or None, 'member': candidate or None,
|
||||
'scheduled_at': scheduled_at and scheduled_at.isoformat()}
|
||||
patch = bool(self.cluster and isinstance(self.cluster.failover, Failover) and self.cluster.failover.index)
|
||||
return self.patch_or_create(self.failover_path, annotations, index, bool(index or patch), False)
|
||||
|
||||
@@ -417,24 +575,6 @@ class Kubernetes(AbstractDCS):
|
||||
self.__do_not_watch = False
|
||||
return True
|
||||
|
||||
if leader_index:
|
||||
end_time = time.time() + timeout
|
||||
w = k8s_watch.Watch()
|
||||
while timeout >= 1:
|
||||
try:
|
||||
for event in w.stream(self._api.list_namespaced_kind, self._namespace,
|
||||
resource_version=leader_index, timeout_seconds=int(timeout + 0.5),
|
||||
field_selector='metadata.name=' + self.leader_path,
|
||||
_request_timeout=(1, timeout + 1)):
|
||||
return event['raw_object'].get('metadata', {}).get('resourceVersion') != leader_index
|
||||
return False
|
||||
except KeyboardInterrupt:
|
||||
raise
|
||||
except Exception:
|
||||
logger.exception('watch')
|
||||
|
||||
timeout = end_time - time.time()
|
||||
|
||||
try:
|
||||
return super(Kubernetes, self).watch(None, timeout)
|
||||
finally:
|
||||
|
||||
@@ -76,7 +76,7 @@ class ZooKeeper(AbstractDCS):
|
||||
|
||||
self._client.start()
|
||||
|
||||
def _kazoo_connect(self, host, port):
|
||||
def _kazoo_connect(self, *args):
|
||||
"""Kazoo is using Ping's to determine health of connection to zookeeper. If there is no
|
||||
response on Ping after Ping interval (1/2 from read_timeout) it will consider current
|
||||
connection dead and try to connect to another node. Without this "magic" it was taking
|
||||
@@ -88,7 +88,7 @@ class ZooKeeper(AbstractDCS):
|
||||
than loop_wait, because we can spend up to 2 seconds when calling `touch_member()` and
|
||||
`write_leader_optime()` methods, which also may hang..."""
|
||||
|
||||
ret = self._orig_kazoo_connect(host, port)
|
||||
ret = self._orig_kazoo_connect(*args)
|
||||
return max(self.loop_wait - 2, 2)*1000, ret[1]
|
||||
|
||||
def session_listener(self, state):
|
||||
|
||||
@@ -27,3 +27,7 @@ class PostgresConnectionException(PostgresException):
|
||||
|
||||
class WatchdogError(PatroniException):
|
||||
pass
|
||||
|
||||
|
||||
class ConfigParseError(PatroniException):
|
||||
pass
|
||||
|
||||
+114
-87
@@ -3,7 +3,6 @@ import functools
|
||||
import json
|
||||
import logging
|
||||
import psycopg2
|
||||
import requests
|
||||
import sys
|
||||
import time
|
||||
import uuid
|
||||
@@ -15,7 +14,7 @@ from patroni.exceptions import DCSError, PostgresConnectionException, PatroniExc
|
||||
from patroni.postgresql import ACTION_ON_START, ACTION_ON_ROLE_CHANGE
|
||||
from patroni.postgresql.misc import postgres_version_to_int
|
||||
from patroni.postgresql.rewind import Rewind
|
||||
from patroni.utils import polling_loop, tzutc
|
||||
from patroni.utils import polling_loop, tzutc, is_standby_cluster as _is_standby_cluster, parse_int
|
||||
from patroni.dcs import RemoteMember
|
||||
from threading import RLock
|
||||
|
||||
@@ -95,6 +94,11 @@ class Ha(object):
|
||||
else:
|
||||
return self.patroni.config.check_mode(mode)
|
||||
|
||||
def master_stop_timeout(self):
|
||||
""" Master stop timeout """
|
||||
ret = parse_int(self.patroni.config['master_stop_timeout'])
|
||||
return ret if ret and ret > 0 and self.is_synchronous_mode() else None
|
||||
|
||||
def is_paused(self):
|
||||
return self.check_mode('pause')
|
||||
|
||||
@@ -109,9 +113,7 @@ class Ha(object):
|
||||
return config.get('standby_cluster')
|
||||
|
||||
def is_standby_cluster(self):
|
||||
config = self.get_standby_cluster_config()
|
||||
# Check whether or not provided configuration describes a standby cluster
|
||||
return isinstance(config, dict) and (config.get('host') or config.get('port') or config.get('restore_command'))
|
||||
return _is_standby_cluster(self.get_standby_cluster_config())
|
||||
|
||||
def is_leader(self):
|
||||
with self._is_leader_lock:
|
||||
@@ -194,10 +196,20 @@ class Ha(object):
|
||||
if self._async_executor.scheduled_action in (None, 'promote') \
|
||||
and data['state'] in ['running', 'restarting', 'starting']:
|
||||
try:
|
||||
timeline, wal_position = self.state_handler.timeline_wal_position()
|
||||
timeline, wal_position, pg_control_timeline = self.state_handler.timeline_wal_position()
|
||||
data['xlog_location'] = wal_position
|
||||
if not timeline:
|
||||
timeline = self.state_handler.replica_cached_timeline(self._leader_timeline)
|
||||
# So far the only way to get the current timeline on the standby is from
|
||||
# the replication connection. In order to avoid opening the replication
|
||||
# connection on every iteration of HA loop we will do it only when noticed
|
||||
# that the timeline on the primary has changed.
|
||||
# Unfortunately such optimization isn't possible on the standby_leader,
|
||||
# therefore we will get the timeline from pg_control, either by calling
|
||||
# pg_control_checkpoint() on 9.6+ or by parsing the output of pg_controldata.
|
||||
if self.state_handler.role == 'standby_leader':
|
||||
timeline = pg_control_timeline or self.state_handler.pg_control_timeline()
|
||||
else:
|
||||
timeline = self.state_handler.replica_cached_timeline(self._leader_timeline)
|
||||
if timeline:
|
||||
data['timeline'] = timeline
|
||||
except Exception:
|
||||
@@ -231,9 +243,8 @@ class Ha(object):
|
||||
clone_member = self.cluster.get_clone_member(self.state_handler.name)
|
||||
member_role = 'leader' if clone_member == self.cluster.leader else 'replica'
|
||||
msg = "from {0} '{1}'".format(member_role, clone_member.name)
|
||||
self._async_executor.schedule('bootstrap {0}'.format(msg))
|
||||
self._async_executor.run_async(self.clone, args=(clone_member, msg))
|
||||
return 'trying to bootstrap {0}'.format(msg)
|
||||
ret = self._async_executor.try_run_async('bootstrap {0}'.format(msg), self.clone, args=(clone_member, msg))
|
||||
return ret or 'trying to bootstrap {0}'.format(msg)
|
||||
|
||||
# no initialize key and node is allowed to be master and has 'bootstrap' section in a configuration file
|
||||
elif self.cluster.initialize is None and not self.patroni.nofailover and 'bootstrap' in self.patroni.config:
|
||||
@@ -242,16 +253,12 @@ class Ha(object):
|
||||
self._post_bootstrap_task = CriticalTask()
|
||||
|
||||
if self.is_standby_cluster():
|
||||
self._async_executor.schedule('bootstrap_standby_leader')
|
||||
self._async_executor.run_async(self.bootstrap_standby_leader)
|
||||
return 'trying to bootstrap a new standby leader'
|
||||
ret = self._async_executor.try_run_async('bootstrap_standby_leader', self.bootstrap_standby_leader)
|
||||
return ret or 'trying to bootstrap a new standby leader'
|
||||
else:
|
||||
self._async_executor.schedule('bootstrap')
|
||||
self._async_executor.run_async(
|
||||
self.state_handler.bootstrap.bootstrap,
|
||||
args=(self.patroni.config['bootstrap'],)
|
||||
)
|
||||
return 'trying to bootstrap a new cluster'
|
||||
ret = self._async_executor.try_run_async('bootstrap', self.state_handler.bootstrap.bootstrap,
|
||||
args=(self.patroni.config['bootstrap'],))
|
||||
return ret or 'trying to bootstrap a new cluster'
|
||||
else:
|
||||
return 'failed to acquire initialize lock'
|
||||
else:
|
||||
@@ -259,9 +266,7 @@ class Ha(object):
|
||||
if self.is_standby_cluster() else None
|
||||
if self.state_handler.can_create_replica_without_replication_connection(create_replica_methods):
|
||||
msg = 'bootstrap (without leader)'
|
||||
self._async_executor.schedule(msg)
|
||||
self._async_executor.run_async(self.clone)
|
||||
return 'trying to ' + msg
|
||||
return self._async_executor.try_run_async(msg, self.clone) or 'trying to ' + msg
|
||||
return 'waiting for {0}leader to bootstrap'.format('standby_' if self.is_standby_cluster() else '')
|
||||
|
||||
def bootstrap_standby_leader(self):
|
||||
@@ -284,15 +289,13 @@ class Ha(object):
|
||||
return None
|
||||
|
||||
if self._rewind.can_rewind:
|
||||
self._async_executor.schedule('running pg_rewind from ' + leader.name)
|
||||
self._async_executor.run_async(self._rewind.execute, (leader,))
|
||||
return True
|
||||
msg = 'running pg_rewind from ' + leader.name
|
||||
return self._async_executor.try_run_async(msg, self._rewind.execute, args=(leader,)) or msg
|
||||
|
||||
# remove_data_directory_on_diverged_timelines is set
|
||||
if not self.is_standby_cluster():
|
||||
self._async_executor.schedule('reinitializing due to diverged timelines')
|
||||
self._async_executor.run_async(self._do_reinitialize, args=(self.cluster, ))
|
||||
return True
|
||||
msg = 'reinitializing due to diverged timelines'
|
||||
return self._async_executor.try_run_async(msg, self._do_reinitialize, args=(self.cluster,)) or msg
|
||||
|
||||
def recover(self):
|
||||
# Postgres is not running and we will restart in standby mode. Watchdog is not needed until we promote.
|
||||
@@ -319,9 +322,8 @@ class Ha(object):
|
||||
and not self._crash_recovery_executed and \
|
||||
(self.cluster.is_unlocked() or self._rewind.can_rewind):
|
||||
self._crash_recovery_executed = True
|
||||
self._async_executor.schedule('doing crash recovery in a single user mode')
|
||||
self._async_executor.run_async(self.state_handler.fix_cluster_state)
|
||||
return self._async_executor.scheduled_action
|
||||
msg = 'doing crash recovery in a single user mode'
|
||||
return self._async_executor.try_run_async(msg, self.state_handler.fix_cluster_state) or msg
|
||||
|
||||
self.load_cluster_from_dcs()
|
||||
|
||||
@@ -329,8 +331,9 @@ class Ha(object):
|
||||
if self.is_standby_cluster() or not self.has_lock():
|
||||
if not self._rewind.executed:
|
||||
self._rewind.trigger_check_diverged_lsn()
|
||||
if self._handle_rewind_or_reinitialize():
|
||||
return self._async_executor.scheduled_action
|
||||
msg = self._handle_rewind_or_reinitialize()
|
||||
if msg:
|
||||
return msg
|
||||
|
||||
if self.has_lock(): # in standby cluster
|
||||
msg = "starting as a standby leader because i had the session lock"
|
||||
@@ -346,23 +349,34 @@ class Ha(object):
|
||||
msg = "starting as readonly because i had the session lock"
|
||||
node_to_follow = None
|
||||
|
||||
self.recovering = True
|
||||
|
||||
self._async_executor.schedule('restarting after failure')
|
||||
self._async_executor.run_async(self.state_handler.follow, (node_to_follow, role, timeout))
|
||||
if self._async_executor.try_run_async('restarting after failure', self.state_handler.follow,
|
||||
args=(node_to_follow, role, timeout)) is None:
|
||||
self.recovering = True
|
||||
return msg
|
||||
|
||||
def _get_node_to_follow(self, cluster):
|
||||
# determine the node to follow. If replicatefrom tag is set,
|
||||
# try to follow the node mentioned there, otherwise, follow the leader.
|
||||
if self.is_standby_cluster() and (self.cluster.is_unlocked() or self.has_lock(False)):
|
||||
standby_config = self.get_standby_cluster_config()
|
||||
is_standby_cluster = _is_standby_cluster(standby_config)
|
||||
if is_standby_cluster and (self.cluster.is_unlocked() or self.has_lock(False)):
|
||||
node_to_follow = self.get_remote_master()
|
||||
elif self.patroni.replicatefrom and self.patroni.replicatefrom != self.state_handler.name:
|
||||
node_to_follow = cluster.get_member(self.patroni.replicatefrom)
|
||||
else:
|
||||
node_to_follow = cluster.leader
|
||||
|
||||
return node_to_follow if node_to_follow and node_to_follow.name != self.state_handler.name else None
|
||||
node_to_follow = node_to_follow if node_to_follow and node_to_follow.name != self.state_handler.name else None
|
||||
|
||||
if node_to_follow and not isinstance(node_to_follow, RemoteMember):
|
||||
# we are going to abuse Member.data to pass following parameters
|
||||
params = ('restore_command', 'archive_cleanup_command')
|
||||
for param in params: # It is highly unlikely to happen, but we want to protect from the case
|
||||
node_to_follow.data.pop(param, None) # when above-mentioned params came from outside.
|
||||
if is_standby_cluster:
|
||||
node_to_follow.data.update({p: standby_config[p] for p in params if standby_config.get(p)})
|
||||
|
||||
return node_to_follow
|
||||
|
||||
def follow(self, demote_reason, follow_reason, refresh=True):
|
||||
if refresh:
|
||||
@@ -384,20 +398,26 @@ class Ha(object):
|
||||
self.demote('immediate-nolock')
|
||||
return demote_reason
|
||||
|
||||
if self._handle_rewind_or_reinitialize():
|
||||
return self._async_executor.scheduled_action
|
||||
msg = self._handle_rewind_or_reinitialize()
|
||||
if msg:
|
||||
return msg
|
||||
|
||||
role = 'standby_leader' if isinstance(node_to_follow, RemoteMember) and self.has_lock(False) else 'replica'
|
||||
# It might happen that leader key in the standby cluster references non-exiting member.
|
||||
# In this case it is safe to continue running without changing recovery.conf
|
||||
if self.is_standby_cluster() and role == 'replica' and not (node_to_follow and node_to_follow.conn_url):
|
||||
return 'continue following the old known standby leader'
|
||||
elif not self.state_handler.config.check_recovery_conf(node_to_follow):
|
||||
self._async_executor.schedule('changing primary_conninfo and restarting')
|
||||
self._async_executor.run_async(self.state_handler.follow, (node_to_follow, role))
|
||||
elif role == 'standby_leader' and self.state_handler.role != role:
|
||||
self.state_handler.set_role(role)
|
||||
self.state_handler.call_nowait(ACTION_ON_ROLE_CHANGE)
|
||||
else:
|
||||
change_required, restart_required = self.state_handler.config.check_recovery_conf(node_to_follow)
|
||||
if change_required:
|
||||
if restart_required:
|
||||
self._async_executor.try_run_async('changing primary_conninfo and restarting',
|
||||
self.state_handler.follow, args=(node_to_follow, role))
|
||||
else:
|
||||
self.state_handler.follow(node_to_follow, role, do_reload=True)
|
||||
elif role == 'standby_leader' and self.state_handler.role != role:
|
||||
self.state_handler.set_role(role)
|
||||
self.state_handler.call_nowait(ACTION_ON_ROLE_CHANGE)
|
||||
|
||||
return follow_reason
|
||||
|
||||
@@ -506,6 +526,7 @@ class Ha(object):
|
||||
cluster_history = {l[0]: l for l in cluster_history or []}
|
||||
history = self.state_handler.get_history(master_timeline)
|
||||
if history:
|
||||
history = history[-self.cluster.config.max_timelines_history:]
|
||||
for line in history:
|
||||
# enrich current history with promotion timestamps stored in DCS
|
||||
if len(line) == 3 and line[0] in cluster_history \
|
||||
@@ -553,21 +574,21 @@ class Ha(object):
|
||||
self._rewind.reset_state()
|
||||
logger.info("cleared rewind state after becoming the leader")
|
||||
|
||||
self._async_executor.schedule('promote')
|
||||
self._async_executor.run_async(self.state_handler.promote,
|
||||
args=(self.dcs.loop_wait, on_success, self._leader_access_is_restricted))
|
||||
self._async_executor.try_run_async('promote', self.state_handler.promote,
|
||||
args=(self.dcs.loop_wait, on_success,
|
||||
self._leader_access_is_restricted))
|
||||
return promote_message
|
||||
|
||||
@staticmethod
|
||||
def fetch_node_status(member):
|
||||
def fetch_node_status(self, member):
|
||||
"""This function perform http get request on member.api_url and fetches its status
|
||||
:returns: `_MemberStatus` object
|
||||
"""
|
||||
|
||||
try:
|
||||
response = requests.get(member.api_url, timeout=2, verify=False)
|
||||
logger.info('Got response from %s %s: %s', member.name, member.api_url, response.content)
|
||||
return _MemberStatus.from_api_response(member, response.json())
|
||||
response = self.patroni.request(member, timeout=2, retries=0)
|
||||
data = response.data.decode('utf-8')
|
||||
logger.info('Got response from %s %s: %s', member.name, member.api_url, data)
|
||||
return _MemberStatus.from_api_response(member, json.loads(data))
|
||||
except Exception as e:
|
||||
logger.warning("Request failed to %s: GET %s (%s)", member.name, member.api_url, e)
|
||||
return _MemberStatus.unknown(member)
|
||||
@@ -591,7 +612,8 @@ class Ha(object):
|
||||
def _is_healthiest_node(self, members, check_replication_lag=True):
|
||||
"""This method tries to determine whether I am healthy enough to became a new leader candidate or not."""
|
||||
|
||||
_, my_wal_position = self.state_handler.timeline_wal_position()
|
||||
# We don't call `last_operation()` here because it returns a string
|
||||
_, my_wal_position, _ = self.state_handler.timeline_wal_position()
|
||||
if check_replication_lag and self.is_lagging(my_wal_position):
|
||||
logger.info('My wal position exceeds maximum replication lag')
|
||||
return False # Too far behind last reported wal position on master
|
||||
@@ -756,7 +778,8 @@ class Ha(object):
|
||||
|
||||
self._rewind.trigger_check_diverged_lsn()
|
||||
self.state_handler.stop(mode_control['stop'], checkpoint=mode_control['checkpoint'],
|
||||
on_safepoint=self.watchdog.disable if self.watchdog.is_running else None)
|
||||
on_safepoint=self.watchdog.disable if self.watchdog.is_running else None,
|
||||
stop_timeout=self.master_stop_timeout())
|
||||
self.state_handler.set_role('demoted')
|
||||
self.set_is_leader(False)
|
||||
|
||||
@@ -774,8 +797,7 @@ class Ha(object):
|
||||
# there could be an async action already running, calling follow from here will lead
|
||||
# to racy state handler state updates.
|
||||
if mode_control['async_req']:
|
||||
self._async_executor.schedule('starting after demotion')
|
||||
self._async_executor.run_async(self.state_handler.follow, (node_to_follow,))
|
||||
self._async_executor.try_run_async('starting after demotion', self.state_handler.follow, (node_to_follow,))
|
||||
else:
|
||||
if self.is_synchronous_mode():
|
||||
self.state_handler.config.set_synchronous_standby(None)
|
||||
@@ -848,9 +870,8 @@ class Ha(object):
|
||||
members = [m for m in self.cluster.members
|
||||
if not failover.candidate or m.name == failover.candidate]
|
||||
if self.is_failover_possible(members): # check that there are healthy members
|
||||
self._async_executor.schedule('manual failover: demote')
|
||||
self._async_executor.run_async(self.demote, ('graceful',))
|
||||
return 'manual failover: demoting myself'
|
||||
ret = self._async_executor.try_run_async('manual failover: demote', self.demote, ('graceful',))
|
||||
return ret or 'manual failover: demoting myself'
|
||||
else:
|
||||
logger.warning('manual failover: no healthy members found, failover is not possible')
|
||||
else:
|
||||
@@ -924,7 +945,7 @@ class Ha(object):
|
||||
return msg
|
||||
|
||||
# check if the node is ready to be used by pg_rewind
|
||||
self._rewind.check_for_checkpoint_after_promote()
|
||||
self._rewind.ensure_checkpoint_after_promote()
|
||||
|
||||
if self.is_standby_cluster():
|
||||
# in case of standby cluster we don't really need to
|
||||
@@ -1069,7 +1090,7 @@ class Ha(object):
|
||||
return (False, 'restart failed')
|
||||
|
||||
def _do_reinitialize(self, cluster):
|
||||
self.state_handler.stop('immediate')
|
||||
self.state_handler.stop('immediate', stop_timeout=self.patroni.config['retry_timeout'])
|
||||
# Commented redundant data directory cleanup here
|
||||
# self.state_handler.remove_data_directory()
|
||||
|
||||
@@ -1129,17 +1150,20 @@ class Ha(object):
|
||||
if not self.state_handler.is_running():
|
||||
self.watchdog.disable()
|
||||
if self.has_lock():
|
||||
self.state_handler.set_role('demoted')
|
||||
if self.state_handler.role in ('master', 'standby_leader'):
|
||||
self.state_handler.set_role('demoted')
|
||||
self._delete_leader()
|
||||
return 'removed leader key after trying and failing to start postgres'
|
||||
return 'failed to start postgres'
|
||||
self._crash_recovery_executed = False
|
||||
if self._rewind.executed and not self._rewind.failed:
|
||||
self._rewind.reset_state()
|
||||
return None
|
||||
|
||||
def cancel_initialization(self):
|
||||
logger.info('removing initialize key after failed attempt to bootstrap the cluster')
|
||||
self.dcs.cancel_initialization()
|
||||
self.state_handler.stop('immediate')
|
||||
self.state_handler.stop('immediate', stop_timeout=self.patroni.config['retry_timeout'])
|
||||
self.state_handler.move_data_directory()
|
||||
raise PatroniException('Failed to bootstrap cluster')
|
||||
|
||||
@@ -1153,16 +1177,16 @@ class Ha(object):
|
||||
return 'waiting for end of recovery after bootstrap'
|
||||
|
||||
self.state_handler.set_role('master')
|
||||
self._async_executor.schedule('post_bootstrap')
|
||||
self._async_executor.run_async(self.state_handler.bootstrap.post_bootstrap,
|
||||
args=(self.patroni.config['bootstrap'], self._post_bootstrap_task))
|
||||
return 'running post_bootstrap'
|
||||
ret = self._async_executor.try_run_async('post_bootstrap', self.state_handler.bootstrap.post_bootstrap,
|
||||
args=(self.patroni.config['bootstrap'], self._post_bootstrap_task))
|
||||
return ret or 'running post_bootstrap'
|
||||
|
||||
self.state_handler.bootstrapping = False
|
||||
self.dcs.set_config_value(json.dumps(self.patroni.config.dynamic_configuration, separators=(',', ':')))
|
||||
if not self.watchdog.activate():
|
||||
logger.error('Cancelling bootstrap because watchdog activation failed')
|
||||
self.cancel_initialization()
|
||||
self.dcs.initialize(create_new=(self.cluster.initialize is None), sysid=self.state_handler.sysid)
|
||||
self.dcs.set_config_value(json.dumps(self.patroni.config.dynamic_configuration, separators=(',', ':')))
|
||||
self.state_handler.slots_handler.sync_replication_slots(self.cluster)
|
||||
self.dcs.take_leader()
|
||||
self.set_is_leader(True)
|
||||
@@ -1262,7 +1286,7 @@ class Ha(object):
|
||||
# is data directory empty?
|
||||
if self.state_handler.data_directory_empty():
|
||||
self.state_handler.set_role('uninitialized')
|
||||
self.state_handler.stop('immediate')
|
||||
self.state_handler.stop('immediate', stop_timeout=self.patroni.config['retry_timeout'])
|
||||
# In case datadir went away while we were master.
|
||||
self.watchdog.disable()
|
||||
|
||||
@@ -1272,26 +1296,28 @@ class Ha(object):
|
||||
return 'released leader key voluntarily as data dir empty and currently leader'
|
||||
|
||||
return self.bootstrap() # new node
|
||||
# "bootstrap", but data directory is not empty
|
||||
elif not self.sysid_valid(self.cluster.initialize) and self.cluster.is_unlocked() and not self.is_paused():
|
||||
if not self.state_handler.cb_called and self.state_handler.is_running() \
|
||||
and not self.state_handler.is_leader():
|
||||
self._join_aborted = True
|
||||
logger.error('No initialize key in DCS and PostgreSQL is running as replica, aborting start')
|
||||
logger.error('Please first start Patroni on the node running as master')
|
||||
sys.exit(1)
|
||||
self.dcs.initialize(create_new=(self.cluster.initialize is None), sysid=self.state_handler.sysid)
|
||||
else:
|
||||
# check if we are allowed to join
|
||||
data_sysid = self.state_handler.sysid
|
||||
if not self.sysid_valid(data_sysid):
|
||||
# data directory is not empty, but no valid sysid, cluster must be broken, suggest reinit
|
||||
return "data dir for the cluster is not empty, but system ID is invalid; consider doing reinitalize"
|
||||
return ("data dir for the cluster is not empty, "
|
||||
"but system ID is invalid; consider doing reinitialize")
|
||||
|
||||
if self.sysid_valid(self.cluster.initialize) and self.cluster.initialize != self.state_handler.sysid:
|
||||
logger.fatal("system ID mismatch, node %s belongs to a different cluster: %s != %s",
|
||||
self.state_handler.name, self.cluster.initialize, self.state_handler.sysid)
|
||||
sys.exit(1)
|
||||
if self.sysid_valid(self.cluster.initialize):
|
||||
if self.cluster.initialize != data_sysid:
|
||||
logger.fatal("system ID mismatch, node %s belongs to a different cluster: %s != %s",
|
||||
self.state_handler.name, self.cluster.initialize, data_sysid)
|
||||
sys.exit(1)
|
||||
elif self.cluster.is_unlocked() and not self.is_paused():
|
||||
# "bootstrap", but data directory is not empty
|
||||
if not self.state_handler.cb_called and self.state_handler.is_running() \
|
||||
and not self.state_handler.is_leader():
|
||||
self._join_aborted = True
|
||||
logger.error('No initialize key in DCS and PostgreSQL is running as replica, aborting start')
|
||||
logger.error('Please first start Patroni on the node running as master')
|
||||
sys.exit(1)
|
||||
self.dcs.initialize(create_new=(self.cluster.initialize is None), sysid=data_sysid)
|
||||
|
||||
if not self.state_handler.is_healthy():
|
||||
if self.is_paused():
|
||||
@@ -1354,7 +1380,8 @@ class Ha(object):
|
||||
# This might not be the desired behavior of users, as a graceful shutdown of the host can mean lost data.
|
||||
# We probably need to something smarter here.
|
||||
disable_wd = self.watchdog.disable if self.watchdog.is_running else None
|
||||
self.while_not_sync_standby(lambda: self.state_handler.stop(checkpoint=False, on_safepoint=disable_wd))
|
||||
self.while_not_sync_standby(lambda: self.state_handler.stop(checkpoint=False, on_safepoint=disable_wd,
|
||||
stop_timeout=self.master_stop_timeout()))
|
||||
if not self.state_handler.is_running():
|
||||
if self.has_lock():
|
||||
self.dcs.delete_leader()
|
||||
|
||||
+51
-15
@@ -11,6 +11,20 @@ from threading import Lock, Thread
|
||||
_LOGGER = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def debug_exception(logger_obj, msg, *args, **kwargs):
|
||||
kwargs.pop("exc_info", False)
|
||||
if logger_obj.isEnabledFor(logging.DEBUG):
|
||||
logger_obj.debug(msg, *args, exc_info=True, **kwargs)
|
||||
else:
|
||||
msg = "{0}, DETAIL: '{1}'".format(msg, sys.exc_info()[1])
|
||||
logger_obj.error(msg, *args, exc_info=False, **kwargs)
|
||||
|
||||
|
||||
def error_exception(logger_obj, msg, *args, **kwargs):
|
||||
exc_info = kwargs.pop("exc_info", True)
|
||||
logger_obj.error(msg, *args, exc_info=exc_info, **kwargs)
|
||||
|
||||
|
||||
class QueueHandler(logging.Handler):
|
||||
|
||||
def __init__(self):
|
||||
@@ -48,9 +62,20 @@ class QueueHandler(logging.Handler):
|
||||
return self._records_lost
|
||||
|
||||
|
||||
class ProxyHandler(logging.Handler):
|
||||
|
||||
def __init__(self, patroni_logger):
|
||||
logging.Handler.__init__(self)
|
||||
self.patroni_logger = patroni_logger
|
||||
|
||||
def emit(self, record):
|
||||
self.patroni_logger.log_handler.handle(record)
|
||||
|
||||
|
||||
class PatroniLogger(Thread):
|
||||
|
||||
DEFAULT_LEVEL = 'INFO'
|
||||
DEFAULT_TRACEBACK_LEVEL = 'ERROR'
|
||||
DEFAULT_FORMAT = '%(asctime)s %(levelname)s: %(message)s'
|
||||
|
||||
NORMAL_LOG_QUEUE_SIZE = 2 # When everything goes normal Patroni writes only 2 messages per HA loop
|
||||
@@ -59,16 +84,18 @@ class PatroniLogger(Thread):
|
||||
|
||||
def __init__(self):
|
||||
super(PatroniLogger, self).__init__()
|
||||
self.daemon = True
|
||||
self._queue_handler = QueueHandler()
|
||||
self._root_logger = logging.getLogger()
|
||||
self._root_logger.addHandler(self._queue_handler)
|
||||
self._config = None
|
||||
self._log_handler = None
|
||||
self.log_handler = None
|
||||
self.log_handler_lock = Lock()
|
||||
self._old_handlers = []
|
||||
self._log_handler_lock = Lock()
|
||||
self.reload_config({'level': 'DEBUG'})
|
||||
self.start()
|
||||
# We will switch to the QueueHandler only when thread was started.
|
||||
# This is necessary to protect from the cases when Patroni constructor
|
||||
# failed and PatroniLogger thread remain running and prevent shutdown.
|
||||
self._proxy_handler = ProxyHandler(self)
|
||||
self._root_logger.addHandler(self._proxy_handler)
|
||||
|
||||
def update_loggers(self):
|
||||
loggers = deepcopy(self._config.get('loggers') or {})
|
||||
@@ -87,18 +114,22 @@ class PatroniLogger(Thread):
|
||||
self._queue_handler.queue.maxsize = config.get('max_queue_size', self.DEFAULT_MAX_QUEUE_SIZE)
|
||||
|
||||
self._root_logger.setLevel(config.get('level', PatroniLogger.DEFAULT_LEVEL))
|
||||
if config.get('traceback_level', PatroniLogger.DEFAULT_TRACEBACK_LEVEL).lower() == 'debug':
|
||||
logging.Logger.exception = debug_exception
|
||||
else:
|
||||
logging.Logger.exception = error_exception
|
||||
|
||||
new_handler = None
|
||||
if 'dir' in config:
|
||||
if not isinstance(self._log_handler, RotatingFileHandler):
|
||||
if not isinstance(self.log_handler, RotatingFileHandler):
|
||||
new_handler = RotatingFileHandler(os.path.join(config['dir'], __name__))
|
||||
handler = new_handler or self._log_handler
|
||||
handler = new_handler or self.log_handler
|
||||
handler.maxBytes = int(config.get('file_size', 25000000))
|
||||
handler.backupCount = int(config.get('file_num', 4))
|
||||
else:
|
||||
if self._log_handler is None or isinstance(self._log_handler, RotatingFileHandler):
|
||||
if self.log_handler is None or isinstance(self.log_handler, RotatingFileHandler):
|
||||
new_handler = logging.StreamHandler()
|
||||
handler = new_handler or self._log_handler
|
||||
handler = new_handler or self.log_handler
|
||||
|
||||
oldlogformat = (self._config or {}).get('format', PatroniLogger.DEFAULT_FORMAT)
|
||||
logformat = config.get('format', PatroniLogger.DEFAULT_FORMAT)
|
||||
@@ -110,17 +141,17 @@ class PatroniLogger(Thread):
|
||||
handler.setFormatter(logging.Formatter(logformat, dateformat))
|
||||
|
||||
if new_handler:
|
||||
with self._log_handler_lock:
|
||||
if self._log_handler:
|
||||
self._old_handlers.append(self._log_handler)
|
||||
self._log_handler = new_handler
|
||||
with self.log_handler_lock:
|
||||
if self.log_handler:
|
||||
self._old_handlers.append(self.log_handler)
|
||||
self.log_handler = new_handler
|
||||
|
||||
self._config = config.copy()
|
||||
self.update_loggers()
|
||||
|
||||
def _close_old_handlers(self):
|
||||
while True:
|
||||
with self._log_handler_lock:
|
||||
with self.log_handler_lock:
|
||||
if not self._old_handlers:
|
||||
break
|
||||
handler = self._old_handlers.pop()
|
||||
@@ -130,6 +161,11 @@ class PatroniLogger(Thread):
|
||||
_LOGGER.exception('Failed to close the old log handler %s', handler)
|
||||
|
||||
def run(self):
|
||||
# switch to QueueHandler only when the thread was started
|
||||
with self.log_handler_lock:
|
||||
self._root_logger.addHandler(self._queue_handler)
|
||||
self._root_logger.removeHandler(self._proxy_handler)
|
||||
|
||||
while True:
|
||||
self._close_old_handlers()
|
||||
|
||||
@@ -137,7 +173,7 @@ class PatroniLogger(Thread):
|
||||
if record is None:
|
||||
break
|
||||
|
||||
self._log_handler.handle(record)
|
||||
self.log_handler.handle(record)
|
||||
self._queue_handler.queue.task_done()
|
||||
|
||||
def shutdown(self):
|
||||
|
||||
+116
-92
@@ -17,9 +17,9 @@ from patroni.postgresql.misc import parse_history, postgres_major_version_to_int
|
||||
from patroni.postgresql.postmaster import PostmasterProcess
|
||||
from patroni.postgresql.slots import SlotsHandler
|
||||
from patroni.exceptions import PostgresConnectionException
|
||||
from patroni.utils import Retry, RetryFailedError, polling_loop
|
||||
from patroni.dcs import slot_name_from_member_name, RemoteMember
|
||||
from patroni.utils import Retry, RetryFailedError, polling_loop, data_directory_is_empty, parse_int
|
||||
from threading import current_thread, Lock
|
||||
from psutil import TimeoutExpired
|
||||
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
@@ -38,15 +38,6 @@ STATE_UNKNOWN = 'unknown'
|
||||
|
||||
STOP_POLLING_INTERVAL = 1
|
||||
|
||||
cluster_info_query = ("SELECT CASE WHEN pg_catalog.pg_is_in_recovery() THEN 0 "
|
||||
"ELSE ('x' || pg_catalog.substr(pg_catalog.pg_{0}file_name("
|
||||
"pg_catalog.pg_current_{0}_{1}()), 1, 8))::bit(32)::int END, "
|
||||
"CASE WHEN pg_catalog.pg_is_in_recovery() THEN GREATEST("
|
||||
" pg_catalog.pg_{0}_{1}_diff(COALESCE("
|
||||
"pg_catalog.pg_last_{0}_receive_{1}(), '0/0'), '0/0')::bigint,"
|
||||
" pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_last_{0}_replay_{1}(), '0/0')::bigint)"
|
||||
"ELSE pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_current_{0}_{1}(), '0/0')::bigint END")
|
||||
|
||||
|
||||
@contextmanager
|
||||
def null_context():
|
||||
@@ -61,6 +52,7 @@ class Postgresql(object):
|
||||
self._data_dir = config['data_dir']
|
||||
self._database = config.get('database', 'postgres')
|
||||
self._version_file = os.path.join(self._data_dir, 'PG_VERSION')
|
||||
self._pg_control = os.path.join(self._data_dir, 'global', 'pg_control')
|
||||
self._major_version = self.get_major_version()
|
||||
|
||||
self._state_lock = Lock()
|
||||
@@ -69,6 +61,7 @@ class Postgresql(object):
|
||||
self._pending_restart = False
|
||||
self._connection = Connection()
|
||||
self.config = ConfigHandler(self, config)
|
||||
self.config.check_directories()
|
||||
|
||||
self._bin_dir = config.get('bin_dir') or ''
|
||||
self.bootstrap = Bootstrap(self)
|
||||
@@ -77,7 +70,6 @@ class Postgresql(object):
|
||||
|
||||
self.slots_handler = SlotsHandler(self)
|
||||
|
||||
self._pgpass = config.get('pgpass') or os.path.join(os.path.expanduser('~'), 'pgpass')
|
||||
self._callback_executor = CallbackExecutor()
|
||||
self.__cb_called = False
|
||||
self.__cb_pending = None
|
||||
@@ -141,6 +133,20 @@ class Postgresql(object):
|
||||
def lsn_name(self):
|
||||
return 'lsn' if self._major_version >= 100000 else 'location'
|
||||
|
||||
@property
|
||||
def cluster_info_query(self):
|
||||
pg_control_timeline = 'timeline_id FROM pg_catalog.pg_control_checkpoint()' \
|
||||
if self._major_version >= 90600 and self.role == 'standby_leader' else '0'
|
||||
return ("SELECT CASE WHEN pg_catalog.pg_is_in_recovery() THEN 0 "
|
||||
"ELSE ('x' || pg_catalog.substr(pg_catalog.pg_{0}file_name("
|
||||
"pg_catalog.pg_current_{0}_{1}()), 1, 8))::bit(32)::int END, "
|
||||
"CASE WHEN pg_catalog.pg_is_in_recovery() THEN GREATEST("
|
||||
" pg_catalog.pg_{0}_{1}_diff(COALESCE("
|
||||
"pg_catalog.pg_last_{0}_receive_{1}(), '0/0'), '0/0')::bigint,"
|
||||
" pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_last_{0}_replay_{1}(), '0/0')::bigint)"
|
||||
"ELSE pg_catalog.pg_{0}_{1}_diff(pg_catalog.pg_current_{0}_{1}(), '0/0')::bigint "
|
||||
"END, {2}").format(self.wal_name, self.lsn_name, pg_control_timeline)
|
||||
|
||||
def _version_file_exists(self):
|
||||
return not self.data_directory_empty() and os.path.isfile(self._version_file)
|
||||
|
||||
@@ -188,8 +194,8 @@ class Postgresql(object):
|
||||
3: STATE_UNKNOWN}
|
||||
return return_codes.get(ret, STATE_UNKNOWN)
|
||||
|
||||
def reload_config(self, config):
|
||||
self.config.reload_config(config)
|
||||
def reload_config(self, config, sighup=False):
|
||||
self.config.reload_config(config, sighup)
|
||||
self._is_leader_retry.deadline = self.retry.deadline = config['retry_timeout']/2.0
|
||||
|
||||
@property
|
||||
@@ -207,7 +213,7 @@ class Postgresql(object):
|
||||
return self._sysid
|
||||
|
||||
def get_postgres_role_from_data_directory(self):
|
||||
if self.data_directory_empty():
|
||||
if self.data_directory_empty() or not self.controldata():
|
||||
return 'uninitialized'
|
||||
elif self.config.recovery_conf_exists():
|
||||
return 'replica'
|
||||
@@ -255,21 +261,13 @@ class Postgresql(object):
|
||||
except RetryFailedError as e:
|
||||
raise PostgresConnectionException(str(e))
|
||||
|
||||
def pg_control_exists(self):
|
||||
return os.path.isfile(self._pg_control)
|
||||
|
||||
def data_directory_empty(self):
|
||||
return not os.path.exists(self._data_dir) or os.listdir(self._data_dir) == []
|
||||
|
||||
def write_pgpass(self, record):
|
||||
if 'user' not in record or 'password' not in record:
|
||||
return os.environ.copy()
|
||||
|
||||
with open(self._pgpass, 'w') as f:
|
||||
if os.name != 'nt':
|
||||
os.fchmod(f.fileno(), 0o600)
|
||||
f.write('{host}:{port}:*:{user}:{password}\n'.format(**record))
|
||||
|
||||
env = os.environ.copy()
|
||||
env['PGPASSFILE'] = self._pgpass
|
||||
return env
|
||||
if self.pg_control_exists():
|
||||
return False
|
||||
return data_directory_is_empty(self._data_dir)
|
||||
|
||||
def replica_method_options(self, method):
|
||||
return deepcopy(self.config.get(method, {}))
|
||||
@@ -290,10 +288,9 @@ class Postgresql(object):
|
||||
|
||||
def _cluster_info_state_get(self, name):
|
||||
if not self._cluster_info_state:
|
||||
stmt = cluster_info_query.format(self.wal_name, self.lsn_name)
|
||||
try:
|
||||
result = self._is_leader_retry(self._query, stmt).fetchone()
|
||||
self._cluster_info_state = dict(zip(['timeline', 'wal_position'], result))
|
||||
result = self._is_leader_retry(self._query, self.cluster_info_query).fetchone()
|
||||
self._cluster_info_state = dict(zip(['timeline', 'wal_position', 'pg_control_timeline'], result))
|
||||
except RetryFailedError as e: # SELECT failed two times
|
||||
self._cluster_info_state = {'error': str(e)}
|
||||
if not self.is_starting() and self.pg_isready() == STATE_REJECT:
|
||||
@@ -307,6 +304,12 @@ class Postgresql(object):
|
||||
def is_leader(self):
|
||||
return bool(self._cluster_info_state_get('timeline'))
|
||||
|
||||
def pg_control_timeline(self):
|
||||
try:
|
||||
return int(self.controldata().get("Latest checkpoint's TimeLineID"))
|
||||
except (TypeError, ValueError):
|
||||
logger.exception('Failed to parse timeline from pg_controldata output')
|
||||
|
||||
def is_running(self):
|
||||
"""Returns PostmasterProcess if one is running on the data directory or None. If most recently seen process
|
||||
is running updates the cached process based on pid file."""
|
||||
@@ -411,7 +414,12 @@ class Postgresql(object):
|
||||
self.set_state('starting')
|
||||
self._pending_restart = False
|
||||
|
||||
configuration = self.config.effective_configuration
|
||||
try:
|
||||
configuration = self.config.effective_configuration
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
self.config.check_directories()
|
||||
self.config.write_postgresql_conf(configuration)
|
||||
self.config.resolve_connection_addresses()
|
||||
self.config.replace_pg_hba()
|
||||
@@ -455,11 +463,13 @@ class Postgresql(object):
|
||||
else:
|
||||
return None
|
||||
|
||||
def checkpoint(self, connect_kwargs=None):
|
||||
def checkpoint(self, connect_kwargs=None, timeout=None):
|
||||
check_not_is_in_recovery = connect_kwargs is not None
|
||||
connect_kwargs = connect_kwargs or self.config.local_connect_kwargs
|
||||
for p in ['connect_timeout', 'options']:
|
||||
connect_kwargs.pop(p, None)
|
||||
if timeout:
|
||||
connect_kwargs['connect_timeout'] = timeout
|
||||
try:
|
||||
with get_connection_cursor(**connect_kwargs) as cur:
|
||||
cur.execute("SET statement_timeout = 0")
|
||||
@@ -472,7 +482,7 @@ class Postgresql(object):
|
||||
logger.exception('Exception during CHECKPOINT')
|
||||
return 'not accessible or not healty'
|
||||
|
||||
def stop(self, mode='fast', block_callbacks=False, checkpoint=None, on_safepoint=None):
|
||||
def stop(self, mode='fast', block_callbacks=False, checkpoint=None, on_safepoint=None, stop_timeout=None):
|
||||
"""Stop PostgreSQL
|
||||
|
||||
Supports a callback when a safepoint is reached. A safepoint is when no user backend can return a successful
|
||||
@@ -484,7 +494,7 @@ class Postgresql(object):
|
||||
if checkpoint is None:
|
||||
checkpoint = False if mode == 'immediate' else True
|
||||
|
||||
success, pg_signaled = self._do_stop(mode, block_callbacks, checkpoint, on_safepoint)
|
||||
success, pg_signaled = self._do_stop(mode, block_callbacks, checkpoint, on_safepoint, stop_timeout)
|
||||
if success:
|
||||
# block_callbacks is used during restart to avoid
|
||||
# running start/stop callbacks in addition to restart ones
|
||||
@@ -497,7 +507,7 @@ class Postgresql(object):
|
||||
self.set_state('stop failed')
|
||||
return success
|
||||
|
||||
def _do_stop(self, mode, block_callbacks, checkpoint, on_safepoint):
|
||||
def _do_stop(self, mode, block_callbacks, checkpoint, on_safepoint, stop_timeout):
|
||||
postmaster = self.is_running()
|
||||
if not postmaster:
|
||||
if on_safepoint:
|
||||
@@ -505,13 +515,13 @@ class Postgresql(object):
|
||||
return True, False
|
||||
|
||||
if checkpoint and not self.is_starting():
|
||||
self.checkpoint()
|
||||
self.checkpoint(timeout=stop_timeout)
|
||||
|
||||
if not block_callbacks:
|
||||
self.set_state('stopping')
|
||||
|
||||
# Send signal to postmaster to stop
|
||||
success = postmaster.signal_stop(mode)
|
||||
success = postmaster.signal_stop(mode, self.pgcommand('pg_ctl'))
|
||||
if success is not None:
|
||||
if success and on_safepoint:
|
||||
on_safepoint()
|
||||
@@ -524,15 +534,32 @@ class Postgresql(object):
|
||||
postmaster.wait_for_user_backends_to_close()
|
||||
on_safepoint()
|
||||
|
||||
postmaster.wait()
|
||||
try:
|
||||
postmaster.wait(timeout=stop_timeout)
|
||||
except TimeoutExpired:
|
||||
logger.warning("Timeout during postmaster stop, aborting Postgres.")
|
||||
if not self.terminate_postmaster(postmaster, mode, stop_timeout):
|
||||
postmaster.wait()
|
||||
|
||||
return True, True
|
||||
|
||||
@staticmethod
|
||||
def terminate_starting_postmaster(postmaster):
|
||||
def terminate_postmaster(self, postmaster, mode, stop_timeout):
|
||||
if mode in ['fast', 'smart']:
|
||||
try:
|
||||
success = postmaster.signal_stop('immediate', self.pgcommand('pg_ctl'))
|
||||
if success:
|
||||
return True
|
||||
postmaster.wait(timeout=stop_timeout)
|
||||
return True
|
||||
except TimeoutExpired:
|
||||
pass
|
||||
logger.warning("Sending SIGKILL to Postmaster and its children")
|
||||
return postmaster.signal_kill()
|
||||
|
||||
def terminate_starting_postmaster(self, postmaster):
|
||||
"""Terminates a postmaster that has not yet opened ports or possibly even written a pid file. Blocks
|
||||
until the process goes away."""
|
||||
postmaster.signal_stop('immediate')
|
||||
postmaster.signal_stop('immediate', self.pgcommand('pg_ctl'))
|
||||
postmaster.wait()
|
||||
|
||||
def _wait_for_connection_close(self, postmaster):
|
||||
@@ -566,10 +593,12 @@ class Postgresql(object):
|
||||
if ready == STATE_REJECT:
|
||||
return False
|
||||
elif ready == STATE_NO_RESPONSE:
|
||||
self.set_state('start failed')
|
||||
self.slots_handler.schedule(False) # TODO: can remove this?
|
||||
self.config.save_configuration_files(True) # TODO: maybe remove this?
|
||||
return True
|
||||
ret = not self.is_running()
|
||||
if ret:
|
||||
self.set_state('start failed')
|
||||
self.slots_handler.schedule(False) # TODO: can remove this?
|
||||
self.config.save_configuration_files(True) # TODO: maybe remove this?
|
||||
return ret
|
||||
else:
|
||||
if ready != STATE_RUNNING:
|
||||
# Bad configuration or unexpected OS error. No idea of PostgreSQL status.
|
||||
@@ -630,13 +659,12 @@ class Postgresql(object):
|
||||
# Don't try to call pg_controldata during backup restore
|
||||
if self._version_file_exists() and self.state != 'creating replica':
|
||||
try:
|
||||
env = {'LANG': 'C', 'LC_ALL': 'C', 'PATH': os.getenv('PATH')}
|
||||
if os.getenv('SYSTEMROOT') is not None:
|
||||
env['SYSTEMROOT'] = os.getenv('SYSTEMROOT')
|
||||
env = os.environ.copy()
|
||||
env.update(LANG='C', LC_ALL='C')
|
||||
data = subprocess.check_output([self.pgcommand('pg_controldata'), self._data_dir], env=env)
|
||||
if data:
|
||||
data = data.decode('utf-8').splitlines()
|
||||
# pg_controldata output depends on major verion. Some of parameters are prefixed by 'Current '
|
||||
# pg_controldata output depends on major version. Some of parameters are prefixed by 'Current '
|
||||
result = {l.split(':')[0].replace('Current ', '', 1): l.split(':', 1)[1].strip() for l in data
|
||||
if l and ':' in l}
|
||||
except subprocess.CalledProcessError:
|
||||
@@ -644,11 +672,11 @@ class Postgresql(object):
|
||||
return result
|
||||
|
||||
@contextmanager
|
||||
def get_replication_connection_cursor(self, host='localhost', port=5432, database=None, **kwargs):
|
||||
replication = self.config.replication
|
||||
with get_connection_cursor(host=host, port=int(port), database=database or self._database, replication=1,
|
||||
user=replication['username'], password=replication.get('password'),
|
||||
connect_timeout=3, options='-c statement_timeout=2000') as cur:
|
||||
def get_replication_connection_cursor(self, host='localhost', port=5432, **kwargs):
|
||||
conn_kwargs = self.config.replication.copy()
|
||||
conn_kwargs.update(host=host, port=int(port) if port else None, user=conn_kwargs.pop('username'),
|
||||
connect_timeout=3, replication=1, options='-c statement_timeout=2000')
|
||||
with get_connection_cursor(**conn_kwargs) as cur:
|
||||
yield cur
|
||||
|
||||
def get_local_timeline_lsn_from_replication_connection(self):
|
||||
@@ -687,43 +715,25 @@ class Postgresql(object):
|
||||
except Exception:
|
||||
logger.exception('Failed to read and parse %s', (history_path,))
|
||||
|
||||
def follow(self, member, role='replica', timeout=None):
|
||||
is_remote_master = isinstance(member, RemoteMember)
|
||||
no_replication_slot = is_remote_master and member.no_replication_slot
|
||||
restore_command = is_remote_master and member.restore_command
|
||||
min_apply_delay = is_remote_master and member.recovery_min_apply_delay
|
||||
archive_cleanup = is_remote_master and member.archive_cleanup_command
|
||||
|
||||
primary_conninfo = self.config.primary_conninfo(member)
|
||||
change_role = self.cb_called and (self.role in ('master', 'demoted') or
|
||||
not {'standby_leader', 'replica'} - {self.role, role})
|
||||
|
||||
recovery_params = self.config.get('recovery_conf', {}).copy()
|
||||
recovery_params.update({'standby_mode': 'on', 'recovery_target_timeline': 'latest'})
|
||||
if primary_conninfo:
|
||||
recovery_params['primary_conninfo'] = primary_conninfo
|
||||
if self.slots_handler.use_slots and not no_replication_slot:
|
||||
required_name = is_remote_master and member.data.get('primary_slot_name')
|
||||
name = required_name or slot_name_from_member_name(self.name)
|
||||
recovery_params['primary_slot_name'] = name
|
||||
if restore_command:
|
||||
recovery_params['restore_command'] = restore_command
|
||||
if min_apply_delay:
|
||||
recovery_params['recovery_min_apply_delay'] = min_apply_delay
|
||||
if archive_cleanup:
|
||||
recovery_params['archive_cleanup_command'] = archive_cleanup
|
||||
|
||||
def follow(self, member, role='replica', timeout=None, do_reload=False):
|
||||
recovery_params = self.config.build_recovery_params(member)
|
||||
self.config.write_recovery_conf(recovery_params)
|
||||
|
||||
# When we demoting the master or standby_leader to replica or promoting replica to a standby_leader
|
||||
# and we know for sure that postgres was already running before, we will only execute on_role_change
|
||||
# callback and prevent execution of on_restart/on_start callback.
|
||||
# If the role remains the same (replica or standby_leader), we will execute on_start or on_restart
|
||||
change_role = self.cb_called and (self.role in ('master', 'demoted') or
|
||||
not {'standby_leader', 'replica'} - {self.role, role})
|
||||
if change_role:
|
||||
self.__cb_pending = ACTION_NOOP
|
||||
|
||||
if self.is_running():
|
||||
self.restart(block_callbacks=change_role, role=role)
|
||||
if do_reload:
|
||||
self.config.write_postgresql_conf()
|
||||
self.reload()
|
||||
else:
|
||||
self.restart(block_callbacks=change_role, role=role)
|
||||
else:
|
||||
self.start(timeout=timeout, block_callbacks=change_role, role=role)
|
||||
|
||||
@@ -755,17 +765,22 @@ class Postgresql(object):
|
||||
# This method could be called from different threads (simultaneously with some other `_query` calls).
|
||||
# If it is called not from main thread we will create a new cursor to execute statement.
|
||||
if current_thread().ident == self.__thread_ident:
|
||||
return self._cluster_info_state_get('timeline'), self._cluster_info_state_get('wal_position')
|
||||
return (self._cluster_info_state_get('timeline'),
|
||||
self._cluster_info_state_get('wal_position'),
|
||||
self._cluster_info_state_get('pg_control_timeline'))
|
||||
|
||||
with self.connection().cursor() as cursor:
|
||||
cursor.execute(cluster_info_query.format(self.wal_name, self.lsn_name))
|
||||
return cursor.fetchone()[:2]
|
||||
cursor.execute(self.cluster_info_query)
|
||||
return cursor.fetchone()[:3]
|
||||
|
||||
def postmaster_start_time(self):
|
||||
try:
|
||||
cursor = self.query("SELECT pg_catalog.to_char(pg_catalog.pg_postmaster_start_time(),"
|
||||
" 'YYYY-MM-DD HH24:MI:SS.MS TZ')")
|
||||
return cursor.fetchone()[0]
|
||||
query = "SELECT pg_catalog.to_char(pg_catalog.pg_postmaster_start_time(), 'YYYY-MM-DD HH24:MI:SS.MS TZ')"
|
||||
if current_thread().ident == self.__thread_ident:
|
||||
return self.query(query).fetchone()[0]
|
||||
with self.connection().cursor() as cursor:
|
||||
cursor.execute(query)
|
||||
return cursor.fetchone()[0]
|
||||
except psycopg2.Error:
|
||||
return None
|
||||
|
||||
@@ -806,6 +821,15 @@ class Postgresql(object):
|
||||
pg_wal_realpath = os.path.realpath(pg_wal_path)
|
||||
logger.info('Removing WAL directory: %s', pg_wal_realpath)
|
||||
shutil.rmtree(pg_wal_realpath)
|
||||
# Remove user defined tablespace directory
|
||||
pg_tblsp_dir = os.path.join(self._data_dir, 'pg_tblspc')
|
||||
if os.path.exists(pg_tblsp_dir):
|
||||
for tsdn in os.listdir(pg_tblsp_dir):
|
||||
pg_tsp_path = os.path.join(pg_tblsp_dir, tsdn)
|
||||
if parse_int(tsdn) and os.path.islink(pg_tsp_path):
|
||||
pg_tsp_rpath = os.path.realpath(pg_tsp_path)
|
||||
logger.info('Removing user defined tablespace directory: %s', pg_tsp_rpath)
|
||||
shutil.rmtree(pg_tsp_rpath, ignore_errors=True)
|
||||
|
||||
shutil.rmtree(self._data_dir)
|
||||
except (IOError, OSError):
|
||||
@@ -834,12 +858,12 @@ class Postgresql(object):
|
||||
if state != 'streaming' or not member or member.tags.get('nosync', False):
|
||||
continue
|
||||
if sync_state == 'sync':
|
||||
return app_name, True
|
||||
return member.name, True
|
||||
if sync_state == 'potential' and app_name == current:
|
||||
# Prefer current even if not the best one any more to avoid indecisivness and spurious swaps.
|
||||
return current, False
|
||||
return cluster.sync.sync_standby, False
|
||||
if sync_state in ('async', 'potential'):
|
||||
candidates.append(app_name)
|
||||
candidates.append(member.name)
|
||||
|
||||
if candidates:
|
||||
return candidates[0], False
|
||||
|
||||
@@ -5,9 +5,8 @@ import tempfile
|
||||
import time
|
||||
|
||||
from patroni.dcs import RemoteMember
|
||||
from patroni.utils import deep_compare, uri
|
||||
from patroni.utils import deep_compare
|
||||
from six import string_types
|
||||
from six.moves.urllib.parse import quote_plus
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -123,28 +122,14 @@ class Bootstrap(object):
|
||||
cmd = config.get('post_bootstrap') or config.get('post_init')
|
||||
if cmd:
|
||||
r = self._postgresql.config.local_connect_kwargs
|
||||
|
||||
if 'host' in r:
|
||||
# '/tmp' => '%2Ftmp' for unix socket path
|
||||
host = quote_plus(r['host']) if r['host'].startswith('/') else r['host']
|
||||
else:
|
||||
host = ''
|
||||
|
||||
connstring = self._postgresql.config.format_dsn(r, True)
|
||||
if 'host' not in r:
|
||||
# https://www.postgresql.org/docs/current/static/libpq-pgpass.html
|
||||
# A host name of localhost matches both TCP (host name localhost) and Unix domain socket
|
||||
# (pghost empty or the default socket directory) connections coming from the local machine.
|
||||
r['host'] = 'localhost' # set it to localhost to write into pgpass
|
||||
|
||||
if 'user' in r:
|
||||
user = r['user']
|
||||
else:
|
||||
user = ''
|
||||
if 'password' in r:
|
||||
import getpass
|
||||
r.setdefault('user', os.environ.get('PGUSER', getpass.getuser()))
|
||||
|
||||
connstring = uri('postgres', (host, r['port']), r['database'], user)
|
||||
env = self._postgresql.write_pgpass(r) if 'password' in r else None
|
||||
env = self._postgresql.config.write_pgpass(r) if 'password' in r else None
|
||||
|
||||
try:
|
||||
ret = self._postgresql.cancellable.call(shlex.split(cmd) + [connstring], env=env)
|
||||
@@ -176,9 +161,9 @@ class Bootstrap(object):
|
||||
|
||||
if clone_member and clone_member.conn_url:
|
||||
r = clone_member.conn_kwargs(self._postgresql.config.replication)
|
||||
connstring = uri('postgres', (r['host'], r['port']), r['database'], r['user'])
|
||||
# add the credentials to connect to the replica origin to pgpass.
|
||||
env = self._postgresql.write_pgpass(r)
|
||||
env = self._postgresql.config.write_pgpass(r)
|
||||
connstring = self._postgresql.config.format_dsn(r, True)
|
||||
else:
|
||||
connstring = ''
|
||||
env = os.environ.copy()
|
||||
@@ -325,7 +310,15 @@ BEGIN
|
||||
CREATE ROLE "{0}" WITH {1};
|
||||
END IF;
|
||||
END;$$""".format(name, ' '.join(options))
|
||||
self._postgresql.query(sql, *params)
|
||||
self._postgresql.query('SET log_statement TO none')
|
||||
self._postgresql.query('SET log_min_duration_statement TO -1')
|
||||
self._postgresql.query("SET log_min_error_statement TO 'log'")
|
||||
try:
|
||||
self._postgresql.query(sql, *params)
|
||||
finally:
|
||||
self._postgresql.query('RESET log_min_error_statement')
|
||||
self._postgresql.query('RESET log_min_duration_statement')
|
||||
self._postgresql.query('RESET log_statement')
|
||||
|
||||
def post_bootstrap(self, config, task):
|
||||
try:
|
||||
|
||||
@@ -1,40 +1,36 @@
|
||||
import logging
|
||||
import subprocess
|
||||
from threading import Event, Lock, Thread
|
||||
|
||||
from patroni.postgresql.cancellable import CancellableExecutor
|
||||
from threading import Condition, Thread
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class CallbackExecutor(Thread):
|
||||
class CallbackExecutor(CancellableExecutor, Thread):
|
||||
|
||||
def __init__(self):
|
||||
super(CallbackExecutor, self).__init__()
|
||||
CancellableExecutor.__init__(self)
|
||||
Thread.__init__(self)
|
||||
self.daemon = True
|
||||
self._lock = Lock()
|
||||
self._cmd = None
|
||||
self._process = None
|
||||
self._callback_event = Event()
|
||||
self._condition = Condition()
|
||||
self.start()
|
||||
|
||||
def call(self, cmd):
|
||||
with self._lock:
|
||||
if self._process and self._process.poll() is None:
|
||||
try:
|
||||
self._process.kill()
|
||||
logger.warning('Killed the old callback process because it was still running: %s', self._cmd)
|
||||
except OSError:
|
||||
logger.exception('Failed to kill the old callback')
|
||||
self._cmd = cmd
|
||||
self._callback_event.set()
|
||||
self._kill_process()
|
||||
with self._condition:
|
||||
self._cmd = cmd
|
||||
self._condition.notify()
|
||||
|
||||
def run(self):
|
||||
while True:
|
||||
self._callback_event.wait()
|
||||
self._callback_event.clear()
|
||||
with self._condition:
|
||||
if self._cmd is None:
|
||||
self._condition.wait()
|
||||
cmd, self._cmd = self._cmd, None
|
||||
|
||||
with self._lock:
|
||||
try:
|
||||
self._process = subprocess.Popen(self._cmd, close_fds=True)
|
||||
except Exception:
|
||||
logger.exception('Failed to execute %s', self._cmd)
|
||||
if not self._start_process(cmd, close_fds=True):
|
||||
continue
|
||||
self._process.wait()
|
||||
self._kill_children()
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
import logging
|
||||
import os
|
||||
import psutil
|
||||
import subprocess
|
||||
|
||||
from patroni.exceptions import PostgresException
|
||||
@@ -10,13 +11,66 @@ from threading import Lock
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class CancellableSubprocess(object):
|
||||
class CancellableExecutor(object):
|
||||
|
||||
def __init__(self):
|
||||
self._is_cancelled = False
|
||||
self._process = None
|
||||
self._process_cmd = None
|
||||
self._process_children = []
|
||||
self._lock = Lock()
|
||||
|
||||
def _start_process(self, cmd, *args, **kwargs):
|
||||
"""This method must be executed only when the `_lock` is acquired"""
|
||||
|
||||
try:
|
||||
self._process_children = []
|
||||
self._process_cmd = cmd
|
||||
self._process = psutil.Popen(cmd, *args, **kwargs)
|
||||
except Exception:
|
||||
return logger.exception('Failed to execute %s', cmd)
|
||||
return True
|
||||
|
||||
def _kill_process(self):
|
||||
with self._lock:
|
||||
if self._process is not None and self._process.is_running() and not self._process_children:
|
||||
try:
|
||||
self._process.suspend() # Suspend the process before getting list of childrens
|
||||
except psutil.Error as e:
|
||||
logger.info('Failed to suspend the process: %s', e.msg)
|
||||
|
||||
try:
|
||||
self._process_children = self._process.children(recursive=True)
|
||||
except psutil.Error:
|
||||
pass
|
||||
|
||||
try:
|
||||
self._process.kill()
|
||||
logger.warning('Killed %s because it was still running', self._process_cmd)
|
||||
except psutil.NoSuchProcess:
|
||||
pass
|
||||
except psutil.AccessDenied as e:
|
||||
logger.warning('Failed to kill the process: %s', e.msg)
|
||||
|
||||
def _kill_children(self):
|
||||
waitlist = []
|
||||
with self._lock:
|
||||
for child in self._process_children:
|
||||
try:
|
||||
child.kill()
|
||||
except psutil.NoSuchProcess:
|
||||
continue
|
||||
except psutil.AccessDenied as e:
|
||||
logger.info('Failed to kill child process: %s', e.msg)
|
||||
waitlist.append(child)
|
||||
psutil.wait_procs(waitlist)
|
||||
|
||||
|
||||
class CancellableSubprocess(CancellableExecutor):
|
||||
|
||||
def __init__(self):
|
||||
super(CancellableSubprocess, self).__init__()
|
||||
self._is_cancelled = False
|
||||
|
||||
def call(self, *args, **kwargs):
|
||||
for s in ('stdin', 'stdout', 'stderr'):
|
||||
kwargs.pop(s, None)
|
||||
@@ -38,17 +92,18 @@ class CancellableSubprocess(object):
|
||||
raise PostgresException('cancelled')
|
||||
|
||||
self._is_cancelled = False
|
||||
self._process = subprocess.Popen(*args, **kwargs)
|
||||
started = self._start_process(*args, **kwargs)
|
||||
|
||||
if communicate_input:
|
||||
if input_data:
|
||||
self._process.communicate(input_data)
|
||||
self._process.stdin.close()
|
||||
|
||||
return self._process.wait()
|
||||
if started:
|
||||
if communicate_input:
|
||||
if input_data:
|
||||
self._process.communicate(input_data)
|
||||
self._process.stdin.close()
|
||||
return self._process.wait()
|
||||
finally:
|
||||
with self._lock:
|
||||
self._process = None
|
||||
self._kill_children()
|
||||
|
||||
def reset_is_cancelled(self):
|
||||
with self._lock:
|
||||
@@ -62,15 +117,13 @@ class CancellableSubprocess(object):
|
||||
def cancel(self):
|
||||
with self._lock:
|
||||
self._is_cancelled = True
|
||||
if self._process is None or self._process.returncode is not None:
|
||||
if self._process is None or not self._process.is_running():
|
||||
return
|
||||
self._process.terminate()
|
||||
|
||||
for _ in polling_loop(10):
|
||||
with self._lock:
|
||||
if self._process is None or self._process.returncode is not None:
|
||||
if self._process is None or not self._process.is_running():
|
||||
return
|
||||
|
||||
with self._lock:
|
||||
if self._process is not None and self._process.returncode is None:
|
||||
self._process.kill()
|
||||
self._kill_process()
|
||||
|
||||
+441
-140
@@ -4,11 +4,15 @@ import re
|
||||
import shutil
|
||||
import socket
|
||||
import stat
|
||||
import time
|
||||
|
||||
from requests.structures import CaseInsensitiveDict
|
||||
from patroni.exceptions import PatroniException
|
||||
from six.moves.urllib_parse import urlparse, parse_qsl, unquote
|
||||
from urllib3.response import HTTPHeaderDict
|
||||
|
||||
from ..utils import compare_values, parse_bool, parse_int, split_host_port, uri
|
||||
from ..dcs import slot_name_from_member_name, RemoteMember
|
||||
from ..utils import compare_values, parse_bool, parse_int, split_host_port, uri, \
|
||||
validate_directory, is_subpath
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -57,14 +61,15 @@ def conninfo_uri_parse(dsn):
|
||||
return ret
|
||||
|
||||
|
||||
def read_param_value(value, is_quoted=False):
|
||||
def read_param_value(value):
|
||||
length = len(value)
|
||||
ret = ''
|
||||
i = 0
|
||||
is_quoted = value[0] == "'"
|
||||
i = int(is_quoted)
|
||||
while i < length:
|
||||
if is_quoted:
|
||||
if value[i] == "'":
|
||||
return ret, i
|
||||
return ret, i + 1
|
||||
elif value[i].isspace():
|
||||
break
|
||||
if value[i] == '\\':
|
||||
@@ -95,12 +100,10 @@ def conninfo_parse(dsn):
|
||||
if i >= length:
|
||||
return
|
||||
|
||||
is_quoted = dsn[i] == "'"
|
||||
i += int(is_quoted)
|
||||
value, end = read_param_value(dsn[i:], is_quoted)
|
||||
value, end = read_param_value(dsn[i:])
|
||||
if value is None:
|
||||
return
|
||||
i += end + int(is_quoted)
|
||||
i += end
|
||||
ret[param] = value
|
||||
return ret
|
||||
|
||||
@@ -117,7 +120,7 @@ def parse_dsn(value):
|
||||
>>> r == {'application_name': 'mya/pp', 'host': ',/host2', 'sslmode': 'require',\
|
||||
'password': 'pass', 'port': '/123', 'user': 'u/se'}
|
||||
True
|
||||
>>> r = parse_dsn(" host = 'host' dbname = db\\ name requiressl=1 ")
|
||||
>>> r = parse_dsn(" host = 'host' dbname = db\\\\ name requiressl=1 ")
|
||||
>>> r == {'host': 'host', 'sslmode': 'require'}
|
||||
True
|
||||
>>> parse_dsn('requiressl = 0\\\\') == {'sslmode': 'prefer'}
|
||||
@@ -147,6 +150,70 @@ def parse_dsn(value):
|
||||
return ret
|
||||
|
||||
|
||||
def strip_comment(value):
|
||||
i = value.find('#')
|
||||
if i > -1:
|
||||
value = value[:i].strip()
|
||||
return value
|
||||
|
||||
|
||||
def read_recovery_param_value(value):
|
||||
"""
|
||||
>>> read_recovery_param_value('') is None
|
||||
True
|
||||
>>> read_recovery_param_value("'") is None
|
||||
True
|
||||
>>> read_recovery_param_value("''a") is None
|
||||
True
|
||||
>>> read_recovery_param_value('a b') is None
|
||||
True
|
||||
>>> read_recovery_param_value("'''") is None
|
||||
True
|
||||
>>> read_recovery_param_value("'\\\\") is None
|
||||
True
|
||||
>>> read_recovery_param_value("'a' s#") is None
|
||||
True
|
||||
>>> read_recovery_param_value("'\\\\'''' #a")
|
||||
"''"
|
||||
>>> read_recovery_param_value('asd')
|
||||
'asd'
|
||||
"""
|
||||
value = value.strip()
|
||||
length = len(value)
|
||||
if length == 0:
|
||||
return None
|
||||
elif value[0] == "'":
|
||||
if length == 1:
|
||||
return None
|
||||
ret = ''
|
||||
i = 1
|
||||
while i < length:
|
||||
if value[i] == '\\':
|
||||
i += 1
|
||||
if i >= length:
|
||||
return None
|
||||
elif value[i] == "'":
|
||||
i += 1
|
||||
if i >= length:
|
||||
break
|
||||
if value[i] in ('#', ' '):
|
||||
if strip_comment(value[i:]):
|
||||
return None
|
||||
break
|
||||
if value[i] != "'":
|
||||
return None
|
||||
ret += value[i]
|
||||
i += 1
|
||||
else:
|
||||
return None
|
||||
return ret
|
||||
else:
|
||||
value = strip_comment(value)
|
||||
if not value or ' ' in value or '\\' in value:
|
||||
return None
|
||||
return value
|
||||
|
||||
|
||||
def mtime(filename):
|
||||
try:
|
||||
return os.stat(filename).st_mtime
|
||||
@@ -154,6 +221,52 @@ def mtime(filename):
|
||||
return None
|
||||
|
||||
|
||||
class ConfigWriter(object):
|
||||
|
||||
def __init__(self, filename):
|
||||
self._filename = filename
|
||||
self._fd = None
|
||||
|
||||
def __enter__(self):
|
||||
self._fd = open(self._filename, 'w')
|
||||
self.writeline('# Do not edit this file manually!\n# It will be overwritten by Patroni!')
|
||||
return self
|
||||
|
||||
def __exit__(self, exc_type, exc_val, exc_tb):
|
||||
if self._fd:
|
||||
self._fd.close()
|
||||
|
||||
def writeline(self, line):
|
||||
self._fd.write(line)
|
||||
self._fd.write('\n')
|
||||
|
||||
def writelines(self, lines):
|
||||
for line in lines:
|
||||
self.writeline(line)
|
||||
|
||||
@staticmethod
|
||||
def escape(value): # Escape (by doubling) any single quotes or backslashes in given string
|
||||
return re.sub(r'([\'\\])', r'\1\1', str(value))
|
||||
|
||||
def write_param(self, param, value):
|
||||
self.writeline("{0} = '{1}'".format(param, self.escape(value)))
|
||||
|
||||
|
||||
class CaseInsensitiveDict(HTTPHeaderDict):
|
||||
|
||||
def add(self, key, val):
|
||||
self[key] = val
|
||||
|
||||
def __getitem__(self, key):
|
||||
return self._container[key.lower()][1]
|
||||
|
||||
def __repr__(self):
|
||||
return str(dict(self.items()))
|
||||
|
||||
def copy(self):
|
||||
return CaseInsensitiveDict(self._container.values())
|
||||
|
||||
|
||||
class ConfigHandler(object):
|
||||
|
||||
# List of parameters which must be always passed to postmaster as command line options
|
||||
@@ -206,8 +319,6 @@ class ConfigHandler(object):
|
||||
'trigger_file'
|
||||
}
|
||||
|
||||
_CONFIG_WARNING_HEADER = '# Do not edit this file manually!\n# It will be overwritten by Patroni!\n'
|
||||
|
||||
def __init__(self, postgresql, config):
|
||||
self._postgresql = postgresql
|
||||
self._config_dir = os.path.abspath(config.get('config_dir') or postgresql.data_dir)
|
||||
@@ -224,9 +335,15 @@ class ConfigHandler(object):
|
||||
self._standby_signal = os.path.join(postgresql.data_dir, 'standby.signal')
|
||||
self._auto_conf = os.path.join(postgresql.data_dir, 'postgresql.auto.conf')
|
||||
self._auto_conf_mtime = None
|
||||
self._pgpass = os.path.abspath(config.get('pgpass') or os.path.join(os.path.expanduser('~'), 'pgpass'))
|
||||
if os.path.exists(self._pgpass) and not os.path.isfile(self._pgpass):
|
||||
raise PatroniException("'{}' exists and it's not a file, check your `postgresql.pgpass` configuration"
|
||||
.format(self._pgpass))
|
||||
self._passfile = None
|
||||
self._passfile_mtime = None
|
||||
self._synchronous_standby_names = None
|
||||
self._postmaster_ctime = None
|
||||
self._primary_conninfo = None
|
||||
self._current_recovery_params = None
|
||||
self._config = {}
|
||||
self._recovery_params = {}
|
||||
self.reload_config(config)
|
||||
@@ -235,6 +352,21 @@ class ConfigHandler(object):
|
||||
self._server_parameters = self.get_server_parameters(self._config)
|
||||
self._adjust_recovery_parameters()
|
||||
|
||||
def try_to_create_dir(self, d, msg):
|
||||
d = os.path.join(self._postgresql._data_dir, d)
|
||||
if (not is_subpath(self._postgresql._data_dir, d) or not self._postgresql.data_directory_empty()):
|
||||
validate_directory(d, msg)
|
||||
|
||||
def check_directories(self):
|
||||
if "unix_socket_directories" in self._server_parameters:
|
||||
for d in self._server_parameters["unix_socket_directories"].split(","):
|
||||
self.try_to_create_dir(d.strip(), "'{}' is defined in unix_socket_directories, {}")
|
||||
if "stats_temp_directory" in self._server_parameters:
|
||||
self.try_to_create_dir(self._server_parameters["stats_temp_directory"],
|
||||
"'{}' is defined in stats_temp_directory, {}")
|
||||
self.try_to_create_dir(os.path.dirname(self._pgpass),
|
||||
"'{}' is defined in `postgresql.pgpass`, {}")
|
||||
|
||||
@property
|
||||
def _configuration_to_save(self):
|
||||
configuration = [os.path.basename(self._postgresql_conf)]
|
||||
@@ -283,28 +415,26 @@ class ConfigHandler(object):
|
||||
if 'custom_conf' not in self._config and not os.path.exists(self._postgresql_base_conf):
|
||||
os.rename(self._postgresql_conf, self._postgresql_base_conf)
|
||||
|
||||
with open(self._postgresql_conf, 'w') as f:
|
||||
os.chmod(self._postgresql_conf, stat.S_IWRITE | stat.S_IREAD)
|
||||
f.write(self._CONFIG_WARNING_HEADER)
|
||||
f.write("include '{0}'\n\n".format(self._config.get('custom_conf') or self._postgresql_base_conf_name))
|
||||
with ConfigWriter(self._postgresql_conf) as f:
|
||||
include = self._config.get('custom_conf') or self._postgresql_base_conf_name
|
||||
f.writeline("include '{0}'\n".format(ConfigWriter.escape(include)))
|
||||
for name, value in sorted((configuration or self._server_parameters).items()):
|
||||
if (not self._postgresql.bootstrap.running_custom_bootstrap or name != 'hba_file') \
|
||||
and name not in self._RECOVERY_PARAMETERS:
|
||||
f.write("{0} = '{1}'\n".format(name, value))
|
||||
f.write_param(name, value)
|
||||
# when we are doing custom bootstrap we assume that we don't know superuser password
|
||||
# and in order to be able to change it, we are opening trust access from a certain address
|
||||
# therefore we need to make sure that hba_file is not overriden
|
||||
# after changing superuser password we will "revert" all these "changes"
|
||||
if self._postgresql.bootstrap.running_custom_bootstrap or 'hba_file' not in self._server_parameters:
|
||||
f.write("hba_file = '{0}'\n".format(self._pg_hba_conf.replace('\\', '\\\\')))
|
||||
f.write_param('hba_file', self._pg_hba_conf)
|
||||
if 'ident_file' not in self._server_parameters:
|
||||
f.write("ident_file = '{0}'\n".format(self._pg_ident_conf.replace('\\', '\\\\')))
|
||||
f.write_param('ident_file', self._pg_ident_conf)
|
||||
|
||||
if self._postgresql.major_version >= 120000:
|
||||
if self._recovery_params:
|
||||
f.write('\n# recovery.conf\n')
|
||||
for name, value in sorted(self._recovery_params.items()):
|
||||
f.write("{0} = '{1}'\n".format(name, value))
|
||||
f.writeline('\n# recovery.conf')
|
||||
self._write_recovery_params(f, self._recovery_params)
|
||||
|
||||
if not self._postgresql.bootstrap.keep_existing_recovery_conf:
|
||||
self._sanitize_auto_conf()
|
||||
@@ -326,24 +456,21 @@ class ConfigHandler(object):
|
||||
# when we are doing custom bootstrap we assume that we don't know superuser password
|
||||
# and in order to be able to change it, we are opening trust access from a certain address
|
||||
if self._postgresql.bootstrap.running_custom_bootstrap:
|
||||
addresses = {'': 'local'}
|
||||
addresses = {} if os.name == 'nt' else {'': 'local'} # windows doesn't yet support unix-domain sockets
|
||||
if 'host' in self.local_replication_address and not self.local_replication_address['host'].startswith('/'):
|
||||
addresses.update({sa[0] + '/32': 'host' for _, _, _, _, sa in socket.getaddrinfo(
|
||||
self.local_replication_address['host'], self.local_replication_address['port'],
|
||||
0, socket.SOCK_STREAM, socket.IPPROTO_TCP)})
|
||||
|
||||
with open(self._pg_hba_conf, 'w') as f:
|
||||
f.write(self._CONFIG_WARNING_HEADER)
|
||||
with ConfigWriter(self._pg_hba_conf) as f:
|
||||
for address, t in addresses.items():
|
||||
f.write((
|
||||
f.writeline((
|
||||
'{0}\treplication\t{1}\t{3}\ttrust\n'
|
||||
'{0}\tall\t{2}\t{3}\ttrust\n'
|
||||
'{0}\tall\t{2}\t{3}\ttrust'
|
||||
).format(t, self.replication['username'], self._superuser.get('username') or 'all', address))
|
||||
elif not self.hba_file and self._config.get('pg_hba'):
|
||||
with open(self._pg_hba_conf, 'w') as f:
|
||||
f.write(self._CONFIG_WARNING_HEADER)
|
||||
for line in self._config['pg_hba']:
|
||||
f.write('{0}\n'.format(line))
|
||||
with ConfigWriter(self._pg_hba_conf) as f:
|
||||
f.writelines(self._config['pg_hba'])
|
||||
return True
|
||||
|
||||
def replace_pg_ident(self):
|
||||
@@ -355,122 +482,263 @@ class ConfigHandler(object):
|
||||
"""
|
||||
|
||||
if not self._server_parameters.get('ident_file') and self._config.get('pg_ident'):
|
||||
with open(self._pg_ident_conf, 'w') as f:
|
||||
f.write(self._CONFIG_WARNING_HEADER)
|
||||
for line in self._config['pg_ident']:
|
||||
f.write('{0}\n'.format(line))
|
||||
with ConfigWriter(self._pg_ident_conf) as f:
|
||||
f.writelines(self._config['pg_ident'])
|
||||
return True
|
||||
|
||||
def primary_conninfo_params(self, member):
|
||||
name = self._postgresql.name
|
||||
if not (member and member.conn_url) or member.name == name:
|
||||
if not (member and member.conn_url) or member.name == self._postgresql.name:
|
||||
return None
|
||||
ret = member.conn_kwargs(self.replication)
|
||||
ret.update(application_name=name, sslmode='prefer')
|
||||
ret['application_name'] = self._postgresql.name
|
||||
ret.setdefault('sslmode', 'prefer')
|
||||
if self._krbsrvname:
|
||||
ret['krbsrvname'] = self._krbsrvname
|
||||
if 'database' in ret:
|
||||
del ret['database']
|
||||
return ret
|
||||
|
||||
def primary_conninfo(self, member):
|
||||
r = self.primary_conninfo_params(member)
|
||||
if not r:
|
||||
return None
|
||||
keywords = 'user password host port sslmode application_name krbsrvname'.split()
|
||||
return ' '.join('{0}={{{0}}}'.format(kw) for kw in keywords if r.get(kw)).format(**r)
|
||||
def format_dsn(self, params, include_dbname=False):
|
||||
# A list of keywords that can be found in a conninfo string. Follows what is acceptable by libpq
|
||||
keywords = ('dbname', 'user', 'passfile' if params.get('passfile') else 'password', 'host', 'port', 'sslmode',
|
||||
'sslcompression', 'sslcert', 'sslkey', 'sslrootcert', 'sslcrl', 'application_name', 'krbsrvname')
|
||||
if include_dbname:
|
||||
params = params.copy()
|
||||
params['dbname'] = params.get('database') or self._postgresql.database
|
||||
# we are abusing information about the necessity of dbname
|
||||
# dsn should contain passfile or password only if there is no dbname in it (it is used in recovery.conf)
|
||||
skip = {'passfile', 'password'}
|
||||
else:
|
||||
skip = {'dbname'}
|
||||
|
||||
def escape(value):
|
||||
return re.sub(r'([\'\\ ])', r'\\\1', str(value))
|
||||
|
||||
return ' '.join('{0}={1}'.format(kw, escape(params[kw])) for kw in keywords
|
||||
if kw not in skip and params.get(kw) is not None)
|
||||
|
||||
def _write_recovery_params(self, fd, recovery_params):
|
||||
for name, value in sorted(recovery_params.items()):
|
||||
if name == 'primary_conninfo':
|
||||
if 'password' in value and self._postgresql.major_version >= 100000:
|
||||
self.write_pgpass(value)
|
||||
value['passfile'] = self._passfile = self._pgpass
|
||||
self._passfile_mtime = mtime(self._pgpass)
|
||||
value = self.format_dsn(value)
|
||||
fd.write_param(name, value)
|
||||
|
||||
def build_recovery_params(self, member):
|
||||
recovery_params = CaseInsensitiveDict({p: v for p, v in self.get('recovery_conf', {}).items()
|
||||
if not p.lower().startswith('recovery_target') and
|
||||
p.lower() not in ('primary_conninfo', 'primary_slot_name')})
|
||||
recovery_params.update({'standby_mode': 'on', 'recovery_target_timeline': 'latest'})
|
||||
if self._postgresql.major_version >= 120000:
|
||||
# on pg12 we want to protect from following params being set in one of included files
|
||||
# not doing so might result in a standby being paused, promoted or shutted down.
|
||||
recovery_params.update({'recovery_target': '', 'recovery_target_name': '', 'recovery_target_time': '',
|
||||
'recovery_target_xid': '', 'recovery_target_lsn': ''})
|
||||
|
||||
is_remote_master = isinstance(member, RemoteMember)
|
||||
primary_conninfo = self.primary_conninfo_params(member)
|
||||
if primary_conninfo:
|
||||
use_slots = self.get('use_slots', True) and self._postgresql.major_version >= 90400
|
||||
if use_slots and not (is_remote_master and member.no_replication_slot):
|
||||
primary_slot_name = member.primary_slot_name if is_remote_master else self._postgresql.name
|
||||
recovery_params['primary_slot_name'] = slot_name_from_member_name(primary_slot_name)
|
||||
recovery_params['primary_conninfo'] = primary_conninfo
|
||||
|
||||
# standby_cluster config might have different parameters, we want to override them
|
||||
standby_cluster_params = ['restore_command', 'archive_cleanup_command']\
|
||||
+ (['recovery_min_apply_delay'] if is_remote_master else [])
|
||||
recovery_params.update({p: member.data.get(p) for p in standby_cluster_params if member and member.data.get(p)})
|
||||
return recovery_params
|
||||
|
||||
def recovery_conf_exists(self):
|
||||
if self._postgresql.major_version >= 120000:
|
||||
return os.path.exists(self._standby_signal) or os.path.exists(self._recovery_signal)
|
||||
return os.path.exists(self._recovery_conf)
|
||||
|
||||
def _read_primary_conninfo(self):
|
||||
@property
|
||||
def _triggerfile_good_name(self):
|
||||
return 'trigger_file' if self._postgresql.major_version < 120000 else 'promote_trigger_file'
|
||||
|
||||
@property
|
||||
def _triggerfile_wrong_name(self):
|
||||
return 'trigger_file' if self._postgresql.major_version >= 120000 else 'promote_trigger_file'
|
||||
|
||||
@property
|
||||
def _recovery_parameters_to_compare(self):
|
||||
skip_params = {'recovery_target_inclusive', 'recovery_target_action', self._triggerfile_wrong_name}
|
||||
return self._RECOVERY_PARAMETERS - skip_params
|
||||
|
||||
def _read_recovery_params(self):
|
||||
pg_conf_mtime = mtime(self._postgresql_conf)
|
||||
auto_conf_mtime = mtime(self._auto_conf)
|
||||
passfile_mtime = mtime(self._passfile) if self._passfile else False
|
||||
postmaster_ctime = self._postgresql.is_running()
|
||||
if postmaster_ctime:
|
||||
postmaster_ctime = postmaster_ctime.create_time()
|
||||
|
||||
if self._postgresql_conf_mtime == pg_conf_mtime and self._auto_conf_mtime == auto_conf_mtime \
|
||||
and self._postmaster_ctime == postmaster_ctime:
|
||||
and self._passfile_mtime == passfile_mtime and self._postmaster_ctime == postmaster_ctime:
|
||||
return None, False
|
||||
|
||||
try:
|
||||
primary_conninfo = self._postgresql.query('SHOW primary_conninfo').fetchone()[0]
|
||||
values = self._get_pg_settings(self._recovery_parameters_to_compare).values()
|
||||
values = {p[0]: [p[1], p[4] == 'postmaster', p[5]] for p in values}
|
||||
self._postgresql_conf_mtime = pg_conf_mtime
|
||||
self._auto_conf_mtime = auto_conf_mtime
|
||||
self._postmaster_ctime = postmaster_ctime
|
||||
except Exception:
|
||||
primary_conninfo = None
|
||||
return primary_conninfo, True
|
||||
values = None
|
||||
return values, True
|
||||
|
||||
def _read_primary_conninfo_pre_v12(self):
|
||||
def _read_recovery_params_pre_v12(self):
|
||||
recovery_conf_mtime = mtime(self._recovery_conf)
|
||||
if recovery_conf_mtime == self._recovery_conf_mtime:
|
||||
passfile_mtime = mtime(self._passfile) if self._passfile else False
|
||||
if recovery_conf_mtime == self._recovery_conf_mtime and passfile_mtime == self._passfile_mtime:
|
||||
return None, False
|
||||
|
||||
primary_conninfo = ''
|
||||
values = {}
|
||||
with open(self._recovery_conf, 'r') as f:
|
||||
for line in f:
|
||||
line = line.strip()
|
||||
if not line or line.startswith('#'):
|
||||
continue
|
||||
value = None
|
||||
match = PARAMETER_RE.match(line)
|
||||
if match and match.group(1) == 'primary_conninfo':
|
||||
i = match.end()
|
||||
if i < len(line):
|
||||
is_quoted = line[i] == "'"
|
||||
i += int(is_quoted)
|
||||
primary_conninfo, _ = read_param_value(line[i:], is_quoted)
|
||||
if match:
|
||||
value = read_recovery_param_value(line[match.end():])
|
||||
if value is None:
|
||||
return None, True
|
||||
values[match.group(1)] = [value, True]
|
||||
self._recovery_conf_mtime = recovery_conf_mtime
|
||||
return primary_conninfo, True
|
||||
values.setdefault('recovery_min_apply_delay', ['0', True])
|
||||
values['recovery_min_apply_delay'][0] = parse_int(values['recovery_min_apply_delay'][0], 'ms')
|
||||
values.update({param: ['', True] for param in self._recovery_parameters_to_compare if param not in values})
|
||||
return values, True
|
||||
|
||||
def check_recovery_conf(self, member): # Name is confusing. In fact it checks the value of primary_conninfo
|
||||
# TODO: recovery.conf could be stale, would be nice to detect that.
|
||||
if self._postgresql.major_version >= 120000:
|
||||
if not os.path.exists(self._standby_signal):
|
||||
return False
|
||||
def _check_passfile(self, passfile, wanted_primary_conninfo):
|
||||
# If there is a passfile in the primary_conninfo try to figure out that
|
||||
# the passfile contains the line allowing connection to the given node.
|
||||
# We assume that the passfile was created by Patroni and therefore doing
|
||||
# the full match and not covering cases when host, port or user are set to '*'
|
||||
passfile_mtime = mtime(passfile)
|
||||
if passfile_mtime:
|
||||
try:
|
||||
with open(passfile) as f:
|
||||
wanted_line = self._pgpass_line(wanted_primary_conninfo).strip()
|
||||
for raw_line in f:
|
||||
if raw_line.strip() == wanted_line:
|
||||
self._passfile = passfile
|
||||
self._passfile_mtime = passfile_mtime
|
||||
return True
|
||||
except Exception:
|
||||
logger.info('Failed to read %s', passfile)
|
||||
return False
|
||||
|
||||
_read_primary_conninfo = self._read_primary_conninfo
|
||||
else:
|
||||
if not self.recovery_conf_exists():
|
||||
return False
|
||||
|
||||
_read_primary_conninfo = self._read_primary_conninfo_pre_v12
|
||||
|
||||
primary_conninfo, updated = _read_primary_conninfo()
|
||||
# updated indicates that mtime of postgresql.conf, postgresql.auto.conf, or recovery.conf was changed
|
||||
# and the primary_conninfo value was read either from config or from the database connection.
|
||||
if updated:
|
||||
# primary_conninfo is one of:
|
||||
# - None (exception or unparsable config)
|
||||
# - '' (not in config)
|
||||
# - or the actual dsn value
|
||||
self._primary_conninfo = primary_conninfo
|
||||
if primary_conninfo:
|
||||
# We will cache parsed value until the next config change.
|
||||
self._primary_conninfo = parse_dsn(primary_conninfo)
|
||||
# If we failed to parse non-empty connection string this indicates that config if broken.
|
||||
if not self._primary_conninfo:
|
||||
return False
|
||||
elif primary_conninfo is not None:
|
||||
self._primary_conninfo = {}
|
||||
else: # primary_conninfo is None, config is probably broken
|
||||
return False
|
||||
|
||||
wanted_primary_conninfo = self.primary_conninfo_params(member)
|
||||
def _check_primary_conninfo(self, primary_conninfo, wanted_primary_conninfo):
|
||||
# first we will cover corner cases, when we are replicating from somewhere while shouldn't
|
||||
# or there is no primary_conninfo but we should replicate from some specific node.
|
||||
if not wanted_primary_conninfo:
|
||||
return not self._primary_conninfo
|
||||
elif not self._primary_conninfo:
|
||||
return not primary_conninfo
|
||||
elif not primary_conninfo:
|
||||
return False
|
||||
|
||||
return all(self._primary_conninfo.get(p) == str(v) for p, v in wanted_primary_conninfo.items())
|
||||
if 'passfile' in primary_conninfo and 'password' not in primary_conninfo \
|
||||
and 'password' in wanted_primary_conninfo:
|
||||
if self._check_passfile(primary_conninfo['passfile'], wanted_primary_conninfo):
|
||||
primary_conninfo['password'] = wanted_primary_conninfo['password']
|
||||
else:
|
||||
return False
|
||||
|
||||
return all(primary_conninfo.get(p) == str(v) for p, v in wanted_primary_conninfo.items() if v is not None)
|
||||
|
||||
def check_recovery_conf(self, member):
|
||||
"""Returns a tuple. The first boolean element indicates that recovery params don't match
|
||||
and the second is set to `True` if the restart is required in order to apply new values"""
|
||||
|
||||
# TODO: recovery.conf could be stale, would be nice to detect that.
|
||||
if self._postgresql.major_version >= 120000:
|
||||
if not os.path.exists(self._standby_signal):
|
||||
return True, True
|
||||
|
||||
_read_recovery_params = self._read_recovery_params
|
||||
else:
|
||||
if not self.recovery_conf_exists():
|
||||
return True, True
|
||||
|
||||
_read_recovery_params = self._read_recovery_params_pre_v12
|
||||
|
||||
params, updated = _read_recovery_params()
|
||||
# updated indicates that mtime of postgresql.conf, postgresql.auto.conf, or recovery.conf
|
||||
# was changed and params were read either from the config or from the database connection.
|
||||
if updated:
|
||||
if params is None: # exception or unparsable config
|
||||
return True, True
|
||||
|
||||
# We will cache parsed value until the next config change.
|
||||
self._current_recovery_params = params
|
||||
primary_conninfo = params['primary_conninfo']
|
||||
if primary_conninfo[0]:
|
||||
primary_conninfo[0] = parse_dsn(params['primary_conninfo'][0])
|
||||
# If we failed to parse non-empty connection string this indicates that config if broken.
|
||||
if not primary_conninfo[0]:
|
||||
return True, True
|
||||
else: # empty string, primary_conninfo is not in the config
|
||||
primary_conninfo[0] = {}
|
||||
|
||||
required = {'restart': 0, 'reload': 0}
|
||||
|
||||
def record_missmatch(mtype):
|
||||
required['restart' if mtype else 'reload'] += 1
|
||||
|
||||
wanted_recovery_params = self.build_recovery_params(member)
|
||||
for param, value in self._current_recovery_params.items():
|
||||
# Skip certain parameters defined in the included postgres config files
|
||||
# if we know that they are not specified in the patroni configuration.
|
||||
if len(value) > 2 and value[2] not in (self._postgresql_conf, self._auto_conf) and \
|
||||
param in ('archive_cleanup_command', 'promote_trigger_file', 'recovery_end_command',
|
||||
'recovery_min_apply_delay', 'restore_command') and param not in wanted_recovery_params:
|
||||
continue
|
||||
if param == 'recovery_min_apply_delay':
|
||||
if not compare_values('integer', 'ms', value[0], wanted_recovery_params.get(param, 0)):
|
||||
record_missmatch(value[1])
|
||||
elif param == 'primary_conninfo':
|
||||
if not self._check_primary_conninfo(value[0], wanted_recovery_params.get('primary_conninfo', {})):
|
||||
record_missmatch(value[1])
|
||||
elif (param != 'primary_slot_name' or wanted_recovery_params.get('primary_conninfo')) \
|
||||
and str(value[0]) != str(wanted_recovery_params.get(param, '')):
|
||||
record_missmatch(value[1])
|
||||
return required['restart'] + required['reload'] > 0, required['restart'] > 0
|
||||
|
||||
@staticmethod
|
||||
def _remove_file_if_exists(name):
|
||||
if os.path.isfile(name) or os.path.islink(name):
|
||||
os.unlink(name)
|
||||
|
||||
@staticmethod
|
||||
def _pgpass_line(record):
|
||||
if 'password' in record:
|
||||
def escape(value):
|
||||
return re.sub(r'([:\\])', r'\\\1', str(value))
|
||||
|
||||
record = {n: escape(record.get(n, '*')) for n in ('host', 'port', 'user', 'password')}
|
||||
return '{host}:{port}:*:{user}:{password}'.format(**record)
|
||||
|
||||
def write_pgpass(self, record):
|
||||
line = self._pgpass_line(record)
|
||||
if not line:
|
||||
return os.environ.copy()
|
||||
|
||||
with open(self._pgpass, 'w') as f:
|
||||
os.chmod(self._pgpass, stat.S_IWRITE | stat.S_IREAD)
|
||||
f.write(line)
|
||||
|
||||
env = os.environ.copy()
|
||||
env['PGPASSFILE'] = self._pgpass
|
||||
return env
|
||||
|
||||
def write_recovery_conf(self, recovery_params):
|
||||
if self._postgresql.major_version >= 120000:
|
||||
if parse_bool(recovery_params.pop('standby_mode', None)):
|
||||
@@ -480,10 +748,9 @@ class ConfigHandler(object):
|
||||
open(self._recovery_signal, 'w').close()
|
||||
self._recovery_params = recovery_params
|
||||
else:
|
||||
with open(self._recovery_conf, 'w') as f:
|
||||
with ConfigWriter(self._recovery_conf) as f:
|
||||
os.chmod(self._recovery_conf, stat.S_IWRITE | stat.S_IREAD)
|
||||
for name, value in recovery_params.items():
|
||||
f.write("{0} = '{1}'\n".format(name, value))
|
||||
self._write_recovery_params(f, recovery_params)
|
||||
|
||||
def remove_recovery_conf(self):
|
||||
for name in (self._recovery_conf, self._standby_signal, self._recovery_signal):
|
||||
@@ -522,13 +789,9 @@ class ConfigHandler(object):
|
||||
self._config['recovery_conf'] = recovery_conf
|
||||
|
||||
if self.get('recovery_conf'):
|
||||
good_name, bad_name = 'trigger_file', 'promote_trigger_file'
|
||||
if self._postgresql.major_version >= 120000:
|
||||
good_name, bad_name = bad_name, good_name
|
||||
|
||||
value = self._config['recovery_conf'].pop(bad_name, None)
|
||||
if good_name not in self._config['recovery_conf'] and value:
|
||||
self._config['recovery_conf'][good_name] = value
|
||||
value = self._config['recovery_conf'].pop(self._triggerfile_wrong_name, None)
|
||||
if self._triggerfile_good_name not in self._config['recovery_conf'] and value:
|
||||
self._config['recovery_conf'][self._triggerfile_good_name] = value
|
||||
|
||||
def get_server_parameters(self, config):
|
||||
parameters = config['parameters'].copy()
|
||||
@@ -568,14 +831,18 @@ class ConfigHandler(object):
|
||||
@property
|
||||
def local_connect_kwargs(self):
|
||||
ret = self._local_address.copy()
|
||||
# add all of the other connection settings that are available
|
||||
ret.update(self._superuser)
|
||||
# if the "username" parameter is present, it actually needs to be "user"
|
||||
# for connecting to PostgreSQL
|
||||
if 'username' in self._superuser:
|
||||
ret['user'] = self._superuser['username']
|
||||
del ret['username']
|
||||
# ensure certain Patroni configurations are available
|
||||
ret.update({'database': self._postgresql.database,
|
||||
'fallback_application_name': 'Patroni',
|
||||
'connect_timeout': 3,
|
||||
'options': '-c statement_timeout=2000'})
|
||||
if 'username' in self._superuser:
|
||||
ret['user'] = self._superuser['username']
|
||||
if 'password' in self._superuser:
|
||||
ret['password'] = self._superuser['password']
|
||||
return ret
|
||||
|
||||
def resolve_connection_addresses(self):
|
||||
@@ -602,52 +869,75 @@ class ConfigHandler(object):
|
||||
|
||||
self._postgresql.set_connection_kwargs(self.local_connect_kwargs)
|
||||
|
||||
def reload_config(self, config):
|
||||
def _get_pg_settings(self, names):
|
||||
return {r[0]: r for r in self._postgresql.query(('SELECT name, setting, unit, vartype, context, sourcefile'
|
||||
+ ' FROM pg_catalog.pg_settings ' +
|
||||
' WHERE pg_catalog.lower(name) = ANY(%s)'),
|
||||
[n.lower() for n in names])}
|
||||
|
||||
@staticmethod
|
||||
def _handle_wal_buffers(old_values, changes):
|
||||
wal_block_size = parse_int(old_values['wal_block_size'][1])
|
||||
wal_segment_size = old_values['wal_segment_size']
|
||||
wal_segment_unit = parse_int(wal_segment_size[2], 'B') if wal_segment_size[2][0].isdigit() else 1
|
||||
wal_segment_size = parse_int(wal_segment_size[1]) * wal_segment_unit / wal_block_size
|
||||
default_wal_buffers = min(max(parse_int(old_values['shared_buffers'][1]) / 32, 8), wal_segment_size)
|
||||
|
||||
wal_buffers = old_values['wal_buffers']
|
||||
new_value = str(changes['wal_buffers'] or -1)
|
||||
|
||||
new_value = default_wal_buffers if new_value == '-1' else parse_int(new_value, wal_buffers[2])
|
||||
old_value = default_wal_buffers if wal_buffers[1] == '-1' else parse_int(*wal_buffers[1:3])
|
||||
|
||||
if new_value == old_value:
|
||||
del changes['wal_buffers']
|
||||
|
||||
def reload_config(self, config, sighup=False):
|
||||
self._superuser = config['authentication'].get('superuser', {})
|
||||
server_parameters = self.get_server_parameters(config)
|
||||
|
||||
conf_changed = hba_changed = ident_changed = local_connection_address_changed = pending_restart = False
|
||||
if self._postgresql.state == 'running':
|
||||
changes = CaseInsensitiveDict({p: v for p, v in server_parameters.items() if '.' not in p})
|
||||
changes.update({p: None for p in self._server_parameters.keys() if not ('.' in p or p in changes)})
|
||||
changes = CaseInsensitiveDict({p: v for p, v in server_parameters.items()
|
||||
if p.lower() not in self._RECOVERY_PARAMETERS})
|
||||
changes.update({p: None for p in self._server_parameters.keys()
|
||||
if not (p in changes or p.lower() in self._RECOVERY_PARAMETERS)})
|
||||
if changes:
|
||||
if 'wal_buffers' in changes: # we need to calculate the default value of wal_buffers
|
||||
undef = [p for p in ('shared_buffers', 'wal_segment_size', 'wal_block_size') if p not in changes]
|
||||
changes.update({p: None for p in undef})
|
||||
# XXX: query can raise an exception
|
||||
for r in self._postgresql.query(('SELECT name, setting, unit, vartype, context '
|
||||
+ 'FROM pg_catalog.pg_settings ' +
|
||||
' WHERE pg_catalog.lower(name) IN ('
|
||||
+ ', '.join(['%s'] * len(changes)) +
|
||||
')'), *(k.lower() for k in changes.keys())):
|
||||
old_values = self._get_pg_settings(changes.keys())
|
||||
if 'wal_buffers' in changes:
|
||||
self._handle_wal_buffers(old_values, changes)
|
||||
for p in undef:
|
||||
del changes[p]
|
||||
|
||||
for r in old_values.values():
|
||||
if r[4] != 'internal' and r[0] in changes:
|
||||
new_value = changes.pop(r[0])
|
||||
if new_value is None or not compare_values(r[3], r[2], r[1], new_value):
|
||||
conf_changed = True
|
||||
if r[4] == 'postmaster':
|
||||
pending_restart = True
|
||||
logger.info('Changed %s from %s to %s (restart required)', r[0], r[1], new_value)
|
||||
logger.info('Changed %s from %s to %s (restart might be required)',
|
||||
r[0], r[1], new_value)
|
||||
if config.get('use_unix_socket') and r[0] == 'unix_socket_directories'\
|
||||
or r[0] in ('listen_addresses', 'port'):
|
||||
local_connection_address_changed = True
|
||||
else:
|
||||
logger.info('Changed %s from %s to %s', r[0], r[1], new_value)
|
||||
conf_changed = True
|
||||
for param in changes:
|
||||
if param in server_parameters:
|
||||
for param, value in changes.items():
|
||||
if '.' in param:
|
||||
# Check that user-defined-paramters have changed (parameters with period in name)
|
||||
if value is None or param not in self._server_parameters \
|
||||
or str(value) != str(self._server_parameters[param]):
|
||||
logger.info('Changed %s from %s to %s', param, self._server_parameters.get(param), value)
|
||||
conf_changed = True
|
||||
elif param in server_parameters:
|
||||
logger.warning('Removing invalid parameter `%s` from postgresql.parameters', param)
|
||||
server_parameters.pop(param)
|
||||
|
||||
# Check that user-defined-paramters have changed (parameters with period in name)
|
||||
if not conf_changed:
|
||||
for p, v in server_parameters.items():
|
||||
if '.' in p and (p not in self._server_parameters or str(v) != str(self._server_parameters[p])):
|
||||
logger.info('Changed %s from %s to %s', p, self._server_parameters.get(p), v)
|
||||
conf_changed = True
|
||||
break
|
||||
if not conf_changed:
|
||||
for p, v in self._server_parameters.items():
|
||||
if '.' in p and (p not in server_parameters or str(v) != str(server_parameters[p])):
|
||||
logger.info('Changed %s from %s to %s', p, v, server_parameters.get(p))
|
||||
conf_changed = True
|
||||
break
|
||||
|
||||
if not server_parameters.get('hba_file') and config.get('pg_hba'):
|
||||
hba_changed = self._config.get('pg_hba', []) != config['pg_hba']
|
||||
|
||||
@@ -658,7 +948,6 @@ class ConfigHandler(object):
|
||||
self._postgresql.set_pending_restart(pending_restart)
|
||||
self._server_parameters = server_parameters
|
||||
self._adjust_recovery_parameters()
|
||||
self._connect_address = config.get('connect_address')
|
||||
self._krbsrvname = config.get('krbsrvname')
|
||||
|
||||
# for not so obvious connection attempts that may happen outside of pyscopg2
|
||||
@@ -677,10 +966,18 @@ class ConfigHandler(object):
|
||||
if ident_changed:
|
||||
self.replace_pg_ident()
|
||||
|
||||
if conf_changed or hba_changed or ident_changed:
|
||||
logger.info('PostgreSQL configuration items changed, reloading configuration.')
|
||||
if sighup or conf_changed or hba_changed or ident_changed:
|
||||
logger.info('Reloading PostgreSQL configuration.')
|
||||
self._postgresql.reload()
|
||||
elif not pending_restart:
|
||||
if self._postgresql.major_version >= 90500:
|
||||
time.sleep(1)
|
||||
try:
|
||||
pending_restart = self._postgresql.query('SELECT COUNT(*) FROM pg_catalog.pg_settings'
|
||||
' WHERE pending_restart').fetchone()[0] > 0
|
||||
self._postgresql.set_pending_restart(pending_restart)
|
||||
except Exception as e:
|
||||
logger.warning('Exception %r when running query', e)
|
||||
else:
|
||||
logger.info('No PostgreSQL configuration items changed, nothing to reload.')
|
||||
|
||||
def set_synchronous_standby(self, name):
|
||||
@@ -728,6 +1025,10 @@ class ConfigHandler(object):
|
||||
|
||||
for name, cname in options_mapping.items():
|
||||
value = parse_int(effective_configuration[name])
|
||||
if cname not in data:
|
||||
logger.warning('%s is missing from pg_controldata output', cname)
|
||||
continue
|
||||
|
||||
cvalue = parse_int(data[cname])
|
||||
if cvalue > value:
|
||||
effective_configuration[name] = cvalue
|
||||
|
||||
@@ -5,13 +5,24 @@ import psutil
|
||||
import re
|
||||
import signal
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
from patroni import PATRONI_ENV_PREFIX, KUBERNETES_ENV_PREFIX
|
||||
|
||||
# avoid spawning the resource tracker process
|
||||
if sys.version_info >= (3, 8): # pragma: no cover
|
||||
import multiprocessing.resource_tracker
|
||||
multiprocessing.resource_tracker.getfd = lambda: 0
|
||||
elif sys.version_info >= (3, 4): # pragma: no cover
|
||||
import multiprocessing.semaphore_tracker
|
||||
multiprocessing.semaphore_tracker.getfd = lambda: 0
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
STOP_SIGNALS = {
|
||||
'smart': signal.SIGTERM,
|
||||
'fast': signal.SIGINT,
|
||||
'immediate': signal.SIGQUIT if os.name != 'nt' else signal.SIGABRT,
|
||||
'smart': 'TERM',
|
||||
'fast': 'INT',
|
||||
'immediate': 'QUIT',
|
||||
}
|
||||
|
||||
|
||||
@@ -94,7 +105,43 @@ class PostmasterProcess(psutil.Process):
|
||||
except psutil.NoSuchProcess:
|
||||
return None
|
||||
|
||||
def signal_stop(self, mode):
|
||||
def signal_kill(self):
|
||||
"""to suspend and kill postmaster and all children
|
||||
|
||||
:returns True if postmaster and children are killed, False if error
|
||||
"""
|
||||
try:
|
||||
self.suspend()
|
||||
except psutil.NoSuchProcess:
|
||||
return True
|
||||
except psutil.Error as e:
|
||||
logger.warning('Failed to suspend postmaster: %s', e)
|
||||
|
||||
try:
|
||||
children = self.children(recursive=True)
|
||||
except psutil.NoSuchProcess:
|
||||
return True
|
||||
except psutil.Error as e:
|
||||
logger.warning('Failed to get a list of postmaster children: %s', e)
|
||||
children = []
|
||||
|
||||
try:
|
||||
self.kill()
|
||||
except psutil.NoSuchProcess:
|
||||
return True
|
||||
except psutil.Error as e:
|
||||
logger.warning('Could not kill postmaster: %s', e)
|
||||
return False
|
||||
|
||||
for child in children:
|
||||
try:
|
||||
child.kill()
|
||||
except psutil.Error:
|
||||
pass
|
||||
psutil.wait_procs(children + [self])
|
||||
return True
|
||||
|
||||
def signal_stop(self, mode, pg_ctl='pg_ctl'):
|
||||
"""Signal postmaster process to stop
|
||||
|
||||
:returns None if signaled, True if process is already gone, False if error
|
||||
@@ -102,8 +149,10 @@ class PostmasterProcess(psutil.Process):
|
||||
if self.is_single_user:
|
||||
logger.warning("Cannot stop server; single-user server is running (PID: {0})".format(self.pid))
|
||||
return False
|
||||
if os.name != 'posix':
|
||||
return self.pg_ctl_kill(mode, pg_ctl)
|
||||
try:
|
||||
self.send_signal(STOP_SIGNALS[mode])
|
||||
self.send_signal(getattr(signal, 'SIG' + STOP_SIGNALS[mode]))
|
||||
except psutil.NoSuchProcess:
|
||||
return True
|
||||
except psutil.AccessDenied as e:
|
||||
@@ -112,6 +161,16 @@ class PostmasterProcess(psutil.Process):
|
||||
|
||||
return None
|
||||
|
||||
def pg_ctl_kill(self, mode, pg_ctl):
|
||||
try:
|
||||
status = subprocess.call([pg_ctl, "kill", STOP_SIGNALS[mode], str(self.pid)])
|
||||
except OSError:
|
||||
return False
|
||||
if status == 0:
|
||||
return None
|
||||
else:
|
||||
return not self.is_running()
|
||||
|
||||
def wait_for_user_backends_to_close(self):
|
||||
# These regexps are cross checked against versions PostgreSQL 9.1 .. 11
|
||||
aux_proc_re = re.compile("(?:postgres:)( .*:)? (?:(?:archiver|startup|autovacuum launcher|autovacuum worker|"
|
||||
@@ -153,7 +212,8 @@ class PostmasterProcess(psutil.Process):
|
||||
# In order to make everything portable we can't use fork&exec approach here, so we will call
|
||||
# ourselves and pass list of arguments which must be used to start postgres.
|
||||
# On Windows, in order to run a side-by-side assembly the specified env must include a valid SYSTEMROOT.
|
||||
env = {p: os.environ[p] for p in ('PATH', 'LD_LIBRARY_PATH', 'LC_ALL', 'LANG', 'SYSTEMROOT') if p in os.environ}
|
||||
env = {p: os.environ[p] for p in os.environ if not p.startswith(
|
||||
PATRONI_ENV_PREFIX) and not p.startswith(KUBERNETES_ENV_PREFIX)}
|
||||
try:
|
||||
proc = PostmasterProcess._from_pidfile(data_dir)
|
||||
if proc and not proc._is_postmaster_process():
|
||||
@@ -170,8 +230,9 @@ class PostmasterProcess(psutil.Process):
|
||||
pass
|
||||
cmdline = [pgcommand, '-D', data_dir, '--config-file={}'.format(conf)] + options
|
||||
logger.debug("Starting postgres: %s", " ".join(cmdline))
|
||||
parent_conn, child_conn = multiprocessing.Pipe(False)
|
||||
proc = multiprocessing.Process(target=pg_ctl_start, args=(child_conn, cmdline, env))
|
||||
ctx = multiprocessing.get_context('spawn') if sys.version_info >= (3, 4) else multiprocessing
|
||||
parent_conn, child_conn = ctx.Pipe(False)
|
||||
proc = ctx.Process(target=pg_ctl_start, args=(child_conn, cmdline, env))
|
||||
proc.start()
|
||||
pid = parent_conn.recv()
|
||||
proc.join()
|
||||
|
||||
@@ -2,9 +2,12 @@ import logging
|
||||
import os
|
||||
import subprocess
|
||||
|
||||
from patroni.dcs import Leader
|
||||
from patroni.postgresql.connection import get_connection_cursor
|
||||
from patroni.postgresql.misc import parse_history, parse_lsn
|
||||
from threading import Lock, Thread
|
||||
|
||||
from .connection import get_connection_cursor
|
||||
from .misc import parse_history, parse_lsn
|
||||
from ..async_executor import CriticalTask
|
||||
from ..dcs import Leader
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -16,6 +19,7 @@ class Rewind(object):
|
||||
|
||||
def __init__(self, postgresql):
|
||||
self._postgresql = postgresql
|
||||
self._checkpoint_task_lock = Lock()
|
||||
self.reset_state()
|
||||
|
||||
@staticmethod
|
||||
@@ -131,31 +135,41 @@ class Rewind(object):
|
||||
self._check_timeline_and_lsn(leader)
|
||||
return leader and leader.conn_url and self._state == REWIND_STATUS.NEED
|
||||
|
||||
def check_for_checkpoint_after_promote(self):
|
||||
def __checkpoint(self, task):
|
||||
try:
|
||||
result = self._postgresql.checkpoint()
|
||||
except Exception as e:
|
||||
result = 'Exception: ' + str(e)
|
||||
with task:
|
||||
task.complete(not bool(result))
|
||||
|
||||
def ensure_checkpoint_after_promote(self):
|
||||
"""After promote issue a CHECKPOINT from a new thread and asynchronously check the result.
|
||||
In case if CHECKPOINT failed, just check that timeline in pg_control was updated."""
|
||||
|
||||
if self._state == REWIND_STATUS.INITIAL and self._postgresql.is_leader():
|
||||
try:
|
||||
timeline = int(self._postgresql.controldata().get("Latest checkpoint's TimeLineID"))
|
||||
if self._postgresql.get_master_timeline() == timeline:
|
||||
self._state = REWIND_STATUS.CHECKPOINT
|
||||
except (TypeError, ValueError):
|
||||
logger.exception('Failed to parse timeline from pg_controldata output')
|
||||
with self._checkpoint_task_lock:
|
||||
if self._checkpoint_task:
|
||||
with self._checkpoint_task:
|
||||
if self._checkpoint_task.result:
|
||||
self._state = REWIND_STATUS.CHECKPOINT
|
||||
if self._checkpoint_task.result is not False:
|
||||
return
|
||||
else:
|
||||
self._checkpoint_task = CriticalTask()
|
||||
return Thread(target=self.__checkpoint, args=(self._checkpoint_task,)).start()
|
||||
|
||||
if self._postgresql.get_master_timeline() == self._postgresql.pg_control_timeline():
|
||||
self._state = REWIND_STATUS.CHECKPOINT
|
||||
|
||||
def checkpoint_after_promote(self):
|
||||
return self._state == REWIND_STATUS.CHECKPOINT
|
||||
|
||||
def pg_rewind(self, r):
|
||||
# prepare pg_rewind connection
|
||||
env = self._postgresql.write_pgpass(r)
|
||||
env = self._postgresql.config.write_pgpass(r)
|
||||
env['PGOPTIONS'] = '-c statement_timeout=0'
|
||||
dsn_attrs = [
|
||||
('user', r.get('user')),
|
||||
('host', r.get('host')),
|
||||
('port', r.get('port')),
|
||||
('dbname', r.get('database') or self._postgresql.database),
|
||||
('sslmode', 'prefer'),
|
||||
('sslcompression', '1'),
|
||||
]
|
||||
dsn = " ".join("{0}={1}".format(k, v) for k, v in dsn_attrs if v is not None)
|
||||
dsn = self._postgresql.config.format_dsn(r, True)
|
||||
logger.info('running pg_rewind from %s', dsn)
|
||||
try:
|
||||
return self._postgresql.cancellable.call([self._postgresql.pgcommand('pg_rewind'), '-D',
|
||||
@@ -203,6 +217,8 @@ class Rewind(object):
|
||||
|
||||
def reset_state(self):
|
||||
self._state = REWIND_STATUS.INITIAL
|
||||
with self._checkpoint_task_lock:
|
||||
self._checkpoint_task = None
|
||||
|
||||
@property
|
||||
def is_needed(self):
|
||||
|
||||
@@ -15,19 +15,14 @@ class SlotsHandler(object):
|
||||
|
||||
def __init__(self, postgresql):
|
||||
self._postgresql = postgresql
|
||||
self._use_slots = postgresql.config.get('use_slots', True)
|
||||
self._replication_slots = {} # already existing replication slots
|
||||
self.schedule()
|
||||
|
||||
@property
|
||||
def use_slots(self):
|
||||
return self._use_slots and self._postgresql.major_version >= 90400
|
||||
|
||||
def _query(self, sql, *params):
|
||||
return self._postgresql.query(sql, *params, retry=False)
|
||||
|
||||
def load_replication_slots(self):
|
||||
if self.use_slots and self._schedule_load_slots:
|
||||
if self._postgresql.major_version >= 90400 and self._schedule_load_slots:
|
||||
replication_slots = {}
|
||||
cursor = self._query('SELECT slot_name, slot_type, plugin, database FROM pg_catalog.pg_replication_slots')
|
||||
for r in cursor:
|
||||
@@ -45,7 +40,7 @@ class SlotsHandler(object):
|
||||
return cursor.rowcount == 1
|
||||
|
||||
def sync_replication_slots(self, cluster):
|
||||
if self.use_slots:
|
||||
if self._postgresql.major_version >= 90400:
|
||||
try:
|
||||
self.load_replication_slots()
|
||||
|
||||
@@ -104,5 +99,5 @@ class SlotsHandler(object):
|
||||
|
||||
def schedule(self, value=None):
|
||||
if value is None:
|
||||
value = self.use_slots
|
||||
value = self._postgresql.major_version >= 90400
|
||||
self._schedule_load_slots = value
|
||||
|
||||
@@ -0,0 +1,58 @@
|
||||
import json
|
||||
import urllib3
|
||||
import six
|
||||
|
||||
from six.moves.urllib_parse import urlparse, urlunparse
|
||||
|
||||
from .utils import USER_AGENT
|
||||
|
||||
|
||||
class PatroniRequest(object):
|
||||
|
||||
def __init__(self, config, insecure=False):
|
||||
cert_reqs = 'CERT_NONE' if insecure or config.get('ctl', {}).get('insecure', False) else 'CERT_REQUIRED'
|
||||
self._pool = urllib3.PoolManager(num_pools=10, maxsize=10, cert_reqs=cert_reqs)
|
||||
self.reload_config(config)
|
||||
|
||||
@staticmethod
|
||||
def _get_cfg_value(config, name):
|
||||
return config.get('ctl', {}).get(name) or config.get('restapi', {}).get(name)
|
||||
|
||||
def _apply_pool_param(self, param, value):
|
||||
if value:
|
||||
self._pool.connection_pool_kw[param] = value
|
||||
else:
|
||||
self._pool.connection_pool_kw.pop(param, None)
|
||||
|
||||
def _apply_ssl_file_param(self, config, name):
|
||||
value = self._get_cfg_value(config, name + 'file')
|
||||
self._apply_pool_param(name + '_file', value)
|
||||
return value
|
||||
|
||||
def reload_config(self, config):
|
||||
self._pool.headers = urllib3.make_headers(basic_auth=self._get_cfg_value(config, 'auth'), user_agent=USER_AGENT)
|
||||
|
||||
if self._apply_ssl_file_param(config, 'cert'):
|
||||
self._apply_ssl_file_param(config, 'key')
|
||||
else:
|
||||
self._pool.connection_pool_kw.pop('key_file', None)
|
||||
|
||||
cacert = config.get('ctl', {}).get('cacert') or config.get('restapi', {}).get('cafile')
|
||||
self._apply_pool_param('ca_certs', cacert)
|
||||
|
||||
def request(self, method, url, body=None, **kwargs):
|
||||
if body is not None and not isinstance(body, six.string_types):
|
||||
body = json.dumps(body)
|
||||
return self._pool.request(method.upper(), url, body=body, **kwargs)
|
||||
|
||||
def __call__(self, member, method='GET', endpoint=None, data=None, **kwargs):
|
||||
url = member.api_url
|
||||
if endpoint:
|
||||
scheme, netloc, _, _, _, _ = urlparse(url)
|
||||
url = urlunparse((scheme, netloc, endpoint, '', '', ''))
|
||||
return self.request(method, url, data, **kwargs)
|
||||
|
||||
|
||||
def get(url, verify=True, **kwargs):
|
||||
http = PatroniRequest({}, not verify)
|
||||
return http.request('GET', url, **kwargs)
|
||||
@@ -1,12 +1,12 @@
|
||||
#!/usr/bin/env python
|
||||
|
||||
import json
|
||||
import logging
|
||||
import requests
|
||||
from requests.exceptions import RequestException
|
||||
import sys
|
||||
import boto.ec2
|
||||
|
||||
from patroni.utils import Retry, RetryFailedError
|
||||
from patroni.request import get as requests_get
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@@ -19,14 +19,14 @@ class AWSConnection(object):
|
||||
self._retry = Retry(deadline=300, max_delay=30, max_tries=-1, retry_exceptions=(boto.exception.StandardError,))
|
||||
try:
|
||||
# get the instance id
|
||||
r = requests.get('http://169.254.169.254/latest/dynamic/instance-identity/document', timeout=2.1)
|
||||
except RequestException:
|
||||
r = requests_get('http://169.254.169.254/latest/dynamic/instance-identity/document', timeout=2.1)
|
||||
except Exception:
|
||||
logger.error('cannot query AWS meta-data')
|
||||
return
|
||||
|
||||
if r.ok:
|
||||
if r.status < 400:
|
||||
try:
|
||||
content = r.json()
|
||||
content = json.loads(r.data.decode('utf-8'))
|
||||
self.instance_id = content['instanceId']
|
||||
self.region = content['region']
|
||||
except Exception:
|
||||
|
||||
+102
-3
@@ -1,15 +1,21 @@
|
||||
import logging
|
||||
import os
|
||||
import platform
|
||||
import random
|
||||
import re
|
||||
import tempfile
|
||||
import time
|
||||
|
||||
from dateutil import tz
|
||||
from patroni.exceptions import PatroniException
|
||||
|
||||
from .exceptions import PatroniException
|
||||
from .version import __version__
|
||||
|
||||
tzutc = tz.tzutc()
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
USER_AGENT = 'Patroni/{0} Python/{1} {2}'.format(__version__, platform.python_version(), platform.system())
|
||||
OCT_RE = re.compile(r'^[-+]?0[0-7]*')
|
||||
DEC_RE = re.compile(r'^[-+]?(0|[1-9][0-9]*)')
|
||||
HEX_RE = re.compile(r'^[-+]?0x[0-9a-fA-F]+')
|
||||
@@ -296,6 +302,17 @@ class Retry(object):
|
||||
max_jitter=self.max_jitter / 100.0, max_delay=self.max_delay, sleep_func=self.sleep_func,
|
||||
deadline=self.deadline, retry_exceptions=self.retry_exceptions)
|
||||
|
||||
@property
|
||||
def sleeptime(self):
|
||||
return self._cur_delay + (random.randint(0, self.max_jitter) / 100.0)
|
||||
|
||||
def update_delay(self):
|
||||
self._cur_delay = min(self._cur_delay * self.backoff, self.max_delay)
|
||||
|
||||
@property
|
||||
def stoptime(self):
|
||||
return self._cur_stoptime
|
||||
|
||||
def __call__(self, func, *args, **kwargs):
|
||||
"""Call a function with arguments until it completes without throwing a `retry_exceptions`
|
||||
|
||||
@@ -317,14 +334,14 @@ class Retry(object):
|
||||
logger.warning('Retry got exception: %s', e)
|
||||
raise RetryFailedError("Too many retry attempts")
|
||||
self._attempts += 1
|
||||
sleeptime = self._cur_delay + (random.randint(0, self.max_jitter) / 100.0)
|
||||
sleeptime = hasattr(e, 'sleeptime') and e.sleeptime or self.sleeptime
|
||||
|
||||
if self._cur_stoptime is not None and time.time() + sleeptime >= self._cur_stoptime:
|
||||
logger.warning('Retry got exception: %s', e)
|
||||
raise RetryFailedError("Exceeded retry deadline")
|
||||
logger.debug('Retry got exception: %s', e)
|
||||
self.sleep_func(sleeptime)
|
||||
self._cur_delay = min(self._cur_delay * self.backoff, self.max_delay)
|
||||
self.update_delay()
|
||||
|
||||
|
||||
def polling_loop(timeout, interval=1):
|
||||
@@ -352,3 +369,85 @@ def uri(proto, netloc, path='', user=None):
|
||||
path = '/{0}'.format(path) if path and not path.startswith('/') else path
|
||||
user = '{0}@'.format(user) if user else ''
|
||||
return '{0}://{1}{2}{3}{4}'.format(proto, user, host, port, path)
|
||||
|
||||
|
||||
def is_standby_cluster(config):
|
||||
# Check whether or not provided configuration describes a standby cluster
|
||||
return isinstance(config, dict) and (config.get('host') or config.get('port') or config.get('restore_command'))
|
||||
|
||||
|
||||
def cluster_as_json(cluster):
|
||||
leader_name = cluster.leader.name if cluster.leader else None
|
||||
xlog_location_cluster = cluster.last_leader_operation or 0
|
||||
|
||||
ret = {'members': []}
|
||||
for m in cluster.members:
|
||||
if m.name == leader_name:
|
||||
config = cluster.config.data if cluster.config and cluster.config.modify_index else {}
|
||||
role = 'standby_leader' if is_standby_cluster(config.get('standby_cluster')) else 'leader'
|
||||
elif m.name == cluster.sync.sync_standby:
|
||||
role = 'sync_standby'
|
||||
else:
|
||||
role = 'replica'
|
||||
|
||||
member = {'name': m.name, 'role': role, 'state': m.data.get('state', ''), 'api_url': m.api_url}
|
||||
conn_kwargs = m.conn_kwargs()
|
||||
if conn_kwargs.get('host'):
|
||||
member['host'] = conn_kwargs['host']
|
||||
if conn_kwargs.get('port'):
|
||||
member['port'] = int(conn_kwargs['port'])
|
||||
optional_attributes = ('timeline', 'pending_restart', 'scheduled_restart', 'tags')
|
||||
member.update({n: m.data[n] for n in optional_attributes if n in m.data})
|
||||
|
||||
if m.name != leader_name:
|
||||
xlog_location = m.data.get('xlog_location')
|
||||
if xlog_location is None:
|
||||
member['lag'] = 'unknown'
|
||||
elif xlog_location_cluster >= xlog_location:
|
||||
member['lag'] = xlog_location_cluster - xlog_location
|
||||
else:
|
||||
member['lag'] = 0
|
||||
|
||||
ret['members'].append(member)
|
||||
|
||||
# sort members by name for consistency
|
||||
ret['members'].sort(key=lambda m: m['name'])
|
||||
if cluster.is_paused():
|
||||
ret['pause'] = True
|
||||
if cluster.failover and cluster.failover.scheduled_at:
|
||||
ret['scheduled_switchover'] = {'at': cluster.failover.scheduled_at.isoformat()}
|
||||
if cluster.failover.leader:
|
||||
ret['scheduled_switchover']['from'] = cluster.failover.leader
|
||||
if cluster.failover.candidate:
|
||||
ret['scheduled_switchover']['to'] = cluster.failover.candidate
|
||||
return ret
|
||||
|
||||
|
||||
def is_subpath(d1, d2):
|
||||
real_d1 = os.path.realpath(d1) + os.path.sep
|
||||
real_d2 = os.path.realpath(os.path.join(real_d1, d2))
|
||||
return os.path.commonprefix([real_d1, real_d2 + os.path.sep]) == real_d1
|
||||
|
||||
|
||||
def validate_directory(d, msg="{} {}"):
|
||||
if not os.path.exists(d):
|
||||
try:
|
||||
os.makedirs(d)
|
||||
except OSError as e:
|
||||
logger.error(e)
|
||||
raise PatroniException(msg.format(d, "couldn't create the directory"))
|
||||
elif os.path.isdir(d):
|
||||
try:
|
||||
fd, tmpfile = tempfile.mkstemp(dir=d)
|
||||
os.close(fd)
|
||||
os.remove(tmpfile)
|
||||
except OSError:
|
||||
raise PatroniException(msg.format(d, "the directory is not writable"))
|
||||
else:
|
||||
raise PatroniException(msg.format(d, "is not a directory"))
|
||||
|
||||
|
||||
def data_directory_is_empty(data_dir):
|
||||
if not os.path.exists(data_dir):
|
||||
return True
|
||||
return all(os.name != 'nt' and (n.startswith('.') or n == 'lost+found') for n in os.listdir(data_dir))
|
||||
|
||||
@@ -0,0 +1,379 @@
|
||||
#!/usr/bin/env python3
|
||||
import os
|
||||
import socket
|
||||
import re
|
||||
import subprocess
|
||||
|
||||
from patroni.utils import split_host_port, data_directory_is_empty
|
||||
from patroni.ctl import find_executable
|
||||
from patroni.dcs import dcs_modules
|
||||
from patroni.exceptions import ConfigParseError
|
||||
from six import string_types
|
||||
|
||||
|
||||
def data_directory_empty(data_dir):
|
||||
if os.path.isfile(os.path.join(data_dir, "global", "pg_control")):
|
||||
return False
|
||||
return data_directory_is_empty(data_dir)
|
||||
|
||||
|
||||
def validate_connect_address(address):
|
||||
try:
|
||||
host, _ = split_host_port(address, 1)
|
||||
except (AttributeError, TypeError, ValueError):
|
||||
raise ConfigParseError("contains a wrong value")
|
||||
if host in ["127.0.0.1", "0.0.0.0", "*", "::1", "localhost"]:
|
||||
raise ConfigParseError('must not contain "127.0.0.1", "0.0.0.0", "*", "::1", "localhost"')
|
||||
return True
|
||||
|
||||
|
||||
def validate_host_port(host_port, listen=False, multiple_hosts=False):
|
||||
try:
|
||||
hosts, port = split_host_port(host_port, None)
|
||||
except (ValueError, TypeError):
|
||||
raise ConfigParseError("contains a wrong value")
|
||||
else:
|
||||
if multiple_hosts:
|
||||
hosts = hosts.split(",")
|
||||
else:
|
||||
hosts = [hosts]
|
||||
for host in hosts:
|
||||
proto = socket.getaddrinfo(host, "", 0, socket.SOCK_STREAM, 0, socket.AI_PASSIVE)
|
||||
s = socket.socket(proto[0][0], socket.SOCK_STREAM)
|
||||
try:
|
||||
if s.connect_ex((host, port)) == 0:
|
||||
if listen:
|
||||
raise ConfigParseError("Port {} is already in use.".format(port))
|
||||
elif not listen:
|
||||
raise ConfigParseError("{} is not reachable".format(host_port))
|
||||
except socket.gaierror as e:
|
||||
raise ConfigParseError(e)
|
||||
finally:
|
||||
s.close()
|
||||
return True
|
||||
|
||||
|
||||
def comma_separated_host_port(string):
|
||||
assert all([validate_host_port(s.strip()) for s in string.split(",")]), "didn't pass the validation"
|
||||
return True
|
||||
|
||||
|
||||
def validate_host_port_listen(host_port):
|
||||
return validate_host_port(host_port, listen=True)
|
||||
|
||||
|
||||
def validate_host_port_listen_multiple_hosts(host_port):
|
||||
return validate_host_port(host_port, listen=True, multiple_hosts=True)
|
||||
|
||||
|
||||
def is_ipv4_address(ip):
|
||||
try:
|
||||
socket.inet_aton(ip)
|
||||
except Exception:
|
||||
raise ConfigParseError("Is not a valid ipv4 address")
|
||||
return True
|
||||
|
||||
|
||||
def is_ipv6_address(ip):
|
||||
try:
|
||||
socket.inet_pton(socket.AF_INET6, ip)
|
||||
except Exception:
|
||||
raise ConfigParseError("Is not a valid ipv6 address")
|
||||
return True
|
||||
|
||||
|
||||
def get_major_version(bin_dir=None):
|
||||
if not bin_dir:
|
||||
binary = 'postgres'
|
||||
else:
|
||||
binary = os.path.join(bin_dir, 'postgres')
|
||||
version = subprocess.check_output([binary, '--version']).decode()
|
||||
version = re.match(r'^[^\s]+ [^\s]+ (\d+)(\.(\d+))?', version)
|
||||
return '.'.join([version.group(1), version.group(3)]) if int(version.group(1)) < 10 else version.group(1)
|
||||
|
||||
|
||||
def validate_data_dir(data_dir):
|
||||
if not data_dir:
|
||||
raise ConfigParseError("is an empty string")
|
||||
elif os.path.exists(data_dir) and not os.path.isdir(data_dir):
|
||||
raise ConfigParseError("is not a directory")
|
||||
elif not data_directory_empty(data_dir):
|
||||
if not os.path.exists(os.path.join(data_dir, "PG_VERSION")):
|
||||
raise ConfigParseError("doesn't look like a valid data directory")
|
||||
else:
|
||||
with open(os.path.join(data_dir, "PG_VERSION"), "r") as version:
|
||||
pgversion = version.read().strip()
|
||||
waldir = ("pg_wal" if float(pgversion) >= 10 else "pg_xlog")
|
||||
if not os.path.isdir(os.path.join(data_dir, waldir)):
|
||||
raise ConfigParseError("data dir for the cluster is not empty, but doesn't contain"
|
||||
" \"{}\" directory".format(waldir))
|
||||
bin_dir = schema.data.get("postgresql", {}).get("bin_dir", None)
|
||||
major_version = get_major_version(bin_dir)
|
||||
if pgversion != major_version:
|
||||
raise ConfigParseError("data_dir directory postgresql version ({}) doesn't match with "
|
||||
"'postgres --version' output ({})".format(pgversion, major_version))
|
||||
return True
|
||||
|
||||
|
||||
class Result(object):
|
||||
def __init__(self, status, error="didn't pass validation", level=0, path="", data=""):
|
||||
self.status = status
|
||||
self.path = path
|
||||
self.data = data
|
||||
self.level = level
|
||||
self._error = error
|
||||
if not self.status:
|
||||
self.error = error
|
||||
else:
|
||||
self.error = None
|
||||
|
||||
def __repr__(self):
|
||||
return self.path + (" " + str(self.data) + " " + self._error if self.error else "")
|
||||
|
||||
|
||||
class Case(object):
|
||||
def __init__(self, schema):
|
||||
self._schema = schema
|
||||
|
||||
|
||||
class Or(object):
|
||||
def __init__(self, *args):
|
||||
self.args = args
|
||||
|
||||
|
||||
class Optional(object):
|
||||
def __init__(self, name):
|
||||
self.name = name
|
||||
|
||||
|
||||
class Directory(object):
|
||||
def __init__(self, contains=None, contains_executable=None):
|
||||
self.contains = contains
|
||||
self.contains_executable = contains_executable
|
||||
|
||||
def validate(self, name):
|
||||
if not name:
|
||||
yield Result(False, "is an empty string")
|
||||
elif not os.path.exists(name):
|
||||
yield Result(False, "Directory '{}' does not exist.".format(name))
|
||||
elif not os.path.isdir(name):
|
||||
yield Result(False, "'{}' is not a directory.".format(name))
|
||||
else:
|
||||
if self.contains:
|
||||
for path in self.contains:
|
||||
if not os.path.exists(os.path.join(name, path)):
|
||||
yield Result(False, "'{}' does not contain '{}'".format(name, path))
|
||||
if self.contains_executable:
|
||||
for program in self.contains_executable:
|
||||
if not find_executable(program, name):
|
||||
yield Result(False, "'{}' does not contain '{}'".format(name, program))
|
||||
|
||||
|
||||
class Schema(object):
|
||||
def __init__(self, validator):
|
||||
self.validator = validator
|
||||
|
||||
def __call__(self, data):
|
||||
for i in self.validate(data):
|
||||
if not i.status:
|
||||
print(i)
|
||||
|
||||
def validate(self, data):
|
||||
self.data = data
|
||||
if isinstance(self.validator, string_types):
|
||||
yield Result(isinstance(self.data, string_types), "is not a string", level=1, data=self.data)
|
||||
elif issubclass(type(self.validator), type):
|
||||
validator = self.validator
|
||||
if self.validator == str:
|
||||
validator = string_types
|
||||
yield Result(isinstance(self.data, validator),
|
||||
"is not {}".format(_get_type_name(self.validator)), level=1, data=self.data)
|
||||
elif callable(self.validator):
|
||||
if hasattr(self.validator, "expected_type"):
|
||||
if not isinstance(data, self.validator.expected_type):
|
||||
yield Result(False, "is not {}"
|
||||
.format(_get_type_name(self.validator.expected_type)), level=1, data=self.data)
|
||||
return
|
||||
try:
|
||||
self.validator(data)
|
||||
yield Result(True, data=self.data)
|
||||
except Exception as e:
|
||||
yield Result(False, "didn't pass validation: {}".format(e), data=self.data)
|
||||
elif isinstance(self.validator, dict):
|
||||
if not len(self.validator):
|
||||
yield Result(isinstance(self.data, dict), "is not a dictionary", level=1, data=self.data)
|
||||
elif isinstance(self.validator, list):
|
||||
if not isinstance(self.data, list):
|
||||
yield Result(isinstance(self.data, list), "is not a list", level=1, data=self.data)
|
||||
return
|
||||
for i in self.iter():
|
||||
yield i
|
||||
|
||||
def iter(self):
|
||||
if isinstance(self.validator, dict):
|
||||
if not isinstance(self.data, dict):
|
||||
yield Result(False, "is not a dictionary.", level=1)
|
||||
else:
|
||||
for i in self.iter_dict():
|
||||
yield i
|
||||
elif isinstance(self.validator, list):
|
||||
if len(self.data) == 0:
|
||||
yield Result(False, "is an empty list", data=self.data)
|
||||
if len(self.validator) > 0:
|
||||
for key, value in enumerate(self.data):
|
||||
for v in Schema(self.validator[0]).validate(value):
|
||||
yield Result(v.status, v.error,
|
||||
path=(str(key) + ("." + v.path if v.path else "")), level=v.level, data=value)
|
||||
elif isinstance(self.validator, Directory):
|
||||
for v in self.validator.validate(self.data):
|
||||
yield v
|
||||
elif isinstance(self.validator, Or):
|
||||
for i in self.iter_or():
|
||||
yield i
|
||||
|
||||
def iter_dict(self):
|
||||
for key in self.validator.keys():
|
||||
for d in self._data_key(key):
|
||||
if d not in self.data and not isinstance(key, Optional):
|
||||
yield Result(False, "is not defined.", path=d)
|
||||
elif d not in self.data and isinstance(key, Optional):
|
||||
continue
|
||||
else:
|
||||
validator = self.validator[key]
|
||||
if isinstance(key, Or) and isinstance(self.validator[key], Case):
|
||||
validator = self.validator[key]._schema[d]
|
||||
for v in Schema(validator).validate(self.data[d]):
|
||||
yield Result(v.status, v.error,
|
||||
path=(d + ("." + v.path if v.path else "")), level=v.level, data=v.data)
|
||||
|
||||
def iter_or(self):
|
||||
results = []
|
||||
for a in self.validator.args:
|
||||
r = []
|
||||
for v in Schema(a).validate(self.data):
|
||||
r.append(v)
|
||||
if any([x.status for x in r]) and not all([x.status for x in r]):
|
||||
results += filter(lambda x: not x.status, r)
|
||||
else:
|
||||
results += r
|
||||
if not any([x.status for x in results]):
|
||||
max_level = 3
|
||||
for v in sorted(results, key=lambda x: x.level):
|
||||
if v.level > max_level:
|
||||
break
|
||||
max_level = v.level
|
||||
yield Result(v.status, v.error, path=v.path, level=v.level, data=v.data)
|
||||
|
||||
def _data_key(self, key):
|
||||
if isinstance(self.data, dict) and isinstance(key, str):
|
||||
yield key
|
||||
elif isinstance(key, Optional):
|
||||
yield key.name
|
||||
elif isinstance(key, Or):
|
||||
if any([i in self.data for i in key.args]):
|
||||
for i in key.args:
|
||||
if i in self.data:
|
||||
yield i
|
||||
else:
|
||||
for i in key.args:
|
||||
yield i
|
||||
|
||||
|
||||
def _get_type_name(python_type):
|
||||
return {str: 'a string', int: 'and integer', float: 'a number', bool: 'a boolean',
|
||||
list: 'an array', dict: 'a dictionary', string_types: "a string"}.get(
|
||||
python_type, getattr(python_type, __name__, "unknown type"))
|
||||
|
||||
|
||||
def assert_(condition, message="Wrong value"):
|
||||
assert condition, message
|
||||
|
||||
|
||||
userattributes = {"username": "", Optional("password"): ""}
|
||||
available_dcs = [m.split(".")[-1] for m in dcs_modules()]
|
||||
comma_separated_host_port.expected_type = string_types
|
||||
validate_connect_address.expected_type = string_types
|
||||
validate_host_port_listen.expected_type = string_types
|
||||
validate_host_port_listen_multiple_hosts.expected_type = string_types
|
||||
validate_data_dir.expected_type = string_types
|
||||
|
||||
schema = Schema({
|
||||
"name": str,
|
||||
"scope": str,
|
||||
"restapi": {
|
||||
"listen": validate_host_port_listen,
|
||||
"connect_address": validate_connect_address
|
||||
},
|
||||
Optional("bootstrap"): {
|
||||
"dcs": {
|
||||
Optional("ttl"): int,
|
||||
Optional("loop_wait"): int,
|
||||
Optional("retry_timeout"): int,
|
||||
Optional("maximum_lag_on_failover"): int
|
||||
},
|
||||
"pg_hba": [str],
|
||||
"initdb": [Or(str, dict)]
|
||||
},
|
||||
Or(*available_dcs): Case({
|
||||
"consul": {
|
||||
Or("host", "url"): Case({
|
||||
"host": validate_host_port,
|
||||
"url": str})
|
||||
},
|
||||
"etcd": {
|
||||
Or("host", "hosts", "srv", "url", "proxy"): Case({
|
||||
"host": validate_host_port,
|
||||
"hosts": Or(comma_separated_host_port, [validate_host_port]),
|
||||
"srv": str,
|
||||
"url": str,
|
||||
"proxy": str})
|
||||
},
|
||||
"exhibitor": {
|
||||
"hosts": [str],
|
||||
"port": lambda i: assert_(int(i) <= 65535),
|
||||
Optional("pool_interval"): int
|
||||
},
|
||||
"zookeeper": {
|
||||
"hosts": Or(comma_separated_host_port, [validate_host_port]),
|
||||
},
|
||||
"kubernetes": {
|
||||
"labels": {},
|
||||
Optional("namespace"): str,
|
||||
Optional("scope_label"): str,
|
||||
Optional("role_label"): str,
|
||||
Optional("use_endpoints"): bool,
|
||||
Optional("pod_ip"): Or(is_ipv4_address, is_ipv6_address),
|
||||
Optional("ports"): [{"name": str, "port": int}],
|
||||
},
|
||||
}),
|
||||
"postgresql": {
|
||||
"listen": validate_host_port_listen_multiple_hosts,
|
||||
"connect_address": validate_connect_address,
|
||||
"authentication": {
|
||||
"replication": userattributes,
|
||||
"superuser": userattributes,
|
||||
"rewind": userattributes
|
||||
},
|
||||
"data_dir": validate_data_dir,
|
||||
Optional("bin_dir"): Directory(contains_executable=["pg_ctl", "initdb", "pg_controldata", "pg_basebackup",
|
||||
"postgres", "pg_isready"]),
|
||||
Optional("parameters"): {
|
||||
Optional("unix_socket_directories"): lambda s: assert_(all([isinstance(s, string_types), len(s)]))
|
||||
},
|
||||
Optional("pg_hba"): [str],
|
||||
Optional("pg_ident"): [str],
|
||||
Optional("pg_ctl_timeout"): int,
|
||||
Optional("use_pg_rewind"): bool
|
||||
},
|
||||
Optional("watchdog"): {
|
||||
Optional("mode"): lambda m: assert_(m in ["off", "automatic", "required"]),
|
||||
Optional("device"): str
|
||||
},
|
||||
Optional("tags"): {
|
||||
Optional("nofailover"): bool,
|
||||
Optional("clonefrom"): bool,
|
||||
Optional("noloadbalance"): bool,
|
||||
Optional("replicatefrom"): str,
|
||||
Optional("nosync"): bool
|
||||
}
|
||||
})
|
||||
+1
-1
@@ -1 +1 @@
|
||||
__version__ = '1.6.0'
|
||||
__version__ = '1.6.5'
|
||||
|
||||
@@ -57,7 +57,7 @@ class WatchdogConfig(object):
|
||||
return not self == other
|
||||
|
||||
def get_impl(self):
|
||||
if self.driver == 'testing':
|
||||
if self.driver == 'testing': # pragma: no cover
|
||||
from patroni.watchdog.linux import TestingWatchdogDevice
|
||||
return TestingWatchdogDevice.from_config(self.driver_config)
|
||||
elif platform.system() == 'Linux' and self.driver == 'default':
|
||||
|
||||
@@ -16,11 +16,11 @@ IOC_DIRBITS = 2
|
||||
|
||||
# Non-generic platform special cases
|
||||
machine = platform.machine()
|
||||
if machine in ['mips', 'sparc', 'powerpc', 'ppc64']:
|
||||
if machine in ['mips', 'sparc', 'powerpc', 'ppc64']: # pragma: no cover
|
||||
IOC_SIZEBITS = 13
|
||||
IOC_DIRBITS = 3
|
||||
IOC_NONE, IOC_WRITE, IOC_READ = 1, 2, 4
|
||||
elif machine == 'parisc':
|
||||
elif machine == 'parisc': # pragma: no cover
|
||||
IOC_WRITE, IOC_READ = 2, 1
|
||||
|
||||
IOC_NRSHIFT = 0
|
||||
@@ -218,7 +218,7 @@ class LinuxWatchdogDevice(WatchdogBase):
|
||||
return timeout.value
|
||||
|
||||
|
||||
class TestingWatchdogDevice(LinuxWatchdogDevice):
|
||||
class TestingWatchdogDevice(LinuxWatchdogDevice): # pragma: no cover
|
||||
"""Converts timeout ioctls to regular writes that can be intercepted from a named pipe."""
|
||||
timeout = 60
|
||||
|
||||
|
||||
@@ -17,7 +17,18 @@ restapi:
|
||||
# cacert: /etc/ssl/certs/ssl-cacert-snakeoil.pem
|
||||
|
||||
etcd:
|
||||
#Provide host to do the initial discovery of the cluster topology:
|
||||
host: 127.0.0.1:2379
|
||||
#Or use "hosts" to provide multiple endpoints
|
||||
#Could be a comma separated string:
|
||||
#hosts: host1:port1,host2:port2
|
||||
#or an actual yaml list:
|
||||
#hosts:
|
||||
#- host1:port1
|
||||
#- host2:port2
|
||||
#Once discovery is complete Patroni will use the list of advertised clientURLs
|
||||
#It is possible to change this behavior through by setting:
|
||||
#use_proxies: true
|
||||
|
||||
bootstrap:
|
||||
# this section will be written into Etcd:/<namespace>/<scope>/config after initializing new cluster
|
||||
|
||||
@@ -17,7 +17,18 @@ restapi:
|
||||
# cacert: /etc/ssl/certs/ssl-cacert-snakeoil.pem
|
||||
|
||||
etcd:
|
||||
#Provide host to do the initial discovery of the cluster topology:
|
||||
host: 127.0.0.1:2379
|
||||
#Or use "hosts" to provide multiple endpoints
|
||||
#Could be a comma separated string:
|
||||
#hosts: host1:port1,host2:port2
|
||||
#or an actual yaml list:
|
||||
#hosts:
|
||||
#- host1:port1
|
||||
#- host2:port2
|
||||
#Once discovery is complete Patroni will use the list of advertised clientURLs
|
||||
#It is possible to change this behavior through by setting:
|
||||
#use_proxies: true
|
||||
|
||||
bootstrap:
|
||||
# this section will be written into Etcd:/<namespace>/<scope>/config after initializing new cluster
|
||||
|
||||
@@ -17,7 +17,18 @@ restapi:
|
||||
# cacert: /etc/ssl/certs/ssl-cacert-snakeoil.pem
|
||||
|
||||
etcd:
|
||||
#Provide host to do the initial discovery of the cluster topology:
|
||||
host: 127.0.0.1:2379
|
||||
#Or use "hosts" to provide multiple endpoints
|
||||
#Could be a comma separated string:
|
||||
#hosts: host1:port1,host2:port2
|
||||
#or an actual yaml list:
|
||||
#hosts:
|
||||
#- host1:port1
|
||||
#- host2:port2
|
||||
#Once discovery is complete Patroni will use the list of advertised clientURLs
|
||||
#It is possible to change this behavior through by setting:
|
||||
#use_proxies: true
|
||||
|
||||
bootstrap:
|
||||
# this section will be written into Etcd:/<namespace>/<scope>/config after initializing new cluster
|
||||
|
||||
+2
-4
@@ -1,15 +1,13 @@
|
||||
urllib3>=1.19.1,!=1.21
|
||||
boto
|
||||
PyYAML
|
||||
requests
|
||||
six >= 1.7
|
||||
kazoo>=1.3.1
|
||||
python-etcd>=0.4.3,<0.5
|
||||
python-consul>=0.7.0
|
||||
python-consul>=0.7.1
|
||||
click>=4.1
|
||||
prettytable>=0.7
|
||||
tzlocal
|
||||
python-dateutil
|
||||
psutil>=2.0.0
|
||||
cdiff
|
||||
kubernetes>=2.0.0,<=7.0.0,!=4.0.*,!=5.0.*
|
||||
kubernetes>=2.0.0,<=10.0.1,!=4.0.*,!=5.0.*
|
||||
|
||||
@@ -8,22 +8,12 @@ import inspect
|
||||
import os
|
||||
import sys
|
||||
|
||||
from patroni import check_psycopg2, fatal
|
||||
from patroni.version import __version__ as VERSION
|
||||
from setuptools.command.test import test as TestCommand
|
||||
from setuptools import find_packages, setup
|
||||
|
||||
if sys.version_info < (2, 7, 0):
|
||||
fatal('patroni needs to be run with Python 2.7+')
|
||||
check_psycopg2()
|
||||
del sys.modules['patroni']
|
||||
del sys.modules['patroni.version']
|
||||
from setuptools import Command, find_packages, setup
|
||||
|
||||
__location__ = os.path.join(os.getcwd(), os.path.dirname(inspect.getfile(inspect.currentframe())))
|
||||
|
||||
NAME = 'patroni'
|
||||
MAIN_PACKAGE = NAME
|
||||
SCRIPTS = 'scripts'
|
||||
DESCRIPTION = 'PostgreSQL High-Available orchestrator and CLI'
|
||||
LICENSE = 'The MIT License'
|
||||
URL = 'https://github.com/zalando/patroni'
|
||||
@@ -32,9 +22,10 @@ AUTHOR_EMAIL = '[email protected], [email protected], alexk
|
||||
KEYWORDS = 'etcd governor patroni postgresql postgres ha haproxy confd' +\
|
||||
' zookeeper exhibitor consul streaming replication kubernetes k8s'
|
||||
|
||||
EXTRAS_REQUIRE = {'aws': ['boto'], 'etcd': ['python-etcd'], 'consul': ['python-consul'],
|
||||
'exhibitor': ['kazoo'], 'zookeeper': ['kazoo'], 'kubernetes': ['kubernetes']}
|
||||
COVERAGE_XML = True
|
||||
COVERAGE_HTML = False
|
||||
JUNIT_XML = True
|
||||
|
||||
# Add here all kinds of additional classifiers as defined under
|
||||
# https://pypi.python.org/pypi?%3Aaction=list_classifiers
|
||||
@@ -64,78 +55,74 @@ CONSOLE_SCRIPTS = ['patroni = patroni:main',
|
||||
"patroni_aws = patroni.scripts.aws:main"]
|
||||
|
||||
|
||||
class PyTest(TestCommand):
|
||||
class PyTest(Command):
|
||||
|
||||
user_options = [('cov=', None, 'Run coverage'), ('cov-xml=', None, 'Generate junit xml report'), ('cov-html=',
|
||||
None, 'Generate junit html report'), ('junitxml=', None, 'Generate xml of test results')]
|
||||
user_options = [('cov=', None, 'Run coverage'), ('cov-xml=', None, 'Generate junit xml report'),
|
||||
('cov-html=', None, 'Generate junit html report')]
|
||||
|
||||
def initialize_options(self):
|
||||
TestCommand.initialize_options(self)
|
||||
self.cov = []
|
||||
self.cov_xml = False
|
||||
self.cov_html = False
|
||||
self.junitxml = None
|
||||
|
||||
def finalize_options(self):
|
||||
TestCommand.finalize_options(self)
|
||||
if self.cov_xml or self.cov_html:
|
||||
self.cov = ['--cov', MAIN_PACKAGE, '--cov-report', 'term-missing']
|
||||
if self.cov_xml:
|
||||
self.cov.extend(['--cov-report', 'xml'])
|
||||
if self.cov_html:
|
||||
self.cov.extend(['--cov-report', 'html'])
|
||||
if self.junitxml is not None:
|
||||
self.junitxml = ['--junitxml', self.junitxml]
|
||||
|
||||
def run_tests(self):
|
||||
try:
|
||||
import pytest
|
||||
except Exception:
|
||||
raise RuntimeError('py.test is not installed, run: pip install pytest')
|
||||
params = {'args': self.test_args}
|
||||
if self.cov:
|
||||
params['args'] += self.cov
|
||||
if self.junitxml:
|
||||
params['args'] += self.junitxml
|
||||
params['args'] += ['--doctest-modules', MAIN_PACKAGE, '-vv']
|
||||
|
||||
import logging
|
||||
silence = logging.WARNING
|
||||
logging.basicConfig(format='%(asctime)s %(levelname)s: %(message)s', level=os.getenv('LOGLEVEL', silence))
|
||||
params['args'] += ['-s' if logging.getLogger().getEffectiveLevel() < silence else '--capture=fd']
|
||||
if not os.getenv('SYSTEMROOT'):
|
||||
os.environ['SYSTEMROOT'] = '/'
|
||||
errno = pytest.main(**params)
|
||||
|
||||
args = ['--verbose', 'tests', '--doctest-modules', MAIN_PACKAGE] +\
|
||||
['-s' if logging.getLogger().getEffectiveLevel() < silence else '--capture=fd']
|
||||
if self.cov:
|
||||
args += self.cov
|
||||
|
||||
errno = pytest.main(args=args)
|
||||
sys.exit(errno)
|
||||
|
||||
def run(self):
|
||||
from pkg_resources import evaluate_marker
|
||||
requirements = self.distribution.install_requires + ['mock>=2.0.0', 'pytest-cov', 'pytest'] +\
|
||||
[v for k, v in self.distribution.extras_require.items() if not k.startswith(':') or evaluate_marker(k[1:])]
|
||||
self.distribution.fetch_build_eggs(requirements)
|
||||
self.run_tests()
|
||||
|
||||
|
||||
def read(fname):
|
||||
with open(os.path.join(__location__, fname)) as fd:
|
||||
return fd.read()
|
||||
|
||||
|
||||
def setup_package():
|
||||
def setup_package(version):
|
||||
# Assemble additional setup commands
|
||||
cmdclass = {'test': PyTest}
|
||||
|
||||
install_requires = []
|
||||
extras_require = {'aws': ['boto'], 'etcd': ['python-etcd'], 'consul': ['python-consul'],
|
||||
'exhibitor': ['kazoo'], 'zookeeper': ['kazoo'], 'kubernetes': ['kubernetes']}
|
||||
|
||||
for r in read('requirements.txt').split('\n'):
|
||||
r = r.strip()
|
||||
if r == '':
|
||||
continue
|
||||
extra = False
|
||||
for e, v in extras_require.items():
|
||||
for e, v in EXTRAS_REQUIRE.items():
|
||||
if r.startswith(v[0]):
|
||||
extras_require[e] = [r]
|
||||
EXTRAS_REQUIRE[e] = [r]
|
||||
extra = True
|
||||
if not extra:
|
||||
install_requires.append(r)
|
||||
|
||||
command_options = {'test': {'test_suite': ('setup.py', 'tests')}}
|
||||
if JUNIT_XML:
|
||||
command_options['test']['junitxml'] = 'setup.py', 'junit.xml'
|
||||
command_options = {'test': {}}
|
||||
if COVERAGE_XML:
|
||||
command_options['test']['cov_xml'] = 'setup.py', True
|
||||
if COVERAGE_HTML:
|
||||
@@ -143,7 +130,7 @@ def setup_package():
|
||||
|
||||
setup(
|
||||
name=NAME,
|
||||
version=VERSION,
|
||||
version=version,
|
||||
url=URL,
|
||||
author=AUTHOR,
|
||||
author_email=AUTHOR_EMAIL,
|
||||
@@ -152,17 +139,28 @@ def setup_package():
|
||||
keywords=KEYWORDS,
|
||||
long_description=read('README.rst'),
|
||||
classifiers=CLASSIFIERS,
|
||||
test_suite='tests',
|
||||
packages=find_packages(exclude=['tests', 'tests.*']),
|
||||
package_data={MAIN_PACKAGE: ["*.json"]},
|
||||
python_requires='>=2.7',
|
||||
install_requires=install_requires,
|
||||
extras_require=extras_require,
|
||||
extras_require=EXTRAS_REQUIRE,
|
||||
setup_requires='flake8',
|
||||
cmdclass=cmdclass,
|
||||
tests_require=['flake8', 'mock>=2.0.0', 'pytest-cov', 'pytest'],
|
||||
command_options=command_options,
|
||||
entry_points={'console_scripts': CONSOLE_SCRIPTS},
|
||||
)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
setup_package()
|
||||
old_modules = sys.modules.copy()
|
||||
try:
|
||||
from patroni import check_psycopg2, fatal, __version__
|
||||
finally:
|
||||
sys.modules.clear()
|
||||
sys.modules.update(old_modules)
|
||||
|
||||
if sys.version_info < (2, 7, 0):
|
||||
fatal('Patroni needs to be run with Python 2.7+')
|
||||
check_psycopg2()
|
||||
|
||||
setup_package(__version__)
|
||||
|
||||
+14
-19
@@ -1,14 +1,12 @@
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import unittest
|
||||
|
||||
from mock import Mock, patch
|
||||
from tempfile import gettempdir
|
||||
|
||||
import psycopg2
|
||||
import requests
|
||||
import urllib3
|
||||
|
||||
from patroni.dcs import Leader, Member
|
||||
from patroni.postgresql import Postgresql
|
||||
@@ -25,19 +23,11 @@ class MockResponse(object):
|
||||
def __init__(self, status_code=200):
|
||||
self.status_code = status_code
|
||||
self.content = '{}'
|
||||
self.ok = True
|
||||
|
||||
def json(self):
|
||||
return json.loads(self.content)
|
||||
|
||||
@property
|
||||
def data(self):
|
||||
return self.content.encode('utf-8')
|
||||
|
||||
@property
|
||||
def text(self):
|
||||
return self.content
|
||||
|
||||
@property
|
||||
def status(self):
|
||||
return self.status_code
|
||||
@@ -52,7 +42,7 @@ def requests_get(url, **kwargs):
|
||||
'"name":"default","clientURLs":["http://localhost:2379","http://localhost:4001"]}]'
|
||||
response = MockResponse()
|
||||
if url.startswith('http://local'):
|
||||
raise requests.exceptions.RequestException()
|
||||
raise urllib3.exceptions.HTTPError()
|
||||
elif ':8011/patroni' in url:
|
||||
response.content = '{"role": "replica", "xlog": {"received_location": 0}, "tags": {}}'
|
||||
elif url.endswith('/members'):
|
||||
@@ -63,11 +53,9 @@ def requests_get(url, **kwargs):
|
||||
data = kwargs.get('data', '')
|
||||
if ' false}' in data:
|
||||
response.status_code = 503
|
||||
response.ok = False
|
||||
response.content = 'restarting after failure already in progress'
|
||||
else:
|
||||
response.status_code = 404
|
||||
response.ok = False
|
||||
return response
|
||||
|
||||
|
||||
@@ -78,6 +66,7 @@ class MockPostmaster(object):
|
||||
self.wait_for_user_backends_to_close = Mock()
|
||||
self.signal_stop = Mock(return_value=None)
|
||||
self.wait = Mock()
|
||||
self.signal_kill = Mock(return_value=False)
|
||||
|
||||
|
||||
class MockCursor(object):
|
||||
@@ -87,6 +76,7 @@ class MockCursor(object):
|
||||
self.closed = False
|
||||
self.rowcount = 0
|
||||
self.results = []
|
||||
self.description = [Mock()]
|
||||
|
||||
def execute(self, sql, *params):
|
||||
if sql.startswith('blabla'):
|
||||
@@ -98,15 +88,18 @@ class MockCursor(object):
|
||||
elif sql.startswith('SELECT slot_name'):
|
||||
self.results = [('blabla', 'physical'), ('foobar', 'physical'), ('ls', 'logical', 'a', 'b')]
|
||||
elif sql.startswith('SELECT CASE WHEN pg_catalog.pg_is_in_recovery()'):
|
||||
self.results = [(1, 2)]
|
||||
self.results = [(1, 2, 1)]
|
||||
elif sql.startswith('SELECT pg_catalog.pg_is_in_recovery()'):
|
||||
self.results = [(False, 2)]
|
||||
elif sql.startswith('WITH replication_info AS ('):
|
||||
elif sql.startswith('SELECT pg_catalog.to_char'):
|
||||
replication_info = '[{"application_name":"walreceiver","client_addr":"1.2.3.4",' +\
|
||||
'"state":"streaming","sync_state":"async","sync_priority":0}]'
|
||||
self.results = [('', 0, '', '', '', '', False, replication_info)]
|
||||
elif sql.startswith('SELECT name, setting'):
|
||||
self.results = [('wal_segment_size', '2048', '8kB', 'integer', 'internal'),
|
||||
('wal_block_size', '8192', None, 'integer', 'internal'),
|
||||
('shared_buffers', '16384', '8kB', 'integer', 'postmaster'),
|
||||
('wal_buffers', '-1', '8kB', 'integer', 'postmaster'),
|
||||
('search_path', 'public', None, 'string', 'user'),
|
||||
('port', '5433', None, 'integer', 'postmaster'),
|
||||
('listen_addresses', '*', None, 'string', 'postmaster'),
|
||||
@@ -173,17 +166,19 @@ class PostgresInit(unittest.TestCase):
|
||||
'search_path': 'public', 'hot_standby': 'on', 'max_wal_senders': 5,
|
||||
'wal_keep_segments': 8, 'wal_log_hints': 'on', 'max_locks_per_transaction': 64,
|
||||
'max_worker_processes': 8, 'max_connections': 100, 'max_prepared_transactions': 0,
|
||||
'track_commit_timestamp': 'off', 'unix_socket_directories': '/tmp', 'trigger_file': 'bla'}
|
||||
'track_commit_timestamp': 'off', 'unix_socket_directories': '/tmp', 'trigger_file': 'bla',
|
||||
'stats_temp_directory': '/tmp'}
|
||||
|
||||
@patch('psycopg2.connect', psycopg2_connect)
|
||||
@patch.object(ConfigHandler, 'write_postgresql_conf', Mock())
|
||||
@patch.object(ConfigHandler, 'replace_pg_hba', Mock())
|
||||
@patch.object(ConfigHandler, 'replace_pg_ident', Mock())
|
||||
@patch.object(Postgresql, 'get_postgres_role_from_data_directory', Mock(return_value='master'))
|
||||
def setUp(self):
|
||||
data_dir = 'data/test0'
|
||||
data_dir = os.path.join('data', 'test0')
|
||||
self.p = Postgresql({'name': 'postgresql0', 'scope': 'batman', 'data_dir': data_dir,
|
||||
'config_dir': data_dir, 'retry_timeout': 10,
|
||||
'krbsrvname': 'postgres', 'pgpass': os.path.join(gettempdir(), 'pgpass0'),
|
||||
'krbsrvname': 'postgres', 'pgpass': os.path.join(data_dir, 'pgpass0'),
|
||||
'listen': '127.0.0.2, 127.0.0.3:5432', 'connect_address': '127.0.0.2:5432',
|
||||
'authentication': {'superuser': {'username': 'foo', 'password': 'test'},
|
||||
'replication': {'username': '', 'password': 'rep-pass'}},
|
||||
|
||||
+26
-6
@@ -12,6 +12,7 @@ from patroni.utils import tzutc
|
||||
from six import BytesIO as IO
|
||||
from six.moves import BaseHTTPServer
|
||||
from . import psycopg2_connect, MockCursor
|
||||
from .test_ha import get_cluster_initialized_without_leader
|
||||
|
||||
|
||||
future_restart_time = datetime.datetime.now(tzutc) + datetime.timedelta(days=5)
|
||||
@@ -145,7 +146,7 @@ class MockRestApiServer(RestApiServer):
|
||||
self.serve_forever = Mock()
|
||||
MockRestApiServer._BaseServer__is_shut_down = Mock()
|
||||
MockRestApiServer._BaseServer__shutdown_request = True
|
||||
config = config or {'listen': '127.0.0.1:8008', 'auth': 'test:test', 'certfile': 'dumb'}
|
||||
config = config or {'listen': '127.0.0.1:8008', 'auth': 'test:test', 'certfile': 'dumb', 'verify_client': 'a'}
|
||||
super(MockRestApiServer, self).__init__(MockPatroni(), config)
|
||||
Handler(MockRequest(request), ('0.0.0.0', 8080), self)
|
||||
|
||||
@@ -159,6 +160,7 @@ class TestRestApiHandler(unittest.TestCase):
|
||||
|
||||
def test_do_GET(self):
|
||||
MockRestApiServer(RestApiHandler, 'GET /replica')
|
||||
MockRestApiServer(RestApiHandler, 'GET /read-only')
|
||||
with patch.object(RestApiHandler, 'get_postgresql_status', Mock(return_value={})):
|
||||
MockRestApiServer(RestApiHandler, 'GET /replica')
|
||||
with patch.object(RestApiHandler, 'get_postgresql_status', Mock(return_value={'role': 'master'})):
|
||||
@@ -197,6 +199,17 @@ class TestRestApiHandler(unittest.TestCase):
|
||||
self.assertIsNotNone(MockRestApiServer(RestApiHandler, 'POST /restart HTTP/1.0'))
|
||||
MockRestApiServer(RestApiHandler, 'POST /restart HTTP/1.0\nAuthorization:')
|
||||
|
||||
@patch.object(MockPatroni, 'dcs')
|
||||
def test_do_GET_cluster(self, mock_dcs):
|
||||
mock_dcs.cluster = get_cluster_initialized_without_leader()
|
||||
mock_dcs.cluster.members[1].data['xlog_location'] = 11
|
||||
self.assertIsNotNone(MockRestApiServer(RestApiHandler, 'GET /cluster'))
|
||||
|
||||
@patch.object(MockPatroni, 'dcs')
|
||||
def test_do_GET_history(self, mock_dcs):
|
||||
mock_dcs.cluster = get_cluster_initialized_without_leader()
|
||||
self.assertIsNotNone(MockRestApiServer(RestApiHandler, 'GET /history'))
|
||||
|
||||
@patch.object(MockPatroni, 'dcs')
|
||||
def test_do_GET_config(self, mock_dcs):
|
||||
mock_dcs.cluster.config.data = {}
|
||||
@@ -233,11 +246,8 @@ class TestRestApiHandler(unittest.TestCase):
|
||||
mock_dcs.get_cluster.return_value.config = ClusterConfig.from_node(1, config)
|
||||
MockRestApiServer(RestApiHandler, request)
|
||||
|
||||
@patch.object(MockPatroni, 'sighup_handler', Mock(side_effect=Exception))
|
||||
@patch.object(MockPatroni, 'sighup_handler', Mock())
|
||||
def test_do_POST_reload(self):
|
||||
with patch.object(MockPatroni, 'config') as mock_config:
|
||||
mock_config.reload_local_configuration.return_value = False
|
||||
MockRestApiServer(RestApiHandler, 'POST /reload HTTP/1.0' + self._authorization)
|
||||
self.assertIsNotNone(MockRestApiServer(RestApiHandler, 'POST /reload HTTP/1.0' + self._authorization))
|
||||
|
||||
@patch.object(MockPatroni, 'dcs')
|
||||
@@ -411,14 +421,24 @@ class TestRestApiServer(unittest.TestCase):
|
||||
def test_reload_config(self):
|
||||
bad_config = {'listen': 'foo'}
|
||||
self.assertRaises(ValueError, MockRestApiServer, None, '', bad_config)
|
||||
srv = MockRestApiServer(lambda a1, a2, a3: None, '')
|
||||
srv = MockRestApiServer(Mock(), '', {'listen': '*:8008', 'certfile': 'a', 'verify_client': 'required'})
|
||||
self.assertRaises(ValueError, srv.reload_config, bad_config)
|
||||
self.assertRaises(ValueError, srv.reload_config, {})
|
||||
with patch.object(socket.socket, 'setsockopt', Mock(side_effect=socket.error)):
|
||||
srv.reload_config({'listen': ':8008'})
|
||||
|
||||
def test_check_auth(self):
|
||||
srv = MockRestApiServer(Mock(), '', {'listen': '*:8008', 'certfile': 'a', 'verify_client': 'required'})
|
||||
mock_rh = Mock()
|
||||
mock_rh.request.getpeercert.return_value = None
|
||||
self.assertIsNot(srv.check_auth(mock_rh), True)
|
||||
|
||||
def test_handle_error(self):
|
||||
try:
|
||||
raise Exception()
|
||||
except Exception:
|
||||
self.assertIsNone(MockRestApiServer.handle_error(None, ('127.0.0.1', 55555)))
|
||||
|
||||
def test_socket_error(self):
|
||||
with patch.object(BaseHTTPServer.HTTPServer, '__init__', Mock(side_effect=socket.error)):
|
||||
self.assertRaises(socket.error, MockRestApiServer, Mock(), '', {'listen': '*:8008'})
|
||||
|
||||
+7
-23
@@ -1,11 +1,11 @@
|
||||
import boto.ec2
|
||||
import sys
|
||||
import unittest
|
||||
import urllib3
|
||||
|
||||
from mock import Mock, patch
|
||||
from collections import namedtuple
|
||||
from patroni.scripts.aws import AWSConnection, main as _main
|
||||
from requests.exceptions import RequestException
|
||||
|
||||
|
||||
class MockEc2Connection(object):
|
||||
@@ -22,28 +22,11 @@ class MockEc2Connection(object):
|
||||
return True
|
||||
|
||||
|
||||
class MockResponse(object):
|
||||
ok = True
|
||||
|
||||
def __init__(self, content):
|
||||
self.content = content
|
||||
|
||||
def json(self):
|
||||
return self.content
|
||||
|
||||
|
||||
def requests_get(url, **kwargs):
|
||||
if url.split('/')[-1] == 'document':
|
||||
result = {"instanceId": "012345", "region": "eu-west-1"}
|
||||
else:
|
||||
result = 'foo'
|
||||
return MockResponse(result)
|
||||
|
||||
|
||||
@patch('boto.ec2.connect_to_region', Mock(return_value=MockEc2Connection()))
|
||||
class TestAWSConnection(unittest.TestCase):
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
@patch('patroni.scripts.aws.requests_get', Mock(return_value=urllib3.HTTPResponse(
|
||||
status=200, body=b'{"instanceId": "012345", "region": "eu-west-1"}')))
|
||||
def setUp(self):
|
||||
self.conn = AWSConnection('test')
|
||||
|
||||
@@ -53,17 +36,18 @@ class TestAWSConnection(unittest.TestCase):
|
||||
self.conn._retry.max_tries = 1
|
||||
self.assertFalse(self.conn.on_role_change('master'))
|
||||
|
||||
@patch('requests.get', Mock(side_effect=RequestException('foo')))
|
||||
@patch('patroni.scripts.aws.requests_get', Mock(side_effect=Exception('foo')))
|
||||
def test_non_aws(self):
|
||||
conn = AWSConnection('test')
|
||||
self.assertFalse(conn.on_role_change("master"))
|
||||
|
||||
@patch('requests.get', Mock(return_value=MockResponse('foo')))
|
||||
@patch('patroni.scripts.aws.requests_get', Mock(return_value=urllib3.HTTPResponse(status=200, body=b'foo')))
|
||||
def test_aws_bizare_response(self):
|
||||
conn = AWSConnection('test')
|
||||
self.assertFalse(conn.aws_available())
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
@patch('patroni.scripts.aws.requests_get', Mock(return_value=urllib3.HTTPResponse(
|
||||
status=200, body=b'{"instanceId": "012345", "region": "eu-west-1"}')))
|
||||
@patch('sys.exit', Mock())
|
||||
def test_main(self):
|
||||
self.assertIsNone(_main())
|
||||
|
||||
+10
-3
@@ -16,6 +16,7 @@ from . import psycopg2_connect, BaseTestPostgresql
|
||||
@patch('os.rename', Mock())
|
||||
class TestBootstrap(BaseTestPostgresql):
|
||||
|
||||
@patch('patroni.postgresql.CallbackExecutor', Mock())
|
||||
def setUp(self):
|
||||
super(TestBootstrap, self).setUp()
|
||||
self.b = self.p.bootstrap
|
||||
@@ -101,6 +102,9 @@ class TestBootstrap(BaseTestPostgresql):
|
||||
@patch.object(CancellableSubprocess, 'call', Mock())
|
||||
@patch.object(Postgresql, 'is_running', Mock(return_value=True))
|
||||
@patch.object(Postgresql, 'data_directory_empty', Mock(return_value=False))
|
||||
@patch.object(Postgresql, 'controldata', Mock(return_value={'max_connections setting': 100,
|
||||
'max_prepared_xacts setting': 0,
|
||||
'max_locks_per_xact setting': 64}))
|
||||
def test_bootstrap(self):
|
||||
with patch('subprocess.call', Mock(return_value=1)):
|
||||
self.assertFalse(self.b.bootstrap({}))
|
||||
@@ -108,7 +112,8 @@ class TestBootstrap(BaseTestPostgresql):
|
||||
config = {'users': {'replicator': {'password': 'rep-pass', 'options': ['replication']}}}
|
||||
|
||||
with patch.object(Postgresql, 'is_running', Mock(return_value=False)),\
|
||||
patch('multiprocessing.Process', Mock(side_effect=Exception)):
|
||||
patch('multiprocessing.Process', Mock(side_effect=Exception)),\
|
||||
patch('multiprocessing.get_context', Mock(side_effect=Exception), create=True):
|
||||
self.assertRaises(Exception, self.b.bootstrap, config)
|
||||
with open(os.path.join(self.p.data_dir, 'pg_hba.conf')) as f:
|
||||
lines = f.readlines()
|
||||
@@ -126,6 +131,7 @@ class TestBootstrap(BaseTestPostgresql):
|
||||
|
||||
@patch.object(CancellableSubprocess, 'call')
|
||||
@patch.object(Postgresql, 'get_major_version', Mock(return_value=90600))
|
||||
@patch.object(Postgresql, 'controldata', Mock(return_value={'Database cluster state': 'in production'}))
|
||||
def test_custom_bootstrap(self, mock_cancellable_subprocess_call):
|
||||
self.p.config._config.pop('pg_hba')
|
||||
config = {'method': 'foo', 'foo': {'command': 'bar'}}
|
||||
@@ -135,6 +141,7 @@ class TestBootstrap(BaseTestPostgresql):
|
||||
|
||||
mock_cancellable_subprocess_call.return_value = 0
|
||||
with patch('multiprocessing.Process', Mock(side_effect=Exception("42"))),\
|
||||
patch('multiprocessing.get_context', Mock(side_effect=Exception("42")), create=True),\
|
||||
patch('os.path.isfile', Mock(return_value=True)),\
|
||||
patch('os.unlink', Mock()),\
|
||||
patch.object(ConfigHandler, 'save_configuration_files', Mock()),\
|
||||
@@ -202,13 +209,13 @@ class TestBootstrap(BaseTestPostgresql):
|
||||
mock_cancellable_subprocess_call.assert_called()
|
||||
args, kwargs = mock_cancellable_subprocess_call.call_args
|
||||
self.assertTrue('PGPASSFILE' in kwargs['env'])
|
||||
self.assertEqual(args[0], ['/bin/false', 'postgres://127.0.0.2:5432/postgres'])
|
||||
self.assertEqual(args[0], ['/bin/false', 'dbname=postgres host=127.0.0.2 port=5432'])
|
||||
|
||||
mock_cancellable_subprocess_call.reset_mock()
|
||||
self.p.config._local_address.pop('host')
|
||||
self.assertTrue(self.b.call_post_bootstrap({'post_init': '/bin/false'}))
|
||||
mock_cancellable_subprocess_call.assert_called()
|
||||
self.assertEqual(mock_cancellable_subprocess_call.call_args[0][0], ['/bin/false', 'postgres://:5432/postgres'])
|
||||
self.assertEqual(mock_cancellable_subprocess_call.call_args[0][0], ['/bin/false', 'dbname=postgres port=5432'])
|
||||
|
||||
mock_cancellable_subprocess_call.side_effect = OSError
|
||||
self.assertFalse(self.b.call_post_bootstrap({'post_init': '/bin/false'}))
|
||||
|
||||
@@ -1,3 +1,4 @@
|
||||
import psutil
|
||||
import unittest
|
||||
|
||||
from mock import Mock, patch
|
||||
@@ -6,22 +7,28 @@ from patroni.postgresql.callback_executor import CallbackExecutor
|
||||
|
||||
class TestCallbackExecutor(unittest.TestCase):
|
||||
|
||||
@patch('subprocess.Popen')
|
||||
@patch('psutil.Popen')
|
||||
def test_callback_executor(self, mock_popen):
|
||||
mock_popen.return_value.wait.side_effect = Exception
|
||||
mock_popen.return_value.poll.return_value = None
|
||||
mock_popen.return_value.children.return_value = []
|
||||
mock_popen.return_value.is_running.return_value = True
|
||||
|
||||
ce = CallbackExecutor()
|
||||
ce._kill_children = Mock(side_effect=Exception)
|
||||
self.assertIsNone(ce.call([]))
|
||||
ce.join()
|
||||
|
||||
self.assertIsNone(ce.call([]))
|
||||
|
||||
mock_popen.return_value.kill.side_effect = OSError
|
||||
mock_popen.return_value.kill.side_effect = psutil.AccessDenied()
|
||||
self.assertIsNone(ce.call([]))
|
||||
|
||||
ce._process_children = []
|
||||
mock_popen.return_value.children.side_effect = psutil.Error()
|
||||
mock_popen.return_value.kill.side_effect = psutil.NoSuchProcess(123)
|
||||
self.assertIsNone(ce.call([]))
|
||||
|
||||
mock_popen.side_effect = Exception
|
||||
ce = CallbackExecutor()
|
||||
ce._callback_event.wait = Mock(side_effect=[None, Exception])
|
||||
ce._condition.wait = Mock(side_effect=[None, Exception])
|
||||
self.assertIsNone(ce.call([]))
|
||||
ce.join()
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
import psutil
|
||||
import unittest
|
||||
|
||||
from mock import Mock, PropertyMock, patch
|
||||
from mock import Mock, patch
|
||||
from patroni.exceptions import PostgresException
|
||||
from patroni.postgresql.cancellable import CancellableSubprocess
|
||||
|
||||
@@ -14,10 +15,20 @@ class TestCancellableSubprocess(unittest.TestCase):
|
||||
self.c.cancel()
|
||||
self.assertRaises(PostgresException, self.c.call, communicate_input=None)
|
||||
|
||||
def test__kill_children(self):
|
||||
self.c._process_children = [Mock()]
|
||||
self.c._kill_children()
|
||||
self.c._process_children[0].kill.side_effect = psutil.AccessDenied()
|
||||
self.c._kill_children()
|
||||
self.c._process_children[0].kill.side_effect = psutil.NoSuchProcess(123)
|
||||
self.c._kill_children()
|
||||
|
||||
@patch('patroni.postgresql.cancellable.polling_loop', Mock(return_value=[0, 0]))
|
||||
def test_cancel(self):
|
||||
self.c._process = Mock()
|
||||
self.c._process.returncode = None
|
||||
self.c._process.is_running.return_value = True
|
||||
self.c._process.children.side_effect = psutil.Error()
|
||||
self.c._process.suspend.side_effect = psutil.Error()
|
||||
self.c.cancel()
|
||||
type(self.c._process).returncode = PropertyMock(side_effect=[None, -15])
|
||||
self.c._process.is_running.side_effect = [True, False]
|
||||
self.c.cancel()
|
||||
|
||||
@@ -15,10 +15,7 @@ class TestConfig(unittest.TestCase):
|
||||
def setUp(self):
|
||||
sys.argv = ['patroni.py']
|
||||
os.environ[Config.PATRONI_CONFIG_VARIABLE] = 'restapi: {}\npostgresql: {data_dir: foo}'
|
||||
self.config = Config()
|
||||
|
||||
def test_no_config(self):
|
||||
self.assertRaises(SystemExit, Config)
|
||||
self.config = Config(None)
|
||||
|
||||
def test_set_dynamic_configuration(self):
|
||||
with patch.object(Config, '_build_effective_configuration', Mock(side_effect=Exception)):
|
||||
@@ -66,12 +63,10 @@ class TestConfig(unittest.TestCase):
|
||||
'PATRONI_admin_PASSWORD': 'admin',
|
||||
'PATRONI_admin_OPTIONS': 'createrole,createdb'
|
||||
})
|
||||
sys.argv = ['patroni.py', 'postgres0.yml']
|
||||
config = Config()
|
||||
config = Config('postgres0.yml')
|
||||
with patch.object(Config, '_load_config_file', Mock(return_value={'restapi': {}})):
|
||||
with patch.object(Config, '_build_effective_configuration', Mock(side_effect=Exception)):
|
||||
self.assertRaises(Exception, config.reload_local_configuration, True)
|
||||
self.assertTrue(config.reload_local_configuration(True))
|
||||
config.reload_local_configuration()
|
||||
self.assertTrue(config.reload_local_configuration())
|
||||
self.assertIsNone(config.reload_local_configuration())
|
||||
|
||||
|
||||
+127
-121
@@ -1,41 +1,38 @@
|
||||
import etcd
|
||||
import os
|
||||
import requests
|
||||
import sys
|
||||
import unittest
|
||||
|
||||
from click.testing import CliRunner
|
||||
from datetime import datetime, timedelta
|
||||
from mock import patch, Mock
|
||||
from patroni.ctl import ctl, store_config, load_config, output_members, request_patroni, get_dcs, parse_dcs, \
|
||||
from patroni.ctl import ctl, store_config, load_config, output_members, get_dcs, parse_dcs, \
|
||||
get_all_members, get_any_member, get_cursor, query_member, configure, PatroniCtlException, apply_config_changes, \
|
||||
format_config_for_editing, show_diff, invoke_editor, format_pg_version, find_executable
|
||||
format_config_for_editing, show_diff, invoke_editor, format_pg_version, find_executable, CONFIG_FILE_PATH
|
||||
from patroni.dcs.etcd import Client, Failover
|
||||
from patroni.utils import tzutc
|
||||
from psycopg2 import OperationalError
|
||||
from urllib3 import PoolManager
|
||||
|
||||
from . import MockConnect, MockCursor, MockResponse, psycopg2_connect, requests_get
|
||||
from . import MockConnect, MockCursor, MockResponse, psycopg2_connect
|
||||
from .test_etcd import etcd_read, socket_getaddrinfo
|
||||
from .test_ha import get_cluster_initialized_without_leader, get_cluster_initialized_with_leader, \
|
||||
get_cluster_initialized_with_only_leader, get_cluster_not_initialized_without_leader, get_cluster, Member
|
||||
|
||||
CONFIG_FILE_PATH = './test-ctl.yaml'
|
||||
|
||||
|
||||
def test_rw_config():
|
||||
runner = CliRunner()
|
||||
with runner.isolated_filesystem():
|
||||
sys.argv = ['patronictl.py', '']
|
||||
load_config(CONFIG_FILE_PATH + '/dummy', None)
|
||||
store_config({'etcd': {'host': 'localhost:2379'}}, CONFIG_FILE_PATH + '/dummy')
|
||||
load_config(CONFIG_FILE_PATH + '/dummy', '0.0.0.0')
|
||||
os.remove(CONFIG_FILE_PATH + '/dummy')
|
||||
os.rmdir(CONFIG_FILE_PATH)
|
||||
load_config(CONFIG_FILE_PATH, None)
|
||||
CONFIG_PATH = './test-ctl.yaml'
|
||||
store_config({'etcd': {'host': 'localhost:2379'}}, CONFIG_PATH + '/dummy')
|
||||
load_config(CONFIG_PATH + '/dummy', '0.0.0.0')
|
||||
os.remove(CONFIG_PATH + '/dummy')
|
||||
os.rmdir(CONFIG_PATH)
|
||||
|
||||
|
||||
@patch('patroni.ctl.load_config',
|
||||
Mock(return_value={'scope': 'alpha', 'postgresql': {'data_dir': '.', 'parameters': {}, 'retry_timeout': 5},
|
||||
'restapi': {'auth': 'u:p', 'listen': ''}, 'etcd': {'host': 'localhost:2379'}}))
|
||||
@patch('patroni.ctl.load_config', Mock(return_value={
|
||||
'scope': 'alpha', 'restapi': {'listen': '::', 'certfile': 'a'}, 'etcd': {'host': 'localhost:2379'},
|
||||
'postgresql': {'data_dir': '.', 'pgpass': './pgpass', 'parameters': {}, 'retry_timeout': 5}}))
|
||||
class TestCtl(unittest.TestCase):
|
||||
|
||||
@patch('socket.getaddrinfo', socket_getaddrinfo)
|
||||
@@ -45,6 +42,12 @@ class TestCtl(unittest.TestCase):
|
||||
self.runner = CliRunner()
|
||||
self.e = get_dcs({'etcd': {'ttl': 30, 'host': 'ok:2379', 'retry_timeout': 10}}, 'foo')
|
||||
|
||||
def test_load_config(self):
|
||||
runner = CliRunner()
|
||||
with runner.isolated_filesystem():
|
||||
self.assertRaises(PatroniCtlException, load_config, './non-existing-config-file', None)
|
||||
self.assertRaises(PatroniCtlException, load_config, './non-existing-config-file', None)
|
||||
|
||||
@patch('psycopg2.connect', psycopg2_connect)
|
||||
def test_get_cursor(self):
|
||||
self.assertIsNone(get_cursor(get_cluster_initialized_without_leader(), {}, role='master'))
|
||||
@@ -69,13 +72,12 @@ class TestCtl(unittest.TestCase):
|
||||
def test_output_members(self):
|
||||
scheduled_at = datetime.now(tzutc) + timedelta(seconds=600)
|
||||
cluster = get_cluster_initialized_with_leader(Failover(1, 'foo', 'bar', scheduled_at))
|
||||
self.assertIsNone(output_members(cluster, name='abc', fmt='pretty'))
|
||||
self.assertIsNone(output_members(cluster, name='abc', fmt='json'))
|
||||
self.assertIsNone(output_members(cluster, name='abc', fmt='yaml'))
|
||||
self.assertIsNone(output_members(cluster, name='abc', fmt='tsv'))
|
||||
del cluster.members[1].data['conn_url']
|
||||
for fmt in ('pretty', 'json', 'yaml', 'tsv'):
|
||||
self.assertIsNone(output_members(cluster, name='abc', fmt=fmt))
|
||||
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
@patch('patroni.ctl.request_patroni', Mock(return_value=MockResponse()))
|
||||
@patch.object(PoolManager, 'request', Mock(return_value=MockResponse()))
|
||||
def test_switchover(self, mock_get_dcs):
|
||||
mock_get_dcs.return_value = self.e
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
@@ -125,18 +127,18 @@ class TestCtl(unittest.TestCase):
|
||||
result = self.runner.invoke(ctl, ['switchover', 'dummy'], input='dummy')
|
||||
assert result.exit_code == 1
|
||||
|
||||
with patch('patroni.ctl.request_patroni', Mock(side_effect=Exception)):
|
||||
with patch.object(PoolManager, 'request', Mock(side_effect=Exception)):
|
||||
# Non-responding patroni
|
||||
result = self.runner.invoke(ctl, ['switchover', 'dummy'], input='leader\nother\n2300-01-01T12:23:00\ny')
|
||||
assert 'falling back to DCS' in result.output
|
||||
|
||||
with patch('patroni.ctl.request_patroni') as mocked:
|
||||
mocked.return_value.status_code = 500
|
||||
with patch.object(PoolManager, 'request') as mocked:
|
||||
mocked.return_value.status = 500
|
||||
result = self.runner.invoke(ctl, ['switchover', 'dummy'], input='leader\nother\n\ny')
|
||||
assert 'Switchover failed' in result.output
|
||||
|
||||
mocked.return_value.status_code = 501
|
||||
mocked.return_value.text = 'Server does not support this operation'
|
||||
mocked.return_value.status = 501
|
||||
mocked.return_value.data = b'Server does not support this operation'
|
||||
result = self.runner.invoke(ctl, ['switchover', 'dummy'], input='leader\nother\n\ny')
|
||||
assert 'Switchover failed' in result.output
|
||||
|
||||
@@ -151,7 +153,7 @@ class TestCtl(unittest.TestCase):
|
||||
assert result.exit_code == 1
|
||||
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
@patch('patroni.ctl.request_patroni', Mock(return_value=MockResponse()))
|
||||
@patch.object(PoolManager, 'request', Mock(return_value=MockResponse()))
|
||||
def test_failover(self, mock_get_dcs):
|
||||
mock_get_dcs.return_value = self.e
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
@@ -159,6 +161,7 @@ class TestCtl(unittest.TestCase):
|
||||
result = self.runner.invoke(ctl, ['failover', 'dummy'], input='\n')
|
||||
assert 'Failover could be performed only to a specific candidate' in result.output
|
||||
|
||||
@patch('patroni.dcs.dcs_modules', Mock(return_value=['patroni.dcs.dummy', 'patroni.dcs.etcd']))
|
||||
def test_get_dcs(self):
|
||||
self.assertRaises(PatroniCtlException, get_dcs, {'dummy': {}}, 'dummy')
|
||||
|
||||
@@ -201,9 +204,6 @@ class TestCtl(unittest.TestCase):
|
||||
rows = query_member(None, None, None, 'master', 'SELECT pg_catalog.pg_is_in_recovery()', {})
|
||||
self.assertTrue('False' in str(rows))
|
||||
|
||||
rows = query_member(None, None, None, 'replica', 'SELECT pg_catalog.pg_is_in_recovery()', {})
|
||||
self.assertEqual(rows, (None, None))
|
||||
|
||||
with patch.object(MockCursor, 'execute', Mock(side_effect=OperationalError('bla'))):
|
||||
rows = query_member(None, None, None, 'replica', 'SELECT pg_catalog.pg_is_in_recovery()', {})
|
||||
|
||||
@@ -231,7 +231,7 @@ class TestCtl(unittest.TestCase):
|
||||
result = self.runner.invoke(ctl, ['dsn', 'alpha', '--member', 'dummy'])
|
||||
assert result.exit_code == 1
|
||||
|
||||
@patch('requests.post')
|
||||
@patch.object(PoolManager, 'request')
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
def test_reload(self, mock_get_dcs, mock_post):
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
@@ -239,18 +239,19 @@ class TestCtl(unittest.TestCase):
|
||||
result = self.runner.invoke(ctl, ['reload', 'alpha'], input='y')
|
||||
assert 'Failed: reload for member' in result.output
|
||||
|
||||
mock_post.return_value.status_code = 200
|
||||
mock_post.return_value.status = 200
|
||||
result = self.runner.invoke(ctl, ['reload', 'alpha'], input='y')
|
||||
assert 'No changes to apply on member' in result.output
|
||||
|
||||
mock_post.return_value.status_code = 202
|
||||
mock_post.return_value.status = 202
|
||||
result = self.runner.invoke(ctl, ['reload', 'alpha'], input='y')
|
||||
assert 'Reload request received for member' in result.output
|
||||
|
||||
@patch('requests.post', requests_get)
|
||||
@patch.object(PoolManager, 'request')
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
def test_restart_reinit(self, mock_get_dcs):
|
||||
def test_restart_reinit(self, mock_get_dcs, mock_post):
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
mock_post.return_value.status = 503
|
||||
result = self.runner.invoke(ctl, ['restart', 'alpha'], input='now\ny\n')
|
||||
assert 'Failed: restart for' in result.output
|
||||
assert result.exit_code == 0
|
||||
@@ -285,58 +286,49 @@ class TestCtl(unittest.TestCase):
|
||||
result = self.runner.invoke(ctl, ['restart', 'alpha', '--pending', '--force', '--timeout', '10min'])
|
||||
assert result.exit_code == 0
|
||||
|
||||
with patch('requests.delete', Mock(return_value=MockResponse(500))):
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
result = self.runner.invoke(ctl, ['restart', 'alpha', 'other', '--force',
|
||||
'--scheduled', '2300-10-01T14:30'])
|
||||
assert 'Failed: flush scheduled restart' in result.output
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
result = self.runner.invoke(ctl, ['restart', 'alpha', 'other', '--force', '--scheduled', '2300-10-01T14:30'])
|
||||
assert 'Failed: flush scheduled restart' in result.output
|
||||
|
||||
with patch('patroni.dcs.Cluster.is_paused', Mock(return_value=True)):
|
||||
result = self.runner.invoke(ctl,
|
||||
['restart', 'alpha', 'other', '--force', '--scheduled', '2300-10-01T14:30'])
|
||||
assert result.exit_code == 1
|
||||
|
||||
with patch('requests.post', Mock(return_value=MockResponse())):
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
result = self.runner.invoke(ctl, ['restart', 'alpha', '--pg-version', '42.0.0',
|
||||
'--scheduled', '2300-10-01T14:30'], input='y')
|
||||
assert result.exit_code == 0
|
||||
|
||||
with patch('requests.post', Mock(return_value=MockResponse(204))):
|
||||
# get restart with the non-200 return code
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
result = self.runner.invoke(ctl, ['restart', 'alpha', '--pg-version', '42.0',
|
||||
'--scheduled', '2300-10-01T14:30'], input='y')
|
||||
assert result.exit_code == 0
|
||||
|
||||
# force restart with restart already present
|
||||
with patch('patroni.ctl.request_patroni', Mock(return_value=MockResponse(204))):
|
||||
result = self.runner.invoke(ctl, ['restart', 'alpha', 'other', '--force',
|
||||
'--scheduled', '2300-10-01T14:30'])
|
||||
assert result.exit_code == 0
|
||||
result = self.runner.invoke(ctl, ['restart', 'alpha', 'other', '--force', '--scheduled', '2300-10-01T14:30'])
|
||||
assert result.exit_code == 0
|
||||
|
||||
with patch('requests.post', Mock(return_value=MockResponse(202))):
|
||||
# get restart with the non-200 return code
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
result = self.runner.invoke(
|
||||
ctl, ['restart', 'alpha', '--pg-version', '99.0.0', '--scheduled', '2300-10-01T14:30'], input='y'
|
||||
)
|
||||
assert 'Success: restart scheduled' in result.output
|
||||
assert result.exit_code == 0
|
||||
ctl_args = ['restart', 'alpha', '--pg-version', '99.0', '--scheduled', '2300-10-01T14:30']
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
mock_post.return_value.status = 200
|
||||
result = self.runner.invoke(ctl, ctl_args, input='y')
|
||||
assert result.exit_code == 0
|
||||
|
||||
with patch('requests.post', Mock(return_value=MockResponse(409))):
|
||||
# get restart with the non-200 return code
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
result = self.runner.invoke(
|
||||
ctl, ['restart', 'alpha', '--pg-version', '99.0.0', '--scheduled', '2300-10-01T14:30'], input='y'
|
||||
)
|
||||
assert 'Failed: another restart is already' in result.output
|
||||
assert result.exit_code == 0
|
||||
# get restart with the non-200 return code
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
mock_post.return_value.status = 204
|
||||
result = self.runner.invoke(ctl, ctl_args, input='y')
|
||||
assert result.exit_code == 0
|
||||
|
||||
# get restart with the non-200 return code
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
mock_post.return_value.status = 202
|
||||
result = self.runner.invoke(ctl, ctl_args, input='y')
|
||||
assert 'Success: restart scheduled' in result.output
|
||||
assert result.exit_code == 0
|
||||
|
||||
# get restart with the non-200 return code
|
||||
# normal restart, the schedule is actually parsed, but not validated in patronictl
|
||||
mock_post.return_value.status = 409
|
||||
result = self.runner.invoke(ctl, ctl_args, input='y')
|
||||
assert 'Failed: another restart is already' in result.output
|
||||
assert result.exit_code == 0
|
||||
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
def test_remove(self, mock_get_dcs):
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
result = self.runner.invoke(ctl, ['remove', 'alpha'], input='alpha\nslave')
|
||||
result = self.runner.invoke(ctl, ['-k', 'remove', 'alpha'], input='alpha\nslave')
|
||||
assert 'Please confirm' in result.output
|
||||
assert 'You are about to remove all' in result.output
|
||||
# Not typing an exact confirmation
|
||||
@@ -353,17 +345,6 @@ class TestCtl(unittest.TestCase):
|
||||
result = self.runner.invoke(ctl, ['remove', 'alpha'], input='alpha\nYes I am aware\nleader')
|
||||
assert result.exit_code == 0
|
||||
|
||||
@patch('requests.post', Mock(side_effect=requests.exceptions.ConnectionError('foo')))
|
||||
@patch('click.get_current_context')
|
||||
def test_request_patroni(self, mock_context):
|
||||
member = get_cluster_initialized_with_leader().leader.member
|
||||
|
||||
mock_context.return_value.obj = {'ctl': {'cacert': 'cert.pem'}}
|
||||
self.assertRaises(requests.exceptions.ConnectionError, request_patroni, member, 'post', 'dummy', {})
|
||||
|
||||
mock_context.return_value.obj = {'ctl': {'insecure': True}}
|
||||
self.assertRaises(requests.exceptions.ConnectionError, request_patroni, member, 'post', 'dummy', {})
|
||||
|
||||
def test_ctl(self):
|
||||
self.runner.invoke(ctl, ['list'])
|
||||
|
||||
@@ -437,7 +418,7 @@ class TestCtl(unittest.TestCase):
|
||||
assert 'Scheduled restart' in result.output
|
||||
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
@patch('requests.delete', Mock(return_value=MockResponse()))
|
||||
@patch.object(PoolManager, 'request', Mock(return_value=MockResponse()))
|
||||
def test_flush(self, mock_get_dcs):
|
||||
mock_get_dcs.return_value = self.e
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
@@ -447,58 +428,58 @@ class TestCtl(unittest.TestCase):
|
||||
|
||||
result = self.runner.invoke(ctl, ['flush', 'dummy', 'restart', '--force'])
|
||||
assert 'Success: flush scheduled restart' in result.output
|
||||
with patch.object(requests, 'delete', return_value=MockResponse(404)):
|
||||
with patch.object(PoolManager, 'request', return_value=MockResponse(404)):
|
||||
result = self.runner.invoke(ctl, ['flush', 'dummy', 'restart', '--force'])
|
||||
assert 'Failed: flush scheduled restart' in result.output
|
||||
|
||||
@patch.object(PoolManager, 'request')
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
@patch('patroni.ctl.polling_loop', Mock(return_value=[1]))
|
||||
def test_pause_cluster(self, mock_get_dcs):
|
||||
def test_pause_cluster(self, mock_get_dcs, mock_post):
|
||||
mock_get_dcs.return_value = self.e
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
|
||||
with patch('requests.patch', Mock(return_value=MockResponse(500))):
|
||||
result = self.runner.invoke(ctl, ['pause', 'dummy'])
|
||||
assert 'Failed' in result.output
|
||||
mock_post.return_value.status = 500
|
||||
result = self.runner.invoke(ctl, ['pause', 'dummy'])
|
||||
assert 'Failed' in result.output
|
||||
|
||||
with patch('requests.patch', Mock(return_value=MockResponse(200))),\
|
||||
patch('patroni.dcs.Cluster.is_paused', Mock(return_value=True)):
|
||||
mock_post.return_value.status = 200
|
||||
with patch('patroni.dcs.Cluster.is_paused', Mock(return_value=True)):
|
||||
result = self.runner.invoke(ctl, ['pause', 'dummy'])
|
||||
assert 'Cluster is already paused' in result.output
|
||||
|
||||
with patch('requests.patch', Mock(return_value=MockResponse(200))):
|
||||
result = self.runner.invoke(ctl, ['pause', 'dummy', '--wait'])
|
||||
assert "'pause' request sent" in result.output
|
||||
mock_get_dcs.return_value.get_cluster = Mock(side_effect=[get_cluster_initialized_with_leader(),
|
||||
get_cluster(None, None, [], None, None)])
|
||||
self.runner.invoke(ctl, ['pause', 'dummy', '--wait'])
|
||||
member = Member(1, 'other', 28, {})
|
||||
mock_get_dcs.return_value.get_cluster = Mock(side_effect=[get_cluster_initialized_with_leader(),
|
||||
get_cluster(None, None, [member], None, None)])
|
||||
self.runner.invoke(ctl, ['pause', 'dummy', '--wait'])
|
||||
result = self.runner.invoke(ctl, ['pause', 'dummy', '--wait'])
|
||||
assert "'pause' request sent" in result.output
|
||||
mock_get_dcs.return_value.get_cluster = Mock(side_effect=[get_cluster_initialized_with_leader(),
|
||||
get_cluster(None, None, [], None, None)])
|
||||
self.runner.invoke(ctl, ['pause', 'dummy', '--wait'])
|
||||
member = Member(1, 'other', 28, {})
|
||||
mock_get_dcs.return_value.get_cluster = Mock(side_effect=[get_cluster_initialized_with_leader(),
|
||||
get_cluster(None, None, [member], None, None)])
|
||||
self.runner.invoke(ctl, ['pause', 'dummy', '--wait'])
|
||||
|
||||
@patch.object(PoolManager, 'request')
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
def test_resume_cluster(self, mock_get_dcs):
|
||||
def test_resume_cluster(self, mock_get_dcs, mock_post):
|
||||
mock_get_dcs.return_value = self.e
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
|
||||
mock_post.return_value.status = 200
|
||||
with patch('patroni.dcs.Cluster.is_paused', Mock(return_value=False)):
|
||||
result = self.runner.invoke(ctl, ['resume', 'dummy'])
|
||||
assert 'Cluster is not paused' in result.output
|
||||
|
||||
with patch('patroni.dcs.Cluster.is_paused', Mock(return_value=True)):
|
||||
with patch('requests.patch', Mock(return_value=MockResponse(200))):
|
||||
result = self.runner.invoke(ctl, ['resume', 'dummy'])
|
||||
assert 'Success' in result.output
|
||||
result = self.runner.invoke(ctl, ['resume', 'dummy'])
|
||||
assert 'Success' in result.output
|
||||
|
||||
with patch('requests.patch', Mock(return_value=MockResponse(500))):
|
||||
result = self.runner.invoke(ctl, ['resume', 'dummy'])
|
||||
assert 'Failed' in result.output
|
||||
mock_post.return_value.status = 500
|
||||
result = self.runner.invoke(ctl, ['resume', 'dummy'])
|
||||
assert 'Failed' in result.output
|
||||
|
||||
with patch('requests.patch', Mock(return_value=MockResponse(200))),\
|
||||
patch('patroni.dcs.Cluster.is_paused', Mock(return_value=False)):
|
||||
result = self.runner.invoke(ctl, ['resume', 'dummy'])
|
||||
assert 'Cluster is not paused' in result.output
|
||||
|
||||
with patch('requests.patch', Mock(side_effect=Exception)):
|
||||
result = self.runner.invoke(ctl, ['resume', 'dummy'])
|
||||
assert 'Can not find accessible cluster member' in result.output
|
||||
mock_post.side_effect = Exception
|
||||
result = self.runner.invoke(ctl, ['resume', 'dummy'])
|
||||
assert 'Can not find accessible cluster member' in result.output
|
||||
|
||||
def test_apply_config_changes(self):
|
||||
config = {"postgresql": {"parameters": {"work_mem": "4MB"}, "use_pg_rewind": True}, "ttl": 30}
|
||||
@@ -574,16 +555,23 @@ class TestCtl(unittest.TestCase):
|
||||
def test_version(self, mock_get_dcs):
|
||||
mock_get_dcs.return_value = self.e
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
with patch('patroni.ctl.request_patroni') as mocked:
|
||||
with patch.object(PoolManager, 'request') as mocked:
|
||||
result = self.runner.invoke(ctl, ['version'])
|
||||
assert 'patronictl version' in result.output
|
||||
mocked.return_value.json = lambda: {'patroni': {'version': '1.2.3'}, 'server_version': 100001}
|
||||
mocked.return_value.data = b'{"patroni":{"version":"1.2.3"},"server_version": 100001}'
|
||||
result = self.runner.invoke(ctl, ['version', 'dummy'])
|
||||
assert '1.2.3' in result.output
|
||||
with patch('requests.get', Mock(side_effect=Exception)):
|
||||
with patch.object(PoolManager, 'request', Mock(side_effect=Exception)):
|
||||
result = self.runner.invoke(ctl, ['version', 'dummy'])
|
||||
assert 'failed to get version' in result.output
|
||||
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
def test_history(self, mock_get_dcs):
|
||||
mock_get_dcs.return_value.get_cluster = Mock()
|
||||
mock_get_dcs.return_value.get_cluster.return_value.history.lines = [[1, 67176, 'no recovery target specified']]
|
||||
result = self.runner.invoke(ctl, ['history'])
|
||||
assert 'Reason' in result.output
|
||||
|
||||
def test_format_pg_version(self):
|
||||
self.assertEqual(format_pg_version(100001), '10.1')
|
||||
self.assertEqual(format_pg_version(90605), '9.6.5')
|
||||
@@ -596,3 +584,21 @@ class TestCtl(unittest.TestCase):
|
||||
self.assertIsNone(find_executable('vim'))
|
||||
with patch('os.path.isfile', Mock(side_effect=[False, True])):
|
||||
self.assertEqual(find_executable('vim', '/'), '/vim.exe')
|
||||
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
def test_get_members(self, mock_get_dcs):
|
||||
mock_get_dcs.return_value = self.e
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_not_initialized_without_leader
|
||||
result = self.runner.invoke(ctl, ['reinit', 'dummy'])
|
||||
assert "cluster doesn\'t have any members" in result.output
|
||||
|
||||
@patch('time.sleep', Mock())
|
||||
@patch('patroni.ctl.get_dcs')
|
||||
def test_reinit_wait(self, mock_get_dcs):
|
||||
mock_get_dcs.return_value.get_cluster = get_cluster_initialized_with_leader
|
||||
with patch.object(PoolManager, 'request') as mocked:
|
||||
mocked.side_effect = [Mock(data=s, status=200) for s in
|
||||
[b"reinitialize", b'{"state":"creating replica"}', b'{"state":"running"}']]
|
||||
result = self.runner.invoke(ctl, ['reinit', 'alpha', 'other', '--wait'], input='y\ny')
|
||||
self.assertIn("Waiting for reinitialize to complete on: other", result.output)
|
||||
self.assertIn("Reinitialize is completed on: other", result.output)
|
||||
|
||||
+30
-14
@@ -7,6 +7,7 @@ from dns.exception import DNSException
|
||||
from mock import Mock, patch
|
||||
from patroni.dcs.etcd import AbstractDCS, Client, Cluster, Etcd, EtcdError, DnsCachingResolver
|
||||
from patroni.exceptions import DCSError
|
||||
from patroni.utils import Retry
|
||||
from urllib3.exceptions import ReadTimeoutError
|
||||
|
||||
from . import SleepException, MockResponse, requests_get
|
||||
@@ -115,12 +116,12 @@ class TestDnsCachingResolver(unittest.TestCase):
|
||||
|
||||
@patch('dns.resolver.query', dns_query)
|
||||
@patch('socket.getaddrinfo', socket_getaddrinfo)
|
||||
@patch('requests.get', requests_get)
|
||||
@patch('patroni.dcs.etcd.requests_get', requests_get)
|
||||
class TestClient(unittest.TestCase):
|
||||
|
||||
@patch('dns.resolver.query', dns_query)
|
||||
@patch('socket.getaddrinfo', socket_getaddrinfo)
|
||||
@patch('requests.get', requests_get)
|
||||
@patch('patroni.dcs.etcd.requests_get', requests_get)
|
||||
def setUp(self):
|
||||
with patch.object(Client, 'machines') as mock_machines:
|
||||
mock_machines.__get__ = Mock(return_value=['http://localhost:2379', 'http://localhost:4001'])
|
||||
@@ -129,12 +130,11 @@ class TestClient(unittest.TestCase):
|
||||
self.client.http.request_encode_body = http_request
|
||||
|
||||
def test_machines(self):
|
||||
self.client._base_uri = 'http://localhost:4001'
|
||||
self.client._machines_cache = ['http://localhost:2379']
|
||||
self.client._base_uri = 'http://localhost:4002'
|
||||
self.client._machines_cache = ['http://localhost:4002', 'http://localhost:2379']
|
||||
self.assertIsNotNone(self.client.machines)
|
||||
self.client._base_uri = 'http://localhost:4001'
|
||||
self.client._machines_cache = []
|
||||
self.assertIsNotNone(self.client.machines)
|
||||
self.client._machines_cache = ['http://localhost:4001']
|
||||
self.client._update_machines_cache = True
|
||||
machines = None
|
||||
try:
|
||||
@@ -145,19 +145,31 @@ class TestClient(unittest.TestCase):
|
||||
|
||||
@patch.object(Client, 'machines')
|
||||
def test_api_execute(self, mock_machines):
|
||||
mock_machines.__get__ = Mock(return_value=['http://localhost:2379'])
|
||||
mock_machines.__get__ = Mock(return_value=['http://localhost:4001', 'http://localhost:2379'])
|
||||
self.assertRaises(ValueError, self.client.api_execute, '', '')
|
||||
self.client._base_uri = 'http://localhost:4001'
|
||||
self.client._machines_cache = ['http://localhost:2379']
|
||||
self.client.api_execute('/', 'POST', timeout=0)
|
||||
mock_machines.__get__ = Mock(return_value=['http://localhost:2379'])
|
||||
self.assertRaises(etcd.EtcdException, self.client.api_execute, '/', 'POST', timeout=0)
|
||||
self.client._base_uri = 'http://localhost:4001'
|
||||
rtry = Retry(deadline=10, max_delay=1, max_tries=-1, retry_exceptions=(etcd.EtcdLeaderElectionInProgress,))
|
||||
rtry(self.client.api_execute, '/', 'POST', timeout=0, params={'retry': rtry})
|
||||
self.client._machines_cache_updated = 0
|
||||
self.client.api_execute('/', 'POST', timeout=0)
|
||||
self.client._machines_cache = [self.client._base_uri]
|
||||
self.assertRaises(etcd.EtcdWatchTimedOut, self.client.api_execute, '/timeout', 'POST', params={'wait': 'true'})
|
||||
self.assertRaises(etcd.EtcdWatchTimedOut, self.client.api_execute, '/timeout', 'POST', params={'wait': 'true'})
|
||||
self.assertRaises(etcd.EtcdException, self.client.api_execute, '/', '')
|
||||
with patch.object(Client, '_load_machines_cache', Mock(side_effect=etcd.EtcdException)):
|
||||
|
||||
with patch.object(Client, '_calculate_timeouts', Mock(side_effect=[(1, 1, 0), (1, 1, 0), (0, 1, 0)])),\
|
||||
patch.object(Client, '_load_machines_cache', Mock(side_effect=Exception)):
|
||||
self.client.http.request = Mock(side_effect=socket.error)
|
||||
self.assertRaises(etcd.EtcdException, rtry, self.client.api_execute, '/', 'GET', params={'retry': rtry})
|
||||
|
||||
with patch.object(Client, '_calculate_timeouts', Mock(side_effect=[(1, 1, 0), (1, 1, 0), (0, 1, 0)])),\
|
||||
patch.object(Client, '_load_machines_cache', Mock(return_value=True)):
|
||||
self.assertRaises(etcd.EtcdException, rtry, self.client.api_execute, '/', 'GET', params={'retry': rtry})
|
||||
|
||||
with patch.object(Client, '_do_http_request', Mock(side_effect=etcd.EtcdException)):
|
||||
self.client._read_timeout = 0.01
|
||||
self.assertRaises(etcd.EtcdException, self.client.api_execute, '/', 'GET')
|
||||
|
||||
def test_get_srv_record(self):
|
||||
@@ -173,8 +185,12 @@ class TestClient(unittest.TestCase):
|
||||
self.client._get_machines_cache_from_dns('error', 2379)
|
||||
|
||||
@patch.object(Client, 'machines')
|
||||
def test__load_machines_cache(self, mock_machines):
|
||||
mock_machines.__get__ = Mock(return_value=['http://localhost:2379'])
|
||||
def test__refresh_machines_cache(self, mock_machines):
|
||||
mock_machines.__get__ = Mock(side_effect=etcd.EtcdConnectionFailed)
|
||||
self.assertIsNone(self.client._refresh_machines_cache())
|
||||
self.assertRaises(etcd.EtcdException, self.client._refresh_machines_cache, True)
|
||||
|
||||
def test__load_machines_cache(self):
|
||||
self.client._config = {}
|
||||
self.assertRaises(Exception, self.client._load_machines_cache)
|
||||
self.client._config = {'srv': 'blabla'}
|
||||
@@ -190,7 +206,7 @@ class TestClient(unittest.TestCase):
|
||||
socket_options=[(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)])
|
||||
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
@patch('patroni.dcs.etcd.requests_get', requests_get)
|
||||
@patch('socket.getaddrinfo', socket_getaddrinfo)
|
||||
@patch.object(etcd.Client, 'write', etcd_write)
|
||||
@patch.object(etcd.Client, 'read', etcd_read)
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
import unittest
|
||||
import urllib3
|
||||
|
||||
from mock import Mock, patch
|
||||
from patroni.dcs.exhibitor import ExhibitorEnsembleProvider, Exhibitor
|
||||
@@ -8,7 +9,7 @@ from . import SleepException, requests_get
|
||||
from .test_zookeeper import MockKazooClient
|
||||
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
@patch('patroni.dcs.exhibitor.requests_get', requests_get)
|
||||
@patch('time.sleep', Mock(side_effect=SleepException))
|
||||
class TestExhibitorEnsembleProvider(unittest.TestCase):
|
||||
|
||||
@@ -21,7 +22,8 @@ class TestExhibitorEnsembleProvider(unittest.TestCase):
|
||||
|
||||
class TestExhibitor(unittest.TestCase):
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
@patch('urllib3.PoolManager.request', Mock(return_value=urllib3.HTTPResponse(
|
||||
status=200, body=b'{"servers":["127.0.0.1","127.0.0.2","127.0.0.3"],"port":2181}')))
|
||||
@patch('patroni.dcs.zookeeper.KazooClient', MockKazooClient)
|
||||
def setUp(self):
|
||||
self.e = Exhibitor({'hosts': ['localhost', 'exhibitor'], 'port': 8181, 'scope': 'test',
|
||||
|
||||
+38
-19
@@ -122,7 +122,7 @@ zookeeper:
|
||||
# all the extra values that are coming from py.test
|
||||
sys.argv = sys.argv[:1]
|
||||
|
||||
self.config = Config()
|
||||
self.config = Config(None)
|
||||
self.config.set_dynamic_configuration({'maximum_lag_on_failover': 5})
|
||||
self.version = '1.5.7'
|
||||
self.postgresql = p
|
||||
@@ -137,23 +137,29 @@ zookeeper:
|
||||
self.scheduled_restart = {'schedule': future_restart_time,
|
||||
'postmaster_start_time': str(postmaster_start_time)}
|
||||
self.watchdog = Watchdog(self.config)
|
||||
self.request = lambda member, **kwargs: requests_get(member.api_url, **kwargs)
|
||||
|
||||
|
||||
def run_async(self, func, args=()):
|
||||
return func(*args) if args else func()
|
||||
self.reset_scheduled_action()
|
||||
if args:
|
||||
func(*args)
|
||||
else:
|
||||
func()
|
||||
|
||||
|
||||
@patch.object(Postgresql, 'is_running', Mock(return_value=MockPostmaster()))
|
||||
@patch.object(Postgresql, 'is_leader', Mock(return_value=True))
|
||||
@patch.object(Postgresql, 'timeline_wal_position', Mock(return_value=(1, 10)))
|
||||
@patch.object(Postgresql, 'timeline_wal_position', Mock(return_value=(1, 10, 1)))
|
||||
@patch.object(Postgresql, '_cluster_info_state_get', Mock(return_value=3))
|
||||
@patch.object(Postgresql, 'call_nowait', Mock(return_value=True))
|
||||
@patch.object(Postgresql, 'data_directory_empty', Mock(return_value=False))
|
||||
@patch.object(Postgresql, 'controldata', Mock(return_value={'Database system identifier': SYSID}))
|
||||
@patch.object(SlotsHandler, 'sync_replication_slots', Mock())
|
||||
@patch.object(ConfigHandler, 'append_pg_hba', Mock())
|
||||
@patch.object(Postgresql, 'write_pgpass', Mock(return_value={}))
|
||||
@patch.object(ConfigHandler, 'write_pgpass', Mock(return_value={}))
|
||||
@patch.object(ConfigHandler, 'write_recovery_conf', Mock())
|
||||
@patch.object(ConfigHandler, 'write_postgresql_conf', Mock())
|
||||
@patch.object(Postgresql, 'query', Mock())
|
||||
@patch.object(Postgresql, 'checkpoint', Mock())
|
||||
@patch.object(CancellableSubprocess, 'call', Mock(return_value=0))
|
||||
@@ -166,12 +172,13 @@ def run_async(self, func, args=()):
|
||||
@patch('patroni.postgresql.polling_loop', Mock(return_value=range(1)))
|
||||
@patch('patroni.async_executor.AsyncExecutor.busy', PropertyMock(return_value=False))
|
||||
@patch('patroni.async_executor.AsyncExecutor.run_async', run_async)
|
||||
@patch('patroni.postgresql.rewind.Thread', Mock())
|
||||
@patch('subprocess.call', Mock(return_value=0))
|
||||
@patch('time.sleep', Mock())
|
||||
class TestHa(PostgresInit):
|
||||
|
||||
@patch('socket.getaddrinfo', socket_getaddrinfo)
|
||||
@patch('patroni.dcs.dcs_modules', Mock(return_value=['patroni.dcs.foo', 'patroni.dcs.etcd']))
|
||||
@patch('patroni.dcs.dcs_modules', Mock(return_value=['patroni.dcs.etcd']))
|
||||
@patch.object(etcd.Client, 'read', etcd_read)
|
||||
def setUp(self):
|
||||
super(TestHa, self).setUp()
|
||||
@@ -193,9 +200,12 @@ class TestHa(PostgresInit):
|
||||
self.assertTrue(self.ha.update_lock(True))
|
||||
|
||||
def test_touch_member(self):
|
||||
self.p.timeline_wal_position = Mock(return_value=(0, 1))
|
||||
self.p.timeline_wal_position = Mock(return_value=(0, 1, 0))
|
||||
self.p.replica_cached_timeline = Mock(side_effect=Exception)
|
||||
self.ha.touch_member()
|
||||
self.p.timeline_wal_position = Mock(return_value=(0, 1, 1))
|
||||
self.p.set_role('standby_leader')
|
||||
self.ha.touch_member()
|
||||
|
||||
def test_is_leader(self):
|
||||
self.assertFalse(self.ha.is_leader())
|
||||
@@ -353,6 +363,7 @@ class TestHa(PostgresInit):
|
||||
self.p.is_leader = false
|
||||
self.assertEqual(self.ha.run_cycle(), 'no action. i am a secondary and i am following a leader')
|
||||
self.ha.patroni.replicatefrom = "foo"
|
||||
self.p.config.check_recovery_conf = Mock(return_value=(True, False))
|
||||
self.assertEqual(self.ha.run_cycle(), 'no action. i am a secondary and i am following a leader')
|
||||
|
||||
def test_follow_in_pause(self):
|
||||
@@ -425,7 +436,7 @@ class TestHa(PostgresInit):
|
||||
|
||||
self.ha.cluster = get_cluster_initialized_with_leader()
|
||||
self.assertIsNone(self.ha.reinitialize(True))
|
||||
|
||||
self.ha._async_executor.schedule('reinitialize')
|
||||
self.assertIsNotNone(self.ha.reinitialize())
|
||||
|
||||
self.ha.state_handler.name = self.ha.cluster.leader.name
|
||||
@@ -439,7 +450,7 @@ class TestHa(PostgresInit):
|
||||
self.p.restart = false
|
||||
self.assertEqual(self.ha.restart({}), (False, 'restart failed'))
|
||||
self.ha.cluster = get_cluster_initialized_with_leader()
|
||||
self.ha.reinitialize()
|
||||
self.ha._async_executor.schedule('reinitialize')
|
||||
self.assertEqual(self.ha.restart({}), (False, 'reinitialize already in progress'))
|
||||
with patch.object(self.ha, "restart_matches", return_value=False):
|
||||
self.assertEqual(self.ha.restart({'foo': 'bar'}), (False, "restart conditions are not satisfied"))
|
||||
@@ -447,7 +458,7 @@ class TestHa(PostgresInit):
|
||||
@patch('os.kill', Mock())
|
||||
def test_restart_in_progress(self):
|
||||
with patch('patroni.async_executor.AsyncExecutor.busy', PropertyMock(return_value=True)):
|
||||
self.ha.restart({}, run_async=True)
|
||||
self.ha._async_executor.schedule('restart')
|
||||
self.assertTrue(self.ha.restart_scheduled())
|
||||
self.assertEqual(self.ha.run_cycle(), 'restart in progress')
|
||||
|
||||
@@ -464,7 +475,6 @@ class TestHa(PostgresInit):
|
||||
self.assertEqual(self.ha.run_cycle(), 'lost leader lock during restart')
|
||||
mock_terminate.assert_called()
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
def test_manual_failover_from_leader(self):
|
||||
self.ha.fetch_node_status = get_node_status()
|
||||
self.ha.has_lock = true
|
||||
@@ -512,7 +522,6 @@ class TestHa(PostgresInit):
|
||||
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, 'blabla', self.p.name, scheduled))
|
||||
self.assertEqual('no action. i am the leader with the lock', self.ha.run_cycle())
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
def test_manual_failover_from_leader_in_pause(self):
|
||||
self.ha.has_lock = true
|
||||
self.ha.is_paused = true
|
||||
@@ -522,7 +531,6 @@ class TestHa(PostgresInit):
|
||||
self.ha.cluster = get_cluster_initialized_with_leader(Failover(0, self.p.name, '', None))
|
||||
self.assertEqual('PAUSE: no action. i am the leader with the lock', self.ha.run_cycle())
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
def test_manual_failover_from_leader_in_synchronous_mode(self):
|
||||
self.p.is_leader = true
|
||||
self.ha.has_lock = true
|
||||
@@ -535,7 +543,6 @@ class TestHa(PostgresInit):
|
||||
self.ha.is_failover_possible = true
|
||||
self.assertEqual('manual failover: demoting myself', self.ha.run_cycle())
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
def test_manual_failover_process_no_leader(self):
|
||||
self.p.is_leader = false
|
||||
self.ha.cluster = get_cluster_initialized_without_leader(failover=Failover(0, '', self.p.name, None))
|
||||
@@ -585,7 +592,6 @@ class TestHa(PostgresInit):
|
||||
self.ha.is_paused = true
|
||||
self.assertFalse(self.ha.is_healthiest_node())
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
def test__is_healthiest_node(self):
|
||||
self.ha.cluster = get_cluster_initialized_without_leader(sync=('postgresql1', self.p.name))
|
||||
self.assertTrue(self.ha._is_healthiest_node(self.ha.old_cluster.members))
|
||||
@@ -599,7 +605,7 @@ class TestHa(PostgresInit):
|
||||
# in synchronous_mode consider itself healthy if the former leader is accessible in read-only and ahead of us
|
||||
with patch.object(Ha, 'is_synchronous_mode', Mock(return_value=True)):
|
||||
self.assertTrue(self.ha._is_healthiest_node(self.ha.old_cluster.members))
|
||||
with patch('patroni.postgresql.Postgresql.timeline_wal_position', return_value=(1, 1)):
|
||||
with patch('patroni.postgresql.Postgresql.timeline_wal_position', return_value=(1, 1, 1)):
|
||||
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
|
||||
with patch('patroni.postgresql.Postgresql.replica_cached_timeline', return_value=1):
|
||||
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
|
||||
@@ -607,19 +613,23 @@ class TestHa(PostgresInit):
|
||||
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
|
||||
self.ha.patroni.nofailover = False
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
def test_fetch_node_status(self):
|
||||
member = Member(0, 'test', 1, {'api_url': 'http://127.0.0.1:8011/patroni'})
|
||||
self.ha.fetch_node_status(member)
|
||||
member = Member(0, 'test', 1, {'api_url': 'http://localhost:8011/patroni'})
|
||||
self.ha.fetch_node_status(member)
|
||||
|
||||
@patch.object(Rewind, 'pg_rewind', true)
|
||||
@patch.object(Rewind, 'check_leader_is_not_in_recovery', true)
|
||||
def test_post_recover(self):
|
||||
self.p.is_running = false
|
||||
self.ha.has_lock = true
|
||||
self.p.set_role('master')
|
||||
self.assertEqual(self.ha.post_recover(), 'removed leader key after trying and failing to start postgres')
|
||||
self.ha.has_lock = false
|
||||
self.assertEqual(self.ha.post_recover(), 'failed to start postgres')
|
||||
leader = Leader(0, 0, Member(0, 'l', 2, {"version": "1.6", "conn_url": "postgres://a", "role": "master"}))
|
||||
self.ha._rewind.execute(leader)
|
||||
self.p.is_running = true
|
||||
self.assertIsNone(self.ha.post_recover())
|
||||
|
||||
@@ -684,7 +694,7 @@ class TestHa(PostgresInit):
|
||||
self.p.is_leader = false
|
||||
self.p.name = 'leader'
|
||||
self.ha.cluster = get_standby_cluster_initialized_with_only_leader()
|
||||
self.p.config.check_recovery_conf = true
|
||||
self.p.config.check_recovery_conf = Mock(return_value=(False, False))
|
||||
self.assertEqual(self.ha.run_cycle(), 'promoted self to a standby leader because i had the session lock')
|
||||
self.assertEqual(self.ha.run_cycle(), 'no action. i am the standby leader with the lock')
|
||||
|
||||
@@ -696,7 +706,6 @@ class TestHa(PostgresInit):
|
||||
with patch.object(Leader, 'conn_url', PropertyMock(return_value='')):
|
||||
self.assertEqual(self.ha.run_cycle(), 'continue following the old known standby leader')
|
||||
|
||||
@patch('requests.get', requests_get)
|
||||
def test_process_unhealthy_standby_cluster_as_standby_leader(self):
|
||||
self.p.is_leader = false
|
||||
self.p.name = 'leader'
|
||||
@@ -809,6 +818,17 @@ class TestHa(PostgresInit):
|
||||
self.assertEqual(self.ha.run_cycle(), 'stopped PostgreSQL to fail over after a crash')
|
||||
demote.assert_called_once()
|
||||
|
||||
def test_master_stop_timeout(self):
|
||||
self.assertEqual(self.ha.master_stop_timeout(), None)
|
||||
self.ha.patroni.config.set_dynamic_configuration({'master_stop_timeout': 30})
|
||||
with patch.object(Ha, 'is_synchronous_mode', Mock(return_value=True)):
|
||||
self.assertEqual(self.ha.master_stop_timeout(), 30)
|
||||
self.ha.patroni.config.set_dynamic_configuration({'master_stop_timeout': 30})
|
||||
with patch.object(Ha, 'is_synchronous_mode', Mock(return_value=False)):
|
||||
self.assertEqual(self.ha.master_stop_timeout(), None)
|
||||
self.ha.patroni.config.set_dynamic_configuration({'master_stop_timeout': None})
|
||||
self.assertEqual(self.ha.master_stop_timeout(), None)
|
||||
|
||||
@patch('patroni.postgresql.Postgresql.follow')
|
||||
def test_demote_immediate(self, follow):
|
||||
self.ha.has_lock = true
|
||||
@@ -1027,7 +1047,6 @@ class TestHa(PostgresInit):
|
||||
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
|
||||
|
||||
@patch('sys.exit', return_value=1)
|
||||
@patch('requests.get', requests_get)
|
||||
def test_abort_join(self, exit_mock):
|
||||
self.ha.cluster = get_cluster_not_initialized_without_leader()
|
||||
self.p.is_leader = false
|
||||
|
||||
+79
-17
@@ -1,8 +1,11 @@
|
||||
import json
|
||||
import time
|
||||
import unittest
|
||||
|
||||
from mock import Mock, patch
|
||||
from patroni.dcs.kubernetes import Kubernetes, KubernetesError, k8s_client, k8s_watch, RetryFailedError
|
||||
from patroni.dcs.kubernetes import Kubernetes, KubernetesError, k8s_client, RetryFailedError
|
||||
from threading import Thread
|
||||
from . import SleepException
|
||||
|
||||
|
||||
def mock_list_namespaced_config_map(self, *args, **kwargs):
|
||||
@@ -16,46 +19,73 @@ def mock_list_namespaced_config_map(self, *args, **kwargs):
|
||||
metadata.update({'name': 'test-sync', 'annotations': {'leader': 'p-0'}})
|
||||
items.append(k8s_client.V1ConfigMap(metadata=k8s_client.V1ObjectMeta(**metadata)))
|
||||
metadata = k8s_client.V1ObjectMeta(resource_version='1')
|
||||
return k8s_client.V1ConfigMapList(metadata=metadata, items=items)
|
||||
return k8s_client.V1ConfigMapList(metadata=metadata, items=items, kind='ConfigMapList')
|
||||
|
||||
|
||||
def mock_list_namespaced_pod(self, *args, **kwargs):
|
||||
metadata = k8s_client.V1ObjectMeta(resource_version='1', name='p-0', annotations={'status': '{}'})
|
||||
items = [k8s_client.V1Pod(metadata=metadata)]
|
||||
return k8s_client.V1PodList(items=items)
|
||||
return k8s_client.V1PodList(items=items, kind='PodList')
|
||||
|
||||
|
||||
@patch.object(k8s_client.CoreV1Api, 'patch_namespaced_config_map', Mock())
|
||||
@patch.object(k8s_client.CoreV1Api, 'create_namespaced_config_map', Mock())
|
||||
def mock_config_map(*args, **kwargs):
|
||||
mock = Mock()
|
||||
mock.metadata.resource_version = '2'
|
||||
return mock
|
||||
|
||||
|
||||
@patch('socket.TCP_KEEPIDLE', 4, create=True)
|
||||
@patch('socket.TCP_KEEPINTVL', 5, create=True)
|
||||
@patch('socket.TCP_KEEPCNT', 6, create=True)
|
||||
@patch.object(k8s_client.CoreV1Api, 'patch_namespaced_config_map', mock_config_map)
|
||||
@patch.object(k8s_client.CoreV1Api, 'create_namespaced_config_map', mock_config_map)
|
||||
@patch('kubernetes.client.api_client.ThreadPool', Mock(), create=True)
|
||||
@patch.object(Thread, 'start', Mock())
|
||||
class TestKubernetes(unittest.TestCase):
|
||||
|
||||
@patch('socket.TCP_KEEPIDLE', 4, create=True)
|
||||
@patch('socket.TCP_KEEPINTVL', 5, create=True)
|
||||
@patch('socket.TCP_KEEPCNT', 6, create=True)
|
||||
@patch('kubernetes.config.load_kube_config', Mock())
|
||||
@patch.object(k8s_client.CoreV1Api, 'list_namespaced_config_map', mock_list_namespaced_config_map)
|
||||
@patch.object(k8s_client.CoreV1Api, 'list_namespaced_pod', mock_list_namespaced_pod)
|
||||
@patch('kubernetes.client.api_client.ThreadPool', Mock(), create=True)
|
||||
@patch.object(Thread, 'start', Mock())
|
||||
def setUp(self):
|
||||
self.k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'retry_timeout': 10, 'labels': {'f': 'b'}})
|
||||
self.k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0',
|
||||
'loop_wait': 10, 'retry_timeout': 10, 'labels': {'f': 'b'}})
|
||||
self.assertRaises(AttributeError, self.k._pods._build_cache)
|
||||
self.k._pods._is_ready = True
|
||||
self.assertRaises(AttributeError, self.k._kinds._build_cache)
|
||||
self.k._kinds._is_ready = True
|
||||
self.k.get_cluster()
|
||||
|
||||
@patch('time.time', Mock(side_effect=[1, 10.9, 100]))
|
||||
def test__wait_caches(self):
|
||||
self.k._pods._is_ready = False
|
||||
with self.k._condition:
|
||||
self.assertRaises(RetryFailedError, self.k._wait_caches)
|
||||
|
||||
def test_get_cluster(self):
|
||||
with patch.object(k8s_client.CoreV1Api, 'list_namespaced_config_map', mock_list_namespaced_config_map), \
|
||||
patch.object(k8s_client.CoreV1Api, 'list_namespaced_pod', mock_list_namespaced_pod), \
|
||||
patch('time.time', Mock(return_value=time.time() + 31)):
|
||||
self.k.get_cluster()
|
||||
|
||||
with patch.object(k8s_client.CoreV1Api, 'list_namespaced_pod', Mock(side_effect=Exception)):
|
||||
with patch.object(Kubernetes, '_wait_caches', Mock(side_effect=Exception)):
|
||||
self.assertRaises(KubernetesError, self.k.get_cluster)
|
||||
|
||||
@patch('kubernetes.config.load_kube_config', Mock())
|
||||
@patch.object(k8s_client.CoreV1Api, 'create_namespaced_endpoints', Mock())
|
||||
def test_update_leader(self):
|
||||
k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'retry_timeout': 10,
|
||||
k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'loop_wait': 10, 'retry_timeout': 10,
|
||||
'labels': {'f': 'b'}, 'use_endpoints': True, 'pod_ip': '10.0.0.0'})
|
||||
self.assertIsNotNone(k.update_leader('123'))
|
||||
|
||||
@patch('kubernetes.config.load_kube_config', Mock())
|
||||
@patch.object(k8s_client.CoreV1Api, 'create_namespaced_endpoints', Mock())
|
||||
def test_update_leader_with_restricted_access(self):
|
||||
k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'retry_timeout': 10,
|
||||
k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'loop_wait': 10, 'retry_timeout': 10,
|
||||
'labels': {'f': 'b'}, 'use_endpoints': True, 'pod_ip': '10.0.0.0'})
|
||||
self.assertIsNotNone(k.update_leader('123', True))
|
||||
|
||||
@@ -97,7 +127,7 @@ class TestKubernetes(unittest.TestCase):
|
||||
@patch.object(k8s_client.CoreV1Api, 'create_namespaced_endpoints',
|
||||
Mock(side_effect=[k8s_client.rest.ApiException(502, ''), k8s_client.rest.ApiException(500, '')]))
|
||||
def test_delete_sync_state(self):
|
||||
k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'retry_timeout': 10,
|
||||
k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'loop_wait': 10, 'retry_timeout': 10,
|
||||
'labels': {'f': 'b'}, 'use_endpoints': True, 'pod_ip': '10.0.0.0'})
|
||||
self.assertFalse(k.delete_sync_state())
|
||||
|
||||
@@ -105,24 +135,56 @@ class TestKubernetes(unittest.TestCase):
|
||||
self.k.set_ttl(10)
|
||||
self.k.watch(None, 0)
|
||||
self.k.watch(None, 0)
|
||||
with patch.object(k8s_watch.Watch, 'stream',
|
||||
Mock(side_effect=[Exception, [], KeyboardInterrupt,
|
||||
[{'raw_object': {'metadata': {'resourceVersion': '2'}}}]])):
|
||||
self.assertFalse(self.k.watch('1', 2))
|
||||
self.assertRaises(KeyboardInterrupt, self.k.watch, '1', 2)
|
||||
self.assertTrue(self.k.watch('1', 2))
|
||||
|
||||
def test_set_history_value(self):
|
||||
self.k.set_history_value('{}')
|
||||
|
||||
@patch('kubernetes.config.load_kube_config', Mock())
|
||||
@patch('patroni.dcs.kubernetes.ObjectCache', Mock())
|
||||
@patch.object(k8s_client.CoreV1Api, 'patch_namespaced_pod', Mock(return_value=True))
|
||||
@patch.object(k8s_client.CoreV1Api, 'create_namespaced_endpoints', Mock())
|
||||
@patch.object(k8s_client.CoreV1Api, 'create_namespaced_service',
|
||||
Mock(side_effect=[True, False, k8s_client.rest.ApiException(500, '')]))
|
||||
def test__create_config_service(self):
|
||||
k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'retry_timeout': 10,
|
||||
k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0', 'loop_wait': 10, 'retry_timeout': 10,
|
||||
'labels': {'f': 'b'}, 'use_endpoints': True, 'pod_ip': '10.0.0.0'})
|
||||
self.assertIsNotNone(k.patch_or_create_config({'foo': 'bar'}))
|
||||
self.assertIsNotNone(k.patch_or_create_config({'foo': 'bar'}))
|
||||
k.touch_member({'state': 'running', 'role': 'replica'})
|
||||
|
||||
|
||||
class TestCacheBuilder(unittest.TestCase):
|
||||
|
||||
@patch('socket.TCP_KEEPIDLE', 4, create=True)
|
||||
@patch('socket.TCP_KEEPINTVL', 5, create=True)
|
||||
@patch('socket.TCP_KEEPCNT', 6, create=True)
|
||||
@patch('kubernetes.config.load_kube_config', Mock())
|
||||
@patch('kubernetes.client.api_client.ThreadPool', Mock(), create=True)
|
||||
@patch.object(Thread, 'start', Mock())
|
||||
def setUp(self):
|
||||
self.k = Kubernetes({'ttl': 30, 'scope': 'test', 'name': 'p-0',
|
||||
'loop_wait': 10, 'retry_timeout': 10, 'labels': {'f': 'b'}})
|
||||
|
||||
@patch.object(k8s_client.CoreV1Api, 'list_namespaced_config_map', mock_list_namespaced_config_map)
|
||||
@patch('patroni.dcs.kubernetes.ObjectCache._watch')
|
||||
def test__build_cache(self, mock_response):
|
||||
mock_response.return_value.read_chunked.return_value = [json.dumps(
|
||||
{'type': 'MODIFIED', 'object': {'metadata': {
|
||||
'name': self.k.config_path, 'resourceVersion': '2', 'annotations': {self.k._CONFIG: 'foo'}}}}
|
||||
).encode('utf-8'), ('\n' + json.dumps(
|
||||
{'type': 'DELETED', 'object': {'metadata': {
|
||||
'name': self.k.config_path, 'resourceVersion': '3'}}}
|
||||
) + '\n' + json.dumps(
|
||||
{'type': 'MDIFIED', 'object': {'metadata': {'name': self.k.config_path}}}
|
||||
) + '\n' + json.dumps({'object': {'code': 410}}) + '\n').encode('utf-8')]
|
||||
self.k._kinds._build_cache()
|
||||
|
||||
@patch('patroni.dcs.kubernetes.logger.error', Mock(side_effect=SleepException))
|
||||
@patch('patroni.dcs.kubernetes.ObjectCache._build_cache', Mock(side_effect=Exception))
|
||||
def test_run(self):
|
||||
self.assertRaises(SleepException, self.k._pods.run)
|
||||
|
||||
@patch('time.sleep', Mock())
|
||||
def test__list(self):
|
||||
self.k._pods._func = Mock(side_effect=Exception)
|
||||
self.assertRaises(Exception, self.k._pods._list)
|
||||
|
||||
+17
-11
@@ -9,6 +9,8 @@ from patroni.config import Config
|
||||
from patroni.log import PatroniLogger
|
||||
from six.moves.queue import Queue, Full
|
||||
|
||||
_LOG = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class TestPatroniLogger(unittest.TestCase):
|
||||
|
||||
@@ -19,10 +21,10 @@ class TestPatroniLogger(unittest.TestCase):
|
||||
logging.getLogger().handlers[:] = self._handlers
|
||||
|
||||
@patch('logging.FileHandler._open', Mock())
|
||||
@patch('logging.Handler.close', Mock(side_effect=Exception))
|
||||
def test_patroni_logger(self):
|
||||
config = {
|
||||
'log': {
|
||||
'traceback_level': 'DEBUG',
|
||||
'max_queue_size': 5,
|
||||
'dir': 'foo',
|
||||
'file_size': 4096,
|
||||
@@ -36,23 +38,27 @@ class TestPatroniLogger(unittest.TestCase):
|
||||
sys.argv = ['patroni.py']
|
||||
os.environ[Config.PATRONI_CONFIG_VARIABLE] = yaml.dump(config, default_flow_style=False)
|
||||
logger = PatroniLogger()
|
||||
patroni_config = Config()
|
||||
patroni_config = Config(None)
|
||||
logger.reload_config(patroni_config['log'])
|
||||
_LOG.exception('test')
|
||||
logger.start()
|
||||
|
||||
with patch.object(logging.Handler, 'format', Mock(side_effect=Exception)):
|
||||
logging.error('test')
|
||||
|
||||
self.assertEqual(logger._log_handler.maxBytes, config['log']['file_size'])
|
||||
self.assertEqual(logger._log_handler.backupCount, config['log']['file_num'])
|
||||
self.assertEqual(logger.log_handler.maxBytes, config['log']['file_size'])
|
||||
self.assertEqual(logger.log_handler.backupCount, config['log']['file_num'])
|
||||
|
||||
config['log']['level'] = 'DEBUG'
|
||||
config['log'].pop('dir')
|
||||
logger.reload_config(config['log'])
|
||||
with patch.object(logging.Logger, 'makeRecord',
|
||||
Mock(side_effect=[logging.LogRecord('', logging.INFO, '', 0, '', (), None), Exception])):
|
||||
with patch('logging.Handler.close', Mock(side_effect=Exception)):
|
||||
logger.reload_config(config['log'])
|
||||
with patch.object(logging.Logger, 'makeRecord',
|
||||
Mock(side_effect=[logging.LogRecord('', logging.INFO, '', 0, '', (), None), Exception])):
|
||||
logging.exception('test')
|
||||
logging.error('test')
|
||||
logging.error('test')
|
||||
with patch.object(Queue, 'put_nowait', Mock(side_effect=Full)):
|
||||
self.assertRaises(SystemExit, logger.shutdown)
|
||||
self.assertRaises(Exception, logger.shutdown)
|
||||
with patch.object(Queue, 'put_nowait', Mock(side_effect=Full)):
|
||||
self.assertRaises(SystemExit, logger.shutdown)
|
||||
self.assertRaises(Exception, logger.shutdown)
|
||||
self.assertLessEqual(logger.queue_size, 2) # "Failed to close the old log handler" could be still in the queue
|
||||
self.assertEqual(logger.records_lost, 0)
|
||||
|
||||
+25
-13
@@ -2,10 +2,10 @@ import etcd
|
||||
import logging
|
||||
import os
|
||||
import signal
|
||||
import sys
|
||||
import time
|
||||
import unittest
|
||||
|
||||
import patroni.config as config
|
||||
from mock import Mock, PropertyMock, patch
|
||||
from patroni.api import RestApiServer
|
||||
from patroni.async_executor import AsyncExecutor
|
||||
@@ -39,23 +39,29 @@ class MockFrozenImporter(object):
|
||||
@patch.object(AsyncExecutor, 'run', Mock())
|
||||
@patch.object(etcd.Client, 'write', etcd_write)
|
||||
@patch.object(etcd.Client, 'read', etcd_read)
|
||||
@patch.object(Thread, 'start', Mock())
|
||||
class TestPatroni(unittest.TestCase):
|
||||
|
||||
@patch('pkgutil.get_importer', Mock(return_value=MockFrozenImporter()))
|
||||
def test_no_config(self):
|
||||
self.assertRaises(SystemExit, patroni_main)
|
||||
|
||||
@patch('sys.argv', ['patroni.py', '--validate-config', 'postgres0.yml'])
|
||||
def test_validate_config(self):
|
||||
self.assertRaises(SystemExit, patroni_main)
|
||||
|
||||
@patch('pkgutil.iter_importers', Mock(return_value=[MockFrozenImporter()]))
|
||||
@patch('sys.frozen', Mock(return_value=True), create=True)
|
||||
@patch.object(BaseHTTPServer.HTTPServer, '__init__', Mock())
|
||||
@patch.object(etcd.Client, 'read', etcd_read)
|
||||
@patch.object(Thread, 'start', Mock())
|
||||
@patch.object(Client, 'machines', PropertyMock(return_value=['http://remotehost:2379']))
|
||||
def setUp(self):
|
||||
self._handlers = logging.getLogger().handlers[:]
|
||||
RestApiServer._BaseServer__is_shut_down = Mock()
|
||||
RestApiServer._BaseServer__shutdown_request = True
|
||||
RestApiServer.socket = 0
|
||||
with patch.object(Client, 'machines') as mock_machines:
|
||||
mock_machines.__get__ = Mock(return_value=['http://remotehost:2379'])
|
||||
sys.argv = ['patroni.py', 'postgres0.yml']
|
||||
os.environ['PATRONI_POSTGRESQL_DATA_DIR'] = 'data/test0'
|
||||
self.p = Patroni()
|
||||
os.environ['PATRONI_POSTGRESQL_DATA_DIR'] = 'data/test0'
|
||||
conf = config.Config('postgres0.yml')
|
||||
self.p = Patroni(conf)
|
||||
|
||||
def tearDown(self):
|
||||
logging.getLogger().handlers[:] = self._handlers
|
||||
@@ -66,19 +72,19 @@ class TestPatroni(unittest.TestCase):
|
||||
self.p.load_dynamic_configuration()
|
||||
self.p.load_dynamic_configuration()
|
||||
|
||||
@patch('sys.argv', ['patroni.py', 'postgres0.yml'])
|
||||
@patch('time.sleep', Mock(side_effect=SleepException))
|
||||
@patch.object(etcd.Client, 'delete', Mock())
|
||||
@patch.object(Client, 'machines')
|
||||
@patch.object(Client, 'machines', PropertyMock(return_value=['http://remotehost:2379']))
|
||||
@patch.object(Thread, 'join', Mock())
|
||||
def test_patroni_patroni_main(self, mock_machines):
|
||||
def test_patroni_patroni_main(self):
|
||||
with patch('subprocess.call', Mock(return_value=1)):
|
||||
sys.argv = ['patroni.py', 'postgres0.yml']
|
||||
|
||||
mock_machines.__get__ = Mock(return_value=['http://remotehost:2379'])
|
||||
with patch.object(Patroni, 'run', Mock(side_effect=SleepException)):
|
||||
os.environ['PATRONI_POSTGRESQL_DATA_DIR'] = 'data/test0'
|
||||
self.assertRaises(SleepException, patroni_main)
|
||||
with patch.object(Patroni, 'run', Mock(side_effect=KeyboardInterrupt())):
|
||||
with patch('patroni.ha.Ha.is_paused', Mock(return_value=True)):
|
||||
os.environ['PATRONI_POSTGRESQL_DATA_DIR'] = 'data/test0'
|
||||
patroni_main()
|
||||
|
||||
@patch('os.getpid')
|
||||
@@ -122,8 +128,12 @@ class TestPatroni(unittest.TestCase):
|
||||
self.p.sighup_handler()
|
||||
self.p.ha.dcs.watch = Mock(side_effect=SleepException)
|
||||
self.p.api.start = Mock()
|
||||
self.p.logger.start = Mock()
|
||||
self.p.config._dynamic_configuration = {}
|
||||
self.assertRaises(SleepException, self.p.run)
|
||||
with patch('patroni.config.Config.reload_local_configuration', Mock(return_value=False)):
|
||||
self.p.sighup_handler()
|
||||
self.assertRaises(SleepException, self.p.run)
|
||||
with patch('patroni.config.Config.set_dynamic_configuration', Mock(return_value=True)):
|
||||
self.assertRaises(SleepException, self.p.run)
|
||||
with patch('patroni.postgresql.Postgresql.data_directory_empty', Mock(return_value=False)):
|
||||
@@ -165,8 +175,10 @@ class TestPatroni(unittest.TestCase):
|
||||
self.p.tags['nosync'] = None
|
||||
self.assertFalse(self.p.nosync)
|
||||
|
||||
@patch.object(Thread, 'join', Mock())
|
||||
def test_shutdown(self):
|
||||
self.p.api.shutdown = Mock(side_effect=Exception)
|
||||
self.p.ha.shutdown = Mock(side_effect=Exception)
|
||||
self.p.shutdown()
|
||||
|
||||
def test_check_psycopg2(self):
|
||||
|
||||
+98
-40
@@ -1,13 +1,15 @@
|
||||
import mock # for the mock.call method, importing it without a namespace breaks python3
|
||||
import os
|
||||
import psutil
|
||||
import psycopg2
|
||||
import re
|
||||
import subprocess
|
||||
import time
|
||||
|
||||
from mock import Mock, MagicMock, PropertyMock, patch, mock_open
|
||||
from patroni.async_executor import CriticalTask
|
||||
from patroni.dcs import Cluster, ClusterConfig, Member, RemoteMember, SyncState
|
||||
from patroni.exceptions import PostgresConnectionException
|
||||
from patroni.exceptions import PostgresConnectionException, PatroniException
|
||||
from patroni.postgresql import Postgresql, STATE_REJECT, STATE_NO_RESPONSE
|
||||
from patroni.postgresql.postmaster import PostmasterProcess
|
||||
from patroni.postgresql.slots import SlotsHandler
|
||||
@@ -90,6 +92,7 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
|
||||
@patch('subprocess.call', Mock(return_value=0))
|
||||
@patch('os.rename', Mock())
|
||||
@patch('patroni.postgresql.CallbackExecutor', Mock())
|
||||
@patch.object(Postgresql, 'get_major_version', Mock(return_value=120000))
|
||||
@patch.object(Postgresql, 'is_running', Mock(return_value=True))
|
||||
def setUp(self):
|
||||
@@ -101,6 +104,7 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
@patch.object(Postgresql, 'wait_for_startup')
|
||||
@patch.object(Postgresql, 'wait_for_port_open')
|
||||
@patch.object(Postgresql, 'is_running')
|
||||
@patch.object(Postgresql, 'controldata', Mock())
|
||||
def test_start(self, mock_is_running, mock_wait_for_port_open, mock_wait_for_startup, mock_popen):
|
||||
mock_is_running.return_value = MockPostmaster()
|
||||
mock_wait_for_port_open.return_value = True
|
||||
@@ -131,6 +135,9 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
|
||||
self.p.cancellable.cancel()
|
||||
self.assertFalse(self.p.start())
|
||||
with patch('patroni.postgresql.config.ConfigHandler.effective_configuration',
|
||||
PropertyMock(side_effect=Exception)):
|
||||
self.assertIsNone(self.p.start())
|
||||
|
||||
@patch.object(Postgresql, 'pg_isready')
|
||||
@patch('patroni.postgresql.polling_loop', Mock(return_value=range(1)))
|
||||
@@ -171,6 +178,17 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
mock_callback.assert_called()
|
||||
mock_postmaster.signal_stop.assert_called()
|
||||
|
||||
# Timed out waiting for fast shutdown triggers immediate shutdown
|
||||
mock_postmaster.wait.side_effect = [psutil.TimeoutExpired(30), psutil.TimeoutExpired(30), Mock()]
|
||||
mock_callback.reset_mock()
|
||||
self.assertTrue(self.p.stop(on_safepoint=mock_callback, stop_timeout=30))
|
||||
mock_callback.assert_called()
|
||||
mock_postmaster.signal_stop.assert_called()
|
||||
|
||||
# Immediate shutdown succeeded
|
||||
mock_postmaster.wait.side_effect = [psutil.TimeoutExpired(30), Mock()]
|
||||
self.assertTrue(self.p.stop(on_safepoint=mock_callback, stop_timeout=30))
|
||||
|
||||
# Stop signal failed
|
||||
mock_postmaster.signal_stop.return_value = False
|
||||
self.assertFalse(self.p.stop())
|
||||
@@ -181,59 +199,81 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
self.assertTrue(self.p.stop(on_safepoint=mock_callback))
|
||||
mock_callback.assert_called()
|
||||
|
||||
# Fast shutdown is timed out but when immediate postmaster is already gone
|
||||
mock_postmaster.wait.side_effect = [psutil.TimeoutExpired(30), Mock()]
|
||||
mock_postmaster.signal_stop.side_effect = [None, True]
|
||||
self.assertTrue(self.p.stop(on_safepoint=mock_callback, stop_timeout=30))
|
||||
|
||||
def test_restart(self):
|
||||
self.p.start = Mock(return_value=False)
|
||||
self.assertFalse(self.p.restart())
|
||||
self.assertEqual(self.p.state, 'restart failed (restarting)')
|
||||
|
||||
@patch('os.chmod', Mock())
|
||||
@patch.object(builtins, 'open', MagicMock())
|
||||
def test_write_pgpass(self):
|
||||
self.p.write_pgpass({'host': 'localhost', 'port': '5432', 'user': 'foo'})
|
||||
self.p.write_pgpass({'host': 'localhost', 'port': '5432', 'user': 'foo', 'password': 'bar'})
|
||||
self.p.config.write_pgpass({'host': 'localhost', 'port': '5432', 'user': 'foo'})
|
||||
self.p.config.write_pgpass({'host': 'localhost', 'port': '5432', 'user': 'foo', 'password': 'bar'})
|
||||
|
||||
def test_checkpoint(self):
|
||||
with patch.object(MockCursor, 'fetchone', Mock(return_value=(True, ))):
|
||||
self.assertEqual(self.p.checkpoint({'user': 'postgres'}), 'is_in_recovery=true')
|
||||
with patch.object(MockCursor, 'execute', Mock(return_value=None)):
|
||||
self.assertIsNone(self.p.checkpoint())
|
||||
self.assertEqual(self.p.checkpoint(), 'not accessible or not healty')
|
||||
self.assertEqual(self.p.checkpoint(timeout=10), 'not accessible or not healty')
|
||||
|
||||
@patch('patroni.postgresql.config.mtime', mock_mtime)
|
||||
@patch.object(MockCursor, 'fetchone', Mock(side_effect=[('foo=bar',), ('',), ('',), ('a=b',)]))
|
||||
def test_check_recovery_conf(self):
|
||||
for version in (120000, 100000):
|
||||
with patch.object(Postgresql, 'major_version', PropertyMock(return_value=version)):
|
||||
self.p.config.write_recovery_conf({'standby_mode': 'on', 'primary_conninfo': 'foo'})
|
||||
self.assertFalse(self.p.config.check_recovery_conf(None))
|
||||
self.p.config.write_recovery_conf({'primary_conninfo': 'foo'})
|
||||
self.p.config.write_postgresql_conf()
|
||||
self.assertFalse(self.p.config.check_recovery_conf(None))
|
||||
self.p.config.write_recovery_conf({'standby_mode': 'on'})
|
||||
self.assertTrue(self.p.config.check_recovery_conf(None))
|
||||
with patch('patroni.postgresql.config.ConfigHandler.primary_conninfo_params',
|
||||
Mock(return_value={'a': 'b'})):
|
||||
self.assertFalse(self.p.config.check_recovery_conf(None))
|
||||
self.p.config.write_recovery_conf({'standby_mode': 'on', 'primary_conninfo': 'a=b'})
|
||||
self.assertTrue(self.p.config.check_recovery_conf(None))
|
||||
@patch('patroni.postgresql.config.ConfigHandler._get_pg_settings')
|
||||
def test_check_recovery_conf(self, mock_get_pg_settings):
|
||||
mock_get_pg_settings.return_value = {
|
||||
'primary_conninfo': ['primary_conninfo', 'foo=', None, 'string', 'postmaster', self.p.config._auto_conf],
|
||||
'recovery_min_apply_delay': ['recovery_min_apply_delay', '0', 'ms', 'integer', 'sighup', 'foo']
|
||||
}
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (True, True))
|
||||
self.p.config.write_recovery_conf({'standby_mode': 'on'})
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (True, True))
|
||||
mock_get_pg_settings.return_value['primary_conninfo'][1] = ''
|
||||
mock_get_pg_settings.return_value['recovery_min_apply_delay'][1] = '1'
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (False, False))
|
||||
mock_get_pg_settings.return_value['recovery_min_apply_delay'][5] = self.p.config._auto_conf
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (True, False))
|
||||
mock_get_pg_settings.return_value['recovery_min_apply_delay'][1] = '0'
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (False, False))
|
||||
conninfo = {'host': '1', 'password': 'bar'}
|
||||
with patch('patroni.postgresql.config.ConfigHandler.primary_conninfo_params', Mock(return_value=conninfo)):
|
||||
mock_get_pg_settings.return_value['recovery_min_apply_delay'][1] = '1'
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (True, True))
|
||||
mock_get_pg_settings.return_value['primary_conninfo'][1] = 'host=1 passfile='\
|
||||
+ re.sub(r'([\'\\ ])', r'\\\1', self.p.config._pgpass)
|
||||
mock_get_pg_settings.return_value['recovery_min_apply_delay'][1] = '0'
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (True, True))
|
||||
self.p.config.write_recovery_conf({'standby_mode': 'on', 'primary_conninfo': conninfo.copy()})
|
||||
self.p.config.write_postgresql_conf()
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (False, False))
|
||||
|
||||
@patch.object(Postgresql, 'major_version', PropertyMock(return_value=120000))
|
||||
@patch.object(Postgresql, 'is_running', MockPostmaster)
|
||||
@patch.object(MockPostmaster, 'create_time', Mock(return_value=1234567), create=True)
|
||||
@patch.object(MockCursor, 'fetchone', Mock(return_value=('',)))
|
||||
def test__read_primary_conninfo(self):
|
||||
self.p.config.write_recovery_conf({'standby_mode': 'on'})
|
||||
@patch('patroni.postgresql.config.ConfigHandler._get_pg_settings')
|
||||
def test__read_recovery_params(self, mock_get_pg_settings):
|
||||
mock_get_pg_settings.return_value = {'primary_conninfo': ['primary_conninfo', '', None, 'string',
|
||||
'postmaster', self.p.config._postgresql_conf]}
|
||||
self.p.config.write_recovery_conf({'standby_mode': 'on', 'primary_conninfo': {'password': 'foo'}})
|
||||
self.p.config.write_postgresql_conf()
|
||||
self.assertTrue(self.p.config.check_recovery_conf(None))
|
||||
self.assertTrue(self.p.config.check_recovery_conf(None))
|
||||
with patch.object(Postgresql, 'query', Mock(side_effect=Exception)),\
|
||||
patch('patroni.postgresql.config.mtime', mock_mtime):
|
||||
self.assertFalse(self.p.config.check_recovery_conf(None))
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (False, False))
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (False, False))
|
||||
mock_get_pg_settings.side_effect = Exception
|
||||
with patch('patroni.postgresql.config.mtime', mock_mtime):
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (True, True))
|
||||
|
||||
@patch.object(Postgresql, 'major_version', PropertyMock(return_value=100000))
|
||||
def test__read_primary_conninfo_pre_v12(self):
|
||||
self.p.config.write_recovery_conf({'standby_mode': 'on'})
|
||||
self.assertTrue(self.p.config.check_recovery_conf(None))
|
||||
self.assertTrue(self.p.config.check_recovery_conf(None))
|
||||
def test__read_recovery_params_pre_v12(self):
|
||||
self.p.config.write_recovery_conf({'standby_mode': 'on', 'primary_conninfo': {'password': 'foo'}})
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (True, True))
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (True, True))
|
||||
self.p.config.write_recovery_conf({'standby_mode': '\n'})
|
||||
with patch('patroni.postgresql.config.mtime', mock_mtime):
|
||||
self.assertEqual(self.p.config.check_recovery_conf(None), (True, True))
|
||||
|
||||
def test_write_postgresql_and_sanitize_auto_conf(self):
|
||||
read_data = 'primary_conninfo = foo\nfoo = bar\n'
|
||||
@@ -255,14 +295,14 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
@patch.object(Postgresql, 'start', Mock())
|
||||
def test_follow(self):
|
||||
self.p.call_nowait('on_start')
|
||||
m = RemoteMember('1', {'restore_command': '2', 'recovery_min_apply_delay': 3, 'archive_cleanup_command': '4'})
|
||||
m = RemoteMember('1', {'restore_command': '2', 'primary_slot_name': 'foo', 'conn_kwargs': {'host': 'bar'}})
|
||||
self.p.follow(m)
|
||||
|
||||
@patch.object(Postgresql, 'is_running', Mock(return_value=True))
|
||||
def test_sync_replication_slots(self):
|
||||
self.p.start()
|
||||
config = ClusterConfig(1, {'slots': {'ls': {'database': 'a', 'plugin': 'b'},
|
||||
'A': 0, 'test_3': 0, 'b': {'type': 'logical', 'plugin': '1'}}}, 1)
|
||||
config = ClusterConfig(1, {'slots': {'test_3': {'database': 'a', 'plugin': 'b'},
|
||||
'A': 0, 'ls': 0, 'b': {'type': 'logical', 'plugin': '1'}}}, 1)
|
||||
cluster = Cluster(True, config, self.leader, 0, [self.me, self.other, self.leadermem], None, None, None)
|
||||
with mock.patch('patroni.postgresql.Postgresql._query', Mock(side_effect=psycopg2.OperationalError)):
|
||||
self.p.slots_handler.sync_replication_slots(cluster)
|
||||
@@ -314,7 +354,7 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
self.assertTrue(self.p.promote(0))
|
||||
|
||||
def test_timeline_wal_position(self):
|
||||
self.assertEqual(self.p.timeline_wal_position(), (1, 2))
|
||||
self.assertEqual(self.p.timeline_wal_position(), (1, 2, 1))
|
||||
Thread(target=self.p.timeline_wal_position).start()
|
||||
|
||||
@patch.object(PostmasterProcess, 'from_pidfile')
|
||||
@@ -359,6 +399,7 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
|
||||
@patch('os.listdir', Mock(return_value=['recovery.conf']))
|
||||
@patch('os.path.exists', Mock(return_value=True))
|
||||
@patch.object(Postgresql, 'controldata', Mock())
|
||||
def test_get_postgres_role_from_data_directory(self):
|
||||
self.assertEqual(self.p.get_postgres_role_from_data_directory(), 'replica')
|
||||
|
||||
@@ -371,6 +412,10 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
pass
|
||||
os.makedirs(os.path.join(self.p.data_dir, 'foo'))
|
||||
_symlink('foo', os.path.join(self.p.data_dir, 'pg_wal'))
|
||||
os.makedirs(os.path.join(self.p.data_dir, 'foo_tsp'))
|
||||
pg_tblspc = os.path.join(self.p.data_dir, 'pg_tblspc')
|
||||
os.makedirs(pg_tblspc)
|
||||
_symlink('../foo_tsp', os.path.join(pg_tblspc, '12345'))
|
||||
self.p.remove_data_directory()
|
||||
open(self.p.data_dir, 'w').close()
|
||||
self.p.remove_data_directory()
|
||||
@@ -421,13 +466,18 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
self.p.config._config['foo'] = {'command': 'bar'}
|
||||
self.assertFalse(self.p.replica_method_can_work_without_replication_connection('foo'))
|
||||
|
||||
@patch('time.sleep', Mock())
|
||||
@patch.object(Postgresql, 'is_running', Mock(return_value=True))
|
||||
def test_reload_config(self):
|
||||
@patch.object(MockCursor, 'fetchone')
|
||||
def test_reload_config(self, mock_fetchone):
|
||||
mock_fetchone.return_value = (1,)
|
||||
parameters = self._PARAMETERS.copy()
|
||||
parameters.pop('f.oo')
|
||||
parameters['wal_buffers'] = '512'
|
||||
config = {'pg_hba': [''], 'pg_ident': [''], 'use_unix_socket': True, 'authentication': {},
|
||||
'retry_timeout': 10, 'listen': '*', 'krbsrvname': 'postgres', 'parameters': parameters}
|
||||
self.p.reload_config(config)
|
||||
mock_fetchone.side_effect = Exception
|
||||
parameters['b.ar'] = 'bar'
|
||||
self.p.reload_config(config)
|
||||
parameters['autovacuum'] = 'on'
|
||||
@@ -450,6 +500,10 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
def test_postmaster_start_time(self):
|
||||
with patch.object(MockCursor, "fetchone", Mock(return_value=('foo', True, '', '', '', '', False))):
|
||||
self.assertEqual(self.p.postmaster_start_time(), 'foo')
|
||||
t = Thread(target=self.p.postmaster_start_time)
|
||||
t.start()
|
||||
t.join()
|
||||
|
||||
with patch.object(MockCursor, "execute", side_effect=psycopg2.Error):
|
||||
self.assertIsNone(self.p.postmaster_start_time())
|
||||
|
||||
@@ -595,7 +649,7 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
config['synchronous_mode_strict'] = True
|
||||
self.p.config.get_server_parameters(config)
|
||||
self.p.config.set_synchronous_standby('foo')
|
||||
self.p.config.get_server_parameters(config)
|
||||
self.assertTrue(str(self.p.config.get_server_parameters(config)).startswith('{'))
|
||||
|
||||
@patch('time.sleep', Mock())
|
||||
def test__wait_for_connection_close(self):
|
||||
@@ -629,7 +683,7 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
data = self.p.read_postmaster_opts()
|
||||
self.assertEqual(data, dict())
|
||||
|
||||
@patch('subprocess.Popen')
|
||||
@patch('psutil.Popen')
|
||||
def test_single_user_mode(self, subprocess_popen_mock):
|
||||
subprocess_popen_mock.return_value.wait.return_value = 0
|
||||
self.assertEqual(self.p.single_user_mode('CHECKPOINT', {'archive_mode': 'on'}), 0)
|
||||
@@ -662,9 +716,13 @@ class TestPostgresql(BaseTestPostgresql):
|
||||
with patch.object(Postgresql, 'controldata',
|
||||
Mock(return_value={'max_connections setting': '200',
|
||||
'max_worker_processes setting': '20',
|
||||
'max_prepared_xacts setting': '100',
|
||||
'max_locks_per_xact setting': '100',
|
||||
'max_wal_senders setting': 10})):
|
||||
self.p.cancellable.cancel()
|
||||
self.assertFalse(self.p.start())
|
||||
self.assertTrue(self.p.pending_restart)
|
||||
|
||||
@patch('os.path.exists', Mock(return_value=True))
|
||||
@patch('os.path.isfile', Mock(return_value=False))
|
||||
def test_pgpass_is_dir(self):
|
||||
self.assertRaises(PatroniException, self.setUp)
|
||||
|
||||
@@ -1,3 +1,4 @@
|
||||
import multiprocessing
|
||||
import psutil
|
||||
import unittest
|
||||
|
||||
@@ -62,9 +63,42 @@ class TestPostmasterProcess(unittest.TestCase):
|
||||
mock_init.side_effect = None
|
||||
self.assertNotEqual(PostmasterProcess.from_pid(123), None)
|
||||
|
||||
@patch('psutil.Process.__init__', Mock())
|
||||
@patch('psutil.wait_procs', Mock())
|
||||
@patch('psutil.Process.suspend')
|
||||
@patch('psutil.Process.children')
|
||||
@patch('psutil.Process.kill')
|
||||
def test_signal_kill(self, mock_kill, mock_children, mock_suspend):
|
||||
proc = PostmasterProcess(123)
|
||||
|
||||
# all processes successfully stopped
|
||||
mock_children.return_value = [Mock()]
|
||||
mock_children.return_value[0].kill.side_effect = psutil.Error
|
||||
self.assertTrue(proc.signal_kill())
|
||||
|
||||
# postmaster has gone before suspend
|
||||
mock_suspend.side_effect = psutil.NoSuchProcess(123)
|
||||
self.assertTrue(proc.signal_kill())
|
||||
|
||||
# postmaster has gone before we got a list of children
|
||||
mock_suspend.side_effect = psutil.Error()
|
||||
mock_children.side_effect = psutil.NoSuchProcess(123)
|
||||
self.assertTrue(proc.signal_kill())
|
||||
|
||||
# postmaster has gone after we got a list of children
|
||||
mock_children.side_effect = psutil.Error()
|
||||
mock_kill.side_effect = psutil.NoSuchProcess(123)
|
||||
self.assertTrue(proc.signal_kill())
|
||||
|
||||
# failed to kill postmaster
|
||||
mock_kill.side_effect = psutil.AccessDenied(123)
|
||||
self.assertFalse(proc.signal_kill())
|
||||
|
||||
@patch('psutil.Process.__init__', Mock())
|
||||
@patch('psutil.Process.send_signal')
|
||||
@patch('psutil.Process.pid', Mock(return_value=123))
|
||||
@patch('os.name', 'posix')
|
||||
@patch('signal.SIGQUIT', 3, create=True)
|
||||
def test_signal_stop(self, mock_send_signal):
|
||||
proc = PostmasterProcess(-123)
|
||||
self.assertEqual(proc.signal_stop('immediate'), False)
|
||||
@@ -75,6 +109,21 @@ class TestPostmasterProcess(unittest.TestCase):
|
||||
self.assertEqual(proc.signal_stop('immediate'), True)
|
||||
self.assertEqual(proc.signal_stop('immediate'), False)
|
||||
|
||||
@patch('psutil.Process.__init__', Mock())
|
||||
@patch('patroni.postgresql.postmaster.os')
|
||||
@patch('subprocess.call', Mock(side_effect=[0, OSError, 1]))
|
||||
@patch('psutil.Process.pid', Mock(return_value=123))
|
||||
@patch('psutil.Process.is_running', Mock(return_value=False))
|
||||
def test_signal_stop_nt(self, mock_os):
|
||||
mock_os.configure_mock(name="nt")
|
||||
proc = PostmasterProcess(-123)
|
||||
self.assertEqual(proc.signal_stop('immediate'), False)
|
||||
|
||||
proc = PostmasterProcess(123)
|
||||
self.assertEqual(proc.signal_stop('immediate'), None)
|
||||
self.assertEqual(proc.signal_stop('immediate'), False)
|
||||
self.assertEqual(proc.signal_stop('immediate'), True)
|
||||
|
||||
@patch('psutil.Process.__init__', Mock())
|
||||
@patch('psutil.wait_procs')
|
||||
def test_wait_for_user_backends_to_close(self, mock_wait):
|
||||
@@ -96,6 +145,7 @@ class TestPostmasterProcess(unittest.TestCase):
|
||||
@patch('subprocess.Popen')
|
||||
@patch('os.setsid', Mock(), create=True)
|
||||
@patch('multiprocessing.Process', MockProcess)
|
||||
@patch('multiprocessing.get_context', Mock(return_value=multiprocessing), create=True)
|
||||
@patch.object(PostmasterProcess, 'from_pid')
|
||||
@patch.object(PostmasterProcess, '_from_pidfile')
|
||||
def test_start(self, mock_frompidfile, mock_frompid, mock_popen):
|
||||
|
||||
+28
-3
@@ -7,6 +7,16 @@ from patroni.postgresql.rewind import Rewind
|
||||
from . import BaseTestPostgresql, MockCursor, psycopg2_connect
|
||||
|
||||
|
||||
class MockThread(object):
|
||||
|
||||
def __init__(self, target, args):
|
||||
self._target = target
|
||||
self._args = args
|
||||
|
||||
def start(self):
|
||||
self._target(*self._args)
|
||||
|
||||
|
||||
@patch('subprocess.call', Mock(return_value=0))
|
||||
@patch('psycopg2.connect', psycopg2_connect)
|
||||
class TestRewind(BaseTestPostgresql):
|
||||
@@ -102,6 +112,21 @@ class TestRewind(BaseTestPostgresql):
|
||||
self.r.check_leader_is_not_in_recovery()
|
||||
self.r.check_leader_is_not_in_recovery()
|
||||
|
||||
@patch.object(Postgresql, 'controldata', Mock(return_value={"Latest checkpoint's TimeLineID": 1}))
|
||||
def test_check_for_checkpoint_after_promote(self):
|
||||
self.r.check_for_checkpoint_after_promote()
|
||||
@patch('patroni.postgresql.rewind.Thread', MockThread)
|
||||
@patch.object(Postgresql, 'controldata')
|
||||
@patch.object(Postgresql, 'checkpoint')
|
||||
def test_ensure_checkpoint_after_promote(self, mock_checkpoint, mock_controldata):
|
||||
mock_checkpoint.return_value = None
|
||||
self.r.ensure_checkpoint_after_promote()
|
||||
self.r.ensure_checkpoint_after_promote()
|
||||
|
||||
self.r.reset_state()
|
||||
mock_controldata.return_value = {"Latest checkpoint's TimeLineID": 1}
|
||||
mock_checkpoint.side_effect = Exception
|
||||
self.r.ensure_checkpoint_after_promote()
|
||||
self.r.ensure_checkpoint_after_promote()
|
||||
|
||||
self.r.reset_state()
|
||||
mock_controldata.side_effect = TypeError
|
||||
self.r.ensure_checkpoint_after_promote()
|
||||
self.r.ensure_checkpoint_after_promote()
|
||||
|
||||
+24
-1
@@ -2,7 +2,7 @@ import unittest
|
||||
|
||||
from mock import Mock, patch
|
||||
from patroni.exceptions import PatroniException
|
||||
from patroni.utils import Retry, RetryFailedError, polling_loop
|
||||
from patroni.utils import Retry, RetryFailedError, polling_loop, validate_directory
|
||||
|
||||
|
||||
class TestUtils(unittest.TestCase):
|
||||
@@ -10,6 +10,29 @@ class TestUtils(unittest.TestCase):
|
||||
def test_polling_loop(self):
|
||||
self.assertEqual(list(polling_loop(0.001, interval=0.001)), [0])
|
||||
|
||||
@patch('os.path.exists', Mock(return_value=True))
|
||||
@patch('os.path.isdir', Mock(return_value=True))
|
||||
@patch('tempfile.mkstemp', Mock(return_value=("", "")))
|
||||
@patch('os.remove', Mock(side_effect=Exception))
|
||||
def test_validate_directory_writable(self):
|
||||
self.assertRaises(Exception, validate_directory, "/tmp")
|
||||
|
||||
@patch('os.path.exists', Mock(return_value=True))
|
||||
@patch('os.path.isdir', Mock(return_value=True))
|
||||
@patch('tempfile.mkstemp', Mock(side_effect=OSError))
|
||||
def test_validate_directory_not_writable(self):
|
||||
self.assertRaises(PatroniException, validate_directory, "/tmp")
|
||||
|
||||
@patch('os.path.exists', Mock(return_value=False))
|
||||
@patch('os.makedirs', Mock(side_effect=OSError))
|
||||
def test_validate_directory_couldnt_create(self):
|
||||
self.assertRaises(PatroniException, validate_directory, "/tmp")
|
||||
|
||||
@patch('os.path.exists', Mock(return_value=True))
|
||||
@patch('os.path.isdir', Mock(return_value=False))
|
||||
def test_validate_directory_is_not_a_directory(self):
|
||||
self.assertRaises(PatroniException, validate_directory, "/tmp")
|
||||
|
||||
|
||||
@patch('time.sleep', Mock())
|
||||
class TestRetrySleeper(unittest.TestCase):
|
||||
|
||||
@@ -0,0 +1,229 @@
|
||||
import copy
|
||||
import os
|
||||
import socket
|
||||
import tempfile
|
||||
import unittest
|
||||
|
||||
from mock import Mock, patch, mock_open
|
||||
from patroni.dcs import dcs_modules
|
||||
from patroni.validator import schema
|
||||
from six import StringIO
|
||||
|
||||
available_dcs = [m.split(".")[-1] for m in dcs_modules()]
|
||||
config = {
|
||||
"name": "string",
|
||||
"scope": "string",
|
||||
"restapi": {
|
||||
"listen": "127.0.0.2:800",
|
||||
"connect_address": "127.0.0.2:800"
|
||||
},
|
||||
"bootstrap": {
|
||||
"dcs": {
|
||||
"ttl": 1000,
|
||||
"loop_wait": 1000,
|
||||
"retry_timeout": 1000,
|
||||
"maximum_lag_on_failover": 1000
|
||||
},
|
||||
"pg_hba": ["string"],
|
||||
"initdb": ["string", {"key": "value"}]
|
||||
},
|
||||
"consul": {
|
||||
"host": "127.0.0.1:5000"
|
||||
},
|
||||
"etcd": {
|
||||
"hosts": "127.0.0.1:2379,127.0.0.1:2380"
|
||||
},
|
||||
"exhibitor": {
|
||||
"hosts": ["string"],
|
||||
"port": 4000,
|
||||
"pool_interval": 1000
|
||||
},
|
||||
"zookeeper": {
|
||||
"hosts": "127.0.0.1:3379,127.0.0.1:3380"
|
||||
},
|
||||
"kubernetes": {
|
||||
"namespace": "string",
|
||||
"labels": {},
|
||||
"scope_label": "string",
|
||||
"role_label": "string",
|
||||
"use_endpoints": False,
|
||||
"pod_ip": "127.0.0.1",
|
||||
"ports": [{"name": "string", "port": 1000}],
|
||||
},
|
||||
"postgresql": {
|
||||
"listen": "127.0.0.2,::1:543",
|
||||
"connect_address": "127.0.0.2:543",
|
||||
"authentication": {
|
||||
"replication": {"username": "user"},
|
||||
"superuser": {"username": "user"},
|
||||
"rewind": {"username": "user"},
|
||||
},
|
||||
"data_dir": os.path.join(tempfile.gettempdir(), "data_dir"),
|
||||
"bin_dir": os.path.join(tempfile.gettempdir(), "bin_dir"),
|
||||
"parameters": {
|
||||
"unix_socket_directories": "."
|
||||
},
|
||||
"pg_hba": [u"string"],
|
||||
"pg_ident": ["string"],
|
||||
"pg_ctl_timeout": 1000,
|
||||
"use_pg_rewind": False
|
||||
},
|
||||
"watchdog": {
|
||||
"mode": "off",
|
||||
"device": "string"
|
||||
},
|
||||
"tags": {
|
||||
"nofailover": False,
|
||||
"clonefrom": False,
|
||||
"noloadbalance": False,
|
||||
"nosync": False
|
||||
}
|
||||
}
|
||||
|
||||
directories = []
|
||||
files = []
|
||||
|
||||
|
||||
def isfile_side_effect(arg):
|
||||
if arg.endswith('.exe'):
|
||||
arg = arg[:-4]
|
||||
return arg in files
|
||||
|
||||
|
||||
def isdir_side_effect(arg):
|
||||
return arg in directories
|
||||
|
||||
|
||||
def exists_side_effect(arg):
|
||||
return isfile_side_effect(arg) or isdir_side_effect(arg)
|
||||
|
||||
|
||||
def connect_side_effect(host_port):
|
||||
_, port = host_port
|
||||
if port < 1000:
|
||||
return 1
|
||||
elif port < 10000:
|
||||
return 0
|
||||
else:
|
||||
raise socket.gaierror()
|
||||
|
||||
|
||||
def parse_output(output):
|
||||
result = []
|
||||
for s in output.split("\n"):
|
||||
x = s.split(" ")[0]
|
||||
if x and x not in result:
|
||||
result.append(x)
|
||||
result.sort()
|
||||
return result
|
||||
|
||||
|
||||
@patch('socket.socket.connect_ex', Mock(side_effect=connect_side_effect))
|
||||
@patch('os.path.exists', Mock(side_effect=exists_side_effect))
|
||||
@patch('os.path.isdir', Mock(side_effect=isdir_side_effect))
|
||||
@patch('os.path.isfile', Mock(side_effect=isfile_side_effect))
|
||||
@patch('sys.stderr', new_callable=StringIO)
|
||||
@patch('sys.stdout', new_callable=StringIO)
|
||||
class TestValidator(unittest.TestCase):
|
||||
|
||||
def setUp(self):
|
||||
del files[:]
|
||||
del directories[:]
|
||||
|
||||
def test_empty_config(self, mock_out, mock_err):
|
||||
schema({})
|
||||
output = mock_out.getvalue()
|
||||
expected = list(sorted(['name', 'postgresql', 'restapi', 'scope'] + available_dcs))
|
||||
self.assertEqual(expected, parse_output(output))
|
||||
|
||||
def test_complete_config(self, mock_out, mock_err):
|
||||
schema(config)
|
||||
output = mock_out.getvalue()
|
||||
self.assertEqual(['postgresql.bin_dir'], parse_output(output))
|
||||
|
||||
def test_bin_dir_is_file(self, mock_out, mock_err):
|
||||
files.append(config["postgresql"]["data_dir"])
|
||||
files.append(config["postgresql"]["bin_dir"])
|
||||
c = copy.deepcopy(config)
|
||||
c["restapi"]["connect_address"] = 'False:blabla'
|
||||
c["etcd"]["hosts"] = ["127.0.0.1:2379", "1244.0.0.1:2379", "127.0.0.1:invalidport"]
|
||||
c["kubernetes"]["pod_ip"] = "127.0.0.1111"
|
||||
schema(c)
|
||||
output = mock_out.getvalue()
|
||||
self.assertEqual(['etcd.hosts.1', 'etcd.hosts.2', 'kubernetes.pod_ip', 'postgresql.bin_dir',
|
||||
'postgresql.data_dir', 'restapi.connect_address'], parse_output(output))
|
||||
|
||||
def test_bin_dir_is_empty(self, mock_out, mock_err):
|
||||
directories.append(config["postgresql"]["data_dir"])
|
||||
directories.append(config["postgresql"]["bin_dir"])
|
||||
files.append(os.path.join(config["postgresql"]["data_dir"], "global", "pg_control"))
|
||||
c = copy.deepcopy(config)
|
||||
c["restapi"]["connect_address"] = "127.0.0.1:8008"
|
||||
c["kubernetes"]["pod_ip"] = "::1"
|
||||
c["consul"]["host"] = "127.0.0.1:50000"
|
||||
c["etcd"]["host"] = "127.0.0.1:237"
|
||||
c["postgresql"]["listen"] = "127.0.0.1:5432"
|
||||
with patch('patroni.validator.open', mock_open(read_data='9')):
|
||||
schema(c)
|
||||
output = mock_out.getvalue()
|
||||
self.assertEqual(['consul.host', 'etcd.host', 'postgresql.bin_dir', 'postgresql.data_dir',
|
||||
'postgresql.listen', 'restapi.connect_address'], parse_output(output))
|
||||
|
||||
@patch('subprocess.check_output', Mock(return_value=b"postgres (PostgreSQL) 12.1"))
|
||||
def test_data_dir_contains_pg_version(self, mock_out, mock_err):
|
||||
directories.append(config["postgresql"]["data_dir"])
|
||||
directories.append(config["postgresql"]["bin_dir"])
|
||||
directories.append(os.path.join(config["postgresql"]["data_dir"], "pg_wal"))
|
||||
files.append(os.path.join(config["postgresql"]["data_dir"], "global", "pg_control"))
|
||||
files.append(os.path.join(config["postgresql"]["data_dir"], "PG_VERSION"))
|
||||
files.append(os.path.join(config["postgresql"]["bin_dir"], "pg_ctl"))
|
||||
files.append(os.path.join(config["postgresql"]["bin_dir"], "initdb"))
|
||||
files.append(os.path.join(config["postgresql"]["bin_dir"], "pg_controldata"))
|
||||
files.append(os.path.join(config["postgresql"]["bin_dir"], "pg_basebackup"))
|
||||
files.append(os.path.join(config["postgresql"]["bin_dir"], "postgres"))
|
||||
files.append(os.path.join(config["postgresql"]["bin_dir"], "pg_isready"))
|
||||
with patch('patroni.validator.open', mock_open(read_data='12')):
|
||||
schema(config)
|
||||
output = mock_out.getvalue()
|
||||
self.assertEqual([], parse_output(output))
|
||||
|
||||
@patch('subprocess.check_output', Mock(return_value=b"postgres (PostgreSQL) 12.1"))
|
||||
def test_pg_version_missmatch(self, mock_out, mock_err):
|
||||
directories.append(config["postgresql"]["data_dir"])
|
||||
directories.append(config["postgresql"]["bin_dir"])
|
||||
directories.append(os.path.join(config["postgresql"]["data_dir"], "pg_wal"))
|
||||
files.append(os.path.join(config["postgresql"]["data_dir"], "global", "pg_control"))
|
||||
files.append(os.path.join(config["postgresql"]["data_dir"], "PG_VERSION"))
|
||||
c = copy.deepcopy(config)
|
||||
c["etcd"]["hosts"] = []
|
||||
del c["postgresql"]["bin_dir"]
|
||||
with patch('patroni.validator.open', mock_open(read_data='11')):
|
||||
schema(c)
|
||||
output = mock_out.getvalue()
|
||||
self.assertEqual(['etcd.hosts', 'postgresql.data_dir'], parse_output(output))
|
||||
|
||||
@patch('subprocess.check_output', Mock(return_value=b"postgres (PostgreSQL) 12.1"))
|
||||
def test_pg_wal_doesnt_exist(self, mock_out, mock_err):
|
||||
directories.append(config["postgresql"]["data_dir"])
|
||||
directories.append(config["postgresql"]["bin_dir"])
|
||||
files.append(os.path.join(config["postgresql"]["data_dir"], "global", "pg_control"))
|
||||
files.append(os.path.join(config["postgresql"]["data_dir"], "PG_VERSION"))
|
||||
c = copy.deepcopy(config)
|
||||
del c["postgresql"]["bin_dir"]
|
||||
with patch('patroni.validator.open', mock_open(read_data='11')):
|
||||
schema(c)
|
||||
output = mock_out.getvalue()
|
||||
self.assertEqual(['postgresql.data_dir'], parse_output(output))
|
||||
|
||||
def test_data_dir_is_empty_string(self, mock_out, mock_err):
|
||||
directories.append(config["postgresql"]["data_dir"])
|
||||
directories.append(config["postgresql"]["bin_dir"])
|
||||
c = copy.deepcopy(config)
|
||||
c["kubernetes"] = False
|
||||
c["postgresql"]["pg_hba"] = ""
|
||||
c["postgresql"]["data_dir"] = ""
|
||||
c["postgresql"]["bin_dir"] = ""
|
||||
schema(c)
|
||||
output = mock_out.getvalue()
|
||||
self.assertEqual(['kubernetes', 'postgresql.bin_dir',
|
||||
'postgresql.data_dir', 'postgresql.pg_hba'], parse_output(output))
|
||||
Reference in New Issue
Block a user